Method and apparatus for executing user task, device, and medium

By acquiring images through robotic devices and using machine learning models to identify target objects and their coupling relationships, this technology solves the problem that existing robotic devices cannot decouple objects, enabling efficient task execution in complex environments.

WO2026016134A1PCT designated stage Publication Date: 2026-01-22BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/106245
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing robotic devices struggle to perform multiple tasks in complex environments according to user needs, especially in decoupling objects to obtain target objects.

Method used

By acquiring images of the physical space through robotic devices, machine learning models are used to identify target objects and decouple the relationships between objects, including language models and motion models, to determine the operation method and ensure safety and accuracy.

Benefits of technology

It improves the flexibility and accuracy of robotic equipment in complex environments, enabling it to complete user tasks more accurately, reduce the risk of object damage, and improve work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024106245_22012026_PF_FP_ABST
    Figure CN2024106245_22012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and apparatus for executing a user task, a device, and a medium. In the method, a user task from a user is received, the user task instructing a robotic device to acquire a first object; on the basis of a first image of a physical space in which the robotic device is located, the first object is positioned in the physical space; the robotic device moves to a position of the first object so as to acquire a second image of the first object; and, in response to determining, on the basis of the second image, that the first object cannot be acquired, the robotic device operates a second object so as to acquire the first object. By means of the exemplary implementation of the present disclosure, the robotic device can execute a user task in a complex physical space, and the flexibility and accuracy of the robotic device when executing a task in a complex environment can be improved, thus completing the expected user task.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices, and media for performing user tasks Technical Field

[0001] The exemplary implementations of this disclosure generally relate to the field of robotics, and more particularly to methods, apparatus, devices, and computer-readable storage media for using robots to perform user tasks. Background Technology

[0002] Robotics technology has developed rapidly and is widely used in many technological fields. Various specialized robotic devices have been developed; for example, in industrial environments, robots can perform a variety of tasks such as processing, grasping, sorting, and packaging. In home environments, for instance, robotic vacuum cleaners and window cleaning robots have been developed. However, robots typically can only perform pre-set, fixed tasks and cannot perform different user-defined tasks according to user needs.

[0003] Summary of the Invention

[0004] In a first aspect of this disclosure, a method for performing a user task is provided. In this method, a user task is received from a user, instructing a robotic device to acquire a first object. The first object is located in the physical space based on a first image of the physical space where the robotic device is located. The robotic device moves to the location of the first object to acquire a second image of the first object. In response to determining, based on the second image, that the first object cannot be acquired, the robotic device manipulates the second object to acquire the first object.

[0005] In a second aspect of this disclosure, an apparatus for performing a user task is provided. The apparatus includes: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to acquire a first object; a positioning module configured to locate the first object in a physical space based on a first image of the physical space where the robotic device is located; an acquisition module configured to move the robotic device to the location of the first object to acquire a second image of the first object; and an execution module configured to, in response to determining based on the second image that the first object cannot be acquired, cause the robotic device to manipulate the second object to acquire the first object.

[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.

[0008] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent in the following detailed description, taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;

[0012] Figure 2 shows a block diagram of some implementations of the present disclosure for performing user tasks;

[0013] Figure 3 shows a block diagram of an image acquisition process according to some implementations of this disclosure;

[0014] Figure 4 shows a flowchart of the process of invoking a language model according to some implementations of this disclosure;

[0015] Figure 5 shows a block diagram of the process of invoking the action model according to some implementations of this disclosure;

[0016] Figures 6A and 6B respectively show block diagrams of the process of obtaining an object according to some implementations of this disclosure;

[0017] Figure 7 shows a flowchart of a method for performing user tasks according to some implementations of this disclosure;

[0018] Figure 8 shows a block diagram of an apparatus for performing user tasks according to some implementations of the present disclosure; and

[0019] Figure 9 shows a block diagram of a device capable of implementing various implementations of the present disclosure. Detailed Implementation

[0020] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0021] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.

[0022] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0023] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0024] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0025] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0026] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0027] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.

[0028] Example Environment

[0029] In recent years, robotics and machine learning technologies have been widely applied in various scenarios. However, robots typically can only perform pre-set, fixed tasks and cannot execute different user tasks according to user needs. In particular, in complex application environments, robotic devices struggle to determine user requirements and thus perform corresponding tasks.

[0030] Simple robotic devices have been developed to perform specific tasks. However, these simple robotic devices cannot understand complex user instructions, nor can they execute the desired tasks according to user instructions within a complex physical space. Therefore, it is desirable to control the robot's operation in an effective way to perform the desired tasks.

[0031] According to an exemplary implementation of this disclosure, a method for performing user tasks is proposed. Referring to Figure 1, which describes an application environment according to an exemplary implementation of this disclosure, Figure 1 shows a block diagram 100 of the application environment according to an exemplary implementation of this disclosure. As shown in Figure 1, a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform various tasks. The physical space 160 can include, but is not limited to, one or more rooms. For example, in a home environment, the physical space 160 can include, but is not limited to, a living room, bedroom, study, kitchen, toilet, etc., or a combination of one or more of the above.

[0032] As shown in Figure 1, the robot device 110 may include multiple parts. For example, the control unit 111 can serve as the control center of the robot device 110, and an application can be loaded into the control unit 111 to control the various parts of the robot device. The user 120 can use the interaction unit 112 to interact with the robot device 110, for example, by inputting control commands to the robot device 110 to perform desired tasks. The robot device 110 may include an arm 113 for performing actions such as grasping and releasing. For example, the arm 113 can grasp an object and move it to a desired position, and so on.

[0033] Alternatively and / or additionally, the robot device 110 may also include a data acquisition unit 114. Here, the data acquisition unit 114 may include various types, such as an image acquisition unit, a sound acquisition unit, etc. Alternatively and / or additionally, the robot device 110 may further include a sensing unit for detecting surrounding objects, for example, detecting the distance between the robot and surrounding objects based on laser light, etc. The robot device 110 may also include a drive unit 115; for example, the robot device 110 may be deployed on a movable base, and the drive unit 115 may drive the wheels of the base to move along a desired path.

[0034] Physical environment 160 may include one or more acquisition units 130, ..., and 132. For example, one or more image acquisition devices may be deployed in a room to acquire images of the room from various angles. Physical environment 160 may include control device 140, which can control one or more acquisition units 130, ..., and 132, etc., via a network (not shown). Alternatively and / or additionally, in a smart home environment, control device 140 can control various electrical devices in physical space 160.

[0035] Alternatively and / or additionally, a machine learning model (e.g., model 150) may be provided to manage physical space 160. It should be understood that although Figure 1 shows model 150 located inside physical space 160, alternatively and / or additionally, model 150 may be located at a remote device outside physical space 160, and control device 140, robotic device 110, or other device may access the remote model 150 via a network.

[0036] Model 150 may include one or more models. If model 150 includes multiple models, these multiple models may include multiple types of models. Model 150 may, for example, include at least a language model (LM) and an action model. The language model, by learning from a large corpus, is capable of question answering. The action model can control the robotic device 110 to perform various actions. Model 150 may also include, for example, an image recognition model, a text recognition model, and so on.

[0037] As shown in Figure 1, user 120 can instruct robot device 110 to manipulate various objects in physical space 160110. Here, objects can be various items in the home environment. For example, user 120 can instruct robot device 110 to find a certain object in physical space 160; or user 120 can instruct robot device 110 to place the found object in a designated location, and so on.

[0038] Summary of the task to be performed

[0039] To at least partially address the shortcomings of the prior art, a method for performing user tasks is proposed according to an exemplary implementation of this disclosure. Referring to Figure 2, which describes an overview of an exemplary implementation of this disclosure, Figure 2 illustrates a block diagram 200 for performing user tasks according to some implementations of this disclosure.

[0040] As shown in Figure 2, the robot device 110 in the physical space 160 can receive a user task 210 from the user 120. The user task 210 can instruct the robot device 110 to retrieve a first object 220. For example, the user 120 can say in natural language, "Get me a plate," and in the example of Figure 2, the first object 220 is "plate." The robot device 110 can acquire a first image of the physical space 160 where the robot device 110 is located. For example, the robot device 110 can acquire the first image via at least one of the acquisition units 114, 130, ..., and 132.

[0041] Based on the first image, the position of the first object 220 in the physical space 160 is determined, and the robot device 110 is instructed to move to the position of the first object 220, thereby acquiring a second image related to the first object 220 at that position. For example, the first image can determine that the plate is located on a kitchen cabinet. Thus, the robot device 110 can be controlled to move to the front of the kitchen cabinet, and a second image can be acquired at the position in front of the cabinet via at least one of the acquisition units 114, 130, ..., and 132.

[0042] The second image includes a first object 220 and a potentially existing second object 230. The second object 230 may prevent the first object 220 from being accessed, in which case it is necessary to manipulate the second object 230 in order to access the first object 220. In the context of this disclosure, the second object 230 and the first object 220 may be coupled, resulting in the inability to access the first object.

[0043] Based on the coupling relationship between the first object 220 and the second object 230, the robot device can decouple the first object 220 and the second object 230 by manipulating the first object 220 and / or the second object 230 in order to acquire the first object 220. For example, if the robot device 110 confirms from the second image acquired that a coffee cup is also placed on the plate (i.e., there is a coupling relationship between the plate and the coffee cup), in this case, the robot device 110 can decouple the coffee cup from the plate by moving the coffee cup, so that the robot device can acquire the plate.

[0044] It should be understood that the coupling relationship described here can represent two objects coupled to each other, and the second object may affect the ability to obtain the first object independently. The coupling relationship described above is merely illustrative; alternatively and / or additionally, coupling relationships may include other situations. For example, if a plate contains other items, the plate cannot be obtained. Or, for example, suppose a user expects to obtain a clothes hanger; the clothes on the hanger are coupled to the hanger, thus preventing the hanger from being obtained, and so on. For ease of description, the following will only use the example of a second object being located above the first object to illustrate the situation where the first object cannot be obtained due to coupling.

[0045] According to some implementations of this disclosure, the methods described above can be executed on any computing device with computing capabilities. For example, the methods described above can be executed using an application deployed on robot device 110. Alternatively and / or additionally, an application can be deployed on control device 140 to execute the methods described above.

[0046] Using the exemplary implementation of this disclosure, a robotic device can travel to a predetermined location in physical space 160 associated with the first object 220 and manipulate the second object 230, that is, decouple the first object 220 and the second object 230 located at the predetermined location, thereby acquiring the first object 220. In the context of this disclosure, the purpose of manipulating the second object is to decouple the two objects. Here, alternatively, phrases such as "manipulating the second object," "decoupling the first object and the second object," and "decoupling" can be used. In this way, the robotic device can also acquire the first object 220 by decoupling the first object 220 and the second object 230. This improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby more accurately completing the expected user tasks.

[0047] Detailed process of executing the task

[0048] Having outlined some implementations according to this disclosure, further details regarding the execution of user tasks will be described below. For ease of description, the following description will use the example of controlling the robotic device 110 to retrieve a plate to illustrate the further details of the execution of user tasks.

[0049] According to some implementations of this disclosure, the first image can come from at least one of the following: an acquisition device at the robot device, or an acquisition device in the first physical space 160. Referring to FIG3 for further details of image acquisition, FIG3 shows a block diagram 300 of the image acquisition process according to some implementations of this disclosure. As shown in FIG3, a first image 310 (e.g., one or more images 310) of the physical space 160 can be acquired from the acquisition unit 114 at the robot device 110. Since the robot device 110 can move freely in the physical space 160, the acquisition unit 114 can acquire images of various locations in the physical space 160, thereby facilitating the search for the target object.

[0050] Alternatively and / or additionally, a first image 310 of the physical space 160 can be acquired from acquisition units 130, ..., and 132. Here, acquisition units 130, ..., and 132 can be pre-deployed at designated locations within the physical space 160, such as a corner of the ceiling, etc. In this way, an image of the physical space 160 taken from a top-down angle can be obtained, facilitating an overall understanding of the layout of the physical space 160 and thus facilitating the location of target objects.

[0051] According to some implementations of this disclosure, a first prompt word can be obtained based on a first image 310 and a first object 220. The first prompt word is used to locate the first object 220 in the first image 310. The prompt word can be, for example, expressed as: "Please identify the 'plate' from the following image," and the acquired first image 310 and the first prompt word are submitted to a machine learning model. After receiving the first image 310 and the prompt word, the machine learning model analyzes the first prompt word and processes the first image 310 to identify the first object 220 from the first image 310. If the first image 310 includes the first object 220, the machine learning model can output a response indicating the location of the first object 220 (e.g., the region coordinates of the first object 220 in the first image 310, and / or directly output the path to the location of the first object 220, etc.). If the first image 310 does not include the target object, the machine learning model can output a response such as "not found."

[0052] After receiving the first response from the machine learning model to the first prompt word, the first object 220 is located based on the first response. In this way, the machine learning model can help the robot device quickly and accurately determine the first object 220 in the first image 310 and further determine the position of the first object 220 in the physical space 160, thereby helping the robot device to quickly and accurately respond to the user's instructions and move to the location of the first object 220.

[0053] According to one embodiment of this disclosure, in response to determining that multiple first objects 220 exist in physical space 160, a first message is provided to the user to indicate the presence of multiple first objects 220 in physical space 160. For example, a machine learning model, through processing and analysis of a first image 310, determines that two plates exist in the first image 310, and sends a response to a robot device indicating that two plates exist in the first image 310. Upon receiving the response, the robot device provides the user with the first message to indicate the presence of two plates in physical space 160. For example, the robot device can present the response "There are two plates in the current scene" to the user via voice. In this way, the corresponding first object can be obtained according to the user's selection, thereby improving the accuracy of the robot device in obtaining objects.

[0054] Alternatively and / or additionally, the robotic device may also present the response to the user via images or videos. For example, an image including two plates may be displayed on the robotic device's screen. This image may be a magnified view of a first image 310, and the specific locations of the two plates may be marked in the image to indicate their positions to the user.

[0055] After receiving the first information, the user can provide a first response. Based on the user's first response to the first message, the first object 220 is located. The user's first response may include a description of the shape, color, position, and state of the first object 220. The first response can be acquired through the acquisition unit 114 of the robot device 110, the acquisition units 130, ..., and 132 of the physical space 160, and sent to the machine learning model. The machine learning model analyzes the first response and locates the first object 220 related to the first response from the first image 310. For example, the user can give the robot device a voice command, "Get me the red plate," and the machine learning model can identify the red plate in the first image 310 as the first object 220 by analyzing the voice command. The user can also give the robot device commands such as "Get me the plate on the left," "Get me the round plate," or "Get me the empty plate," and the machine learning model can determine the first object 220 from the first image 310 based on these commands and help the robot device locate the first object 220.

[0056] Users can also give an initial response to the robot device through actions. For example, when there are two plates in the physical space 160, the user can point to one of the plates to give an initial response. The acquisition unit 114 of the robot device 110, the acquisition units 130, ..., and 132 of the physical space 160 can obtain the initial response by capturing the user's actions.

[0057] According to one embodiment of this disclosure, a second prompt word is obtained based on a second image 320 and a first object 220. The second prompt word is used to determine the operation mode of the coupling relationship. For example, the second prompt word could be "how to get the plate". After obtaining the second prompt word, the machine learning model analyzes the second image 320 and generates a second response to decouple the first object 220 and the second object 230. After receiving the second response from the machine learning model for the second prompt word, the robot device determines the operation mode for the coupling relationship between the first object 220 and the second object 230. The robot device is instructed to decouple the relationship based on the operation mode. The second prompt word enables the machine learning model to further analyze the second image to obtain the operation mode for the coupling relationship between the first object 220 and the second object 230. This allows for rapid decoupling of the first object 220 and the second object 230, improving the response speed and service efficiency of the robot device.

[0058] In some specific implementations, a coupling relationship exists between a first object 220 (e.g., a plate) and a second object 230 (e.g., a coffee cup) in physical space 160 (e.g., the coffee cup is placed on the plate), and this coupling relationship is described in the acquired second image 320. A machine learning model can simulate operations on the first object 220 and the second object 230 in the second image 320 based on the second image 320 and a second prompt ("how to get the plate") obtained from the second image 320 and the first object 220, and output a second response to determine the operation method for decoupling the first object 220 and the second object 230. The robotic device controls the first object 220 and / or the second object 230 based on the operation method; for example, the robotic device can decouple the coffee cup from the plate by moving the coffee cup.

[0059] According to one embodiment of this disclosure, safety constraints related to the operating mode can also be determined, and the actions of the robot device can be determined based on the safety constraints and the operating mode. Installation constraints may include control constraints of the robot device on the first object 220 and / or the second object 230, and physical constraints between the robot device and other entities in the physical space 160. As an example, during the process of the robot device 110 grasping or releasing the first object 220 and / or the second object 230 via the arm 113, the robot can determine the material, weight, state, etc., of the first object 220 and / or the second object 230 based on the first image 310 and / or the second image 320. This allows for the determination of the gripping force of the arm 113, as well as physical quantities such as speed and acceleration when the arm 113 moves the first object 220 and / or the second object 230.

[0060] For example, the second object 230 that the robot needs to move is a coffee cup filled with coffee. Based on the machine learning model's assessment of the second image 320, it is determined that the robot should maintain the coffee cup's rim vertically upwards and move it with a low acceleration. As another example, the second object 230 that the robot needs to move is a banana. In this case, the machine learning model determines the gripping force of the robot's arm on the banana based on the assessment of the second image 320. Through these methods, the accuracy of the robot in moving the first object 220 and / or the second object 230 can be improved, and the risk of damage to the first object 220 and / or the second object 230 can be reduced.

[0061] Safety constraints may also include physical constraints between the robotic device (e.g., the robotic arm) and other entities in the physical space 160, such as the first object 220 and / or the second object 230. For example, a plate (first object 220) that the robotic device wants to pick up is partially obscured behind another object (e.g., a coffee pot). The machine learning model determines, using the second image 320, that the robotic arm needs to avoid the coffee pot when picking up the plate. As another example, when the robotic device moves a coffee cup placed on a plate, it needs to reposition the cup on the cabinet surface after picking it up and moving it. This process requires not only the robotic device to move the cup but also to control the cup to re-establish constraints with entities in the physical scene (e.g., the cabinet) (i.e., the coffee cup is placed on the cabinet).

[0062] According to one embodiment of this disclosure, the operation may include moving the first object 220 and / or the second object 230 to decouple the first object 220 and the second object 230. In some simple coupling relationships, such as a coffee cup placed on a saucer, the robotic device can grasp the coffee cup and move it vertically upward, horizontally, or vertically downward to place the coffee cup on a cabinet and decouple it from the saucer.

[0063] In some other ways disclosed herein, the coupling between the first object 220 and the second object 230 can also be decoupled by transforming the shape of the first object 220 and / or the second object 230. For example, the first object 220 and the second object 230 are fixed to each other by a snap-fit ​​mechanism. The robot can decouple from the first object 220 by manipulating the deformation of the second object 230, thereby acquiring the first object 220. By determining safety constraints on the operation method and determining the action based on these safety constraints, damage to the second object 230 during operation by the robot is reduced, improving the safety of the robot when acquiring objects and enhancing its adaptability to complex environments.

[0064] According to one embodiment of this disclosure, a second message can be provided to a user, indicating a coupling relationship between a first object 220 and a second object 230. A second response from the user to the second message is received. A robotic device is instructed to process the first object 220 based on the second response. A machine learning model determines the coupling relationship between the first object 220 and the second object 230 based on a second image 320, and the machine learning model can control the robotic device to send the second message to the user.

[0065] For example, a robot might issue a voice message, "There's a coffee cup on the plate, what should I do?" After receiving this second message, the user responds. The robot collects the second response through acquisition units 114, 130, ..., and 132, and then sends it to a machine learning model. The machine learning model processes the user's second response and sends the result back to the robot, enabling it to handle the first object 220 based on the user's second response. In this way, the robot can continuously handle problems based on the situation, and because it can request instructions from the user regarding the physical space, the accuracy of its work is improved.

[0066] According to one embodiment of this disclosure, in response to determining that a second answer indicates decoupling, the robot device is instructed to decouple in order to acquire the first object 220. In response to determining that a second answer indicates direct acquisition of the first object 220, the robot device is instructed to acquire the first object 220. After the robot device sends a second message to the user, and the user responds with a second answer indicating decoupling based on the second message, the robot device can decouple the first object 220 from the second object 230 based on the second answer. For example, if the robot device asks, "What should I do if there's a coffee cup on the plate?" and the user answers, "I want an empty plate," the robot device removes the coffee cup (i.e., decouples the coffee cup from the plate) and acquires the plate. If the user answers, "Bring the plate and the coffee cup together," the robot device can acquire both the plate and the coffee cup simultaneously based on this answer.

[0067] According to one embodiment of this disclosure, in response to determining a coupling relationship between a first object 220 and a second object 230 based on a second image 320, a candidate object for replacing the first object 220 can be determined. A third message is provided to the user to indicate the candidate object. A third response from the user to the third message can be received, and the candidate object can be selected as the first object 220. In some embodiments, the machine learning model confirms from the second image 320 that a coupling relationship exists between the first object 220 and the second object 230 (e.g., a coffee cup on a plate), and the machine learning model also discovers a candidate object near the first object 220 in the second image 320 (e.g., a bowl near the plate). The machine learning model can instruct the robot device to send a third message to ask the user, "There is a coffee cup on the plate, and an empty bowl next to the plate. Do you want the plate or the bowl?" In this way, the flexibility of the robot device in picking up the first object 220 can be improved, so that the robot device can make more flexible judgments in complex operating environments, thereby improving the success rate of the robot device in completing the task.

[0068] According to one embodiment of this disclosure, in response to determining that a first object 220 is in an unavailable state, a robotic device is instructed to acquire a third object to adjust the unavailable state. For example, a machine learning model determines the presence of a first object 220 (e.g., a plate) based on a second image 320, but the plate is currently wet. In this case, the machine learning model instructs the robotic device to acquire a rag or paper towel while acquiring the plate. Alternatively, a robot learning model may instruct the robotic device to acquire the first object 220 (e.g., wine), while simultaneously instructing the robotic device to acquire a corkscrew. Thus, by determining the current state of the first object 220 or considering its potential use case, more candidate objects can be acquired, thereby improving the intelligence of the robotic device's operation, reducing the number of times the user issues commands to the robotic device, and increasing the robot's efficiency.

[0069] It should be understood that the machine learning model can acquire images within the physical space 160 through acquisition units 114, 130, ..., 132 before instructing the robotic device to pick up candidate objects (such as tissues, bottle openers, etc.) to search for potential candidate objects within the physical space 160. Of course, the machine learning model can also determine candidate objects from the first image 310 or the second image 320. If the machine learning model identifies a second object 230 in the current physical space 160 that is associated with the first object 220, the machine learning model can instruct the robotic device to pick up the candidate object.

[0070] According to one embodiment of this disclosure, if the machine device acquires the first object 220 and the candidate object (if it exists), it instructs the robot device to move the first object 220 and the candidate object (if it exists) to the user's location.

[0071] According to one embodiment of this disclosure, the machine learning model may include a language model and an action model. The language model is a trained and fine-tuned model with rich knowledge of performing tasks across multiple domains. It can process related questions based on prompts, images, and other information, and output processing results. The action model is adapted to determine the specific actions to be performed by the robot device based on the processing results of the language model, so that the robot device performs the corresponding actions. The language model and action model will be described in more detail below.

[0072] Referring to Figure 4 for further details, Figure 4 shows a block diagram 400 of the process of invoking a language model according to some implementations of this disclosure. As shown in Figure 4, the first object 220 to be acquired can be determined from the user task 210 as a "plate". At this time, a prompt word 410 can be generated based on the first image 310 and the first object 220. The prompt word 410 can be expressed as: "Locate the plate in the following image", etc. The first prompt word 410 and the first image 310 can be input to the language model 420 so that the language model 420 can find the plate from the image 310.

[0073] According to some implementations of this disclosure, the language model 420 is a trained and fine-tuned model with rich knowledge of performing tasks across multiple domains. Figure 4 is merely illustrative; the language model 420 can process one or more images from different acquisition devices and identify a first object 220, a second object 230, candidate objects, etc. For example, the language model 420 may identify a second object 230 from the second image 320 that is coupled with the first object 220, or the language model 420 may identify candidate objects from the second image 320 that are related to the first object 220 in terms of usage.

[0074] According to some implementations of this disclosure, fine-tuning operations can be performed on the model. For example, the robot device can be instructed to pre-collect images of various parts of the physical space 160, and then the collected images can be used to fine-tune the model. In this way, the model can determine the specific storage location of each object.

[0075] By utilizing some implementation methods of this disclosure, the powerful processing capabilities and rich knowledge of the model can be leveraged to identify the first object 220, the second object 230, and candidate objects. Compared to existing technical solutions that can only search for target objects in physical space 160 through image recognition, the technical solution of this disclosure can obtain the first object 220 that is coupled with the second object 230 in physical space 160. Furthermore, the technical solution of this disclosure can also associate the current state and usage scenario factors of the first object 220 to further obtain candidate objects related to the first object 220. In this way, the efficiency of obtaining the first object 220 can be improved, thereby improving the overall efficiency of performing user tasks.

[0076] According to some implementations of this disclosure, a knowledge base can be used to determine the association between candidate objects and the first object 220. Specifically, in determining the association between candidate objects and the first object 220, multiple objects can be identified from images in the physical space 160 (e.g., the first image 310 and the second image 320), and candidate objects can be selected from these multiple objects based on the knowledge base. Here, the knowledge base can be predefined and includes the first object 220 and multiple candidate objects that are associated with the first object 220. For example, the knowledge base can include: (bowl, chopsticks), (red wine, corkscrew), etc.

[0077] According to some implementations of this disclosure, a motion model can be used to determine the specific actions to be performed by the robot device. See Figure 5 for further details, which shows a block diagram 500 illustrating the process of invoking a motion model according to some implementations of this disclosure. As shown in Figure 5, a motion model 530 can be provided, which can determine the specific actions to be performed by the robot device based on the current state and instructions of the robot device. This motion model can be a pre-trained and fine-tuned model.

[0078] It should be understood that the current state may include data from multiple aspects, such as an image of the robot device, an image of the robot device's environment, pose data of the robot arm (e.g., the positions of the robot arm's joints (POS1, ...)), and the state of the tool (e.g., a gripper, a cutting tool, etc.) fixed to the end of the robot arm. For example, 0 can be used to represent the gripper's closed state, and 1 can be used to represent the gripper's open state. Instructions and the current state can be input into the motion model 530, which then uses the motion model to determine the action to be performed by the robot device based on the instructions and the current state. Here, the action can represent the difference between the robot device's current pose and the next pose, and the difference between the tool's current state and the next state, etc.

[0079] An instruction 510 (e.g., "remove the coffee cup") can be input to the motion model 530. Here, the instruction 510 can be expressed in natural language, and the instruction 510 can be determined from the response of the language model. Furthermore, the current state of the robot device can be obtained, and the motion model 530 can determine the corresponding action 540 based on the input data. For example, the orientation, position, speed, acceleration, etc., of each joint in the arm, and / or the wheels and / or other movable devices of the robot device at the next time point can be determined. Furthermore, the determined action 540 can be used to control the state of the robot device at the next time point.

[0080] Using some implementation methods disclosed herein, a relationship can be established between the language model and the action model, and the user's initial input, expressed in natural language, can be converted into specific actions that can be performed by the robotic device. In this way, the actions of the robotic device can be precisely controlled, thereby executing the user task with higher efficiency.

[0081] According to some implementations of this disclosure, more details are described with reference to Figures 6 and 7B. Figure 6A shows a schematic diagram of the coupling relationship between the first object 220 and the second object 230 according to some implementations of this disclosure, and Figure 6B shows a schematic diagram of decoupling the first object 220 and the second object 230 according to some embodiments of this disclosure.

[0082] After the user gives a second response instructing the first object 220 and the second object 230 to decouple, instructions related to the second response can be input into the action model 530. Furthermore, the action model 530 can generate new actions to further decouple the first object 220 and the second object 230. Using some implementation methods of this disclosure, new instructions and states can be continuously input into the action model to determine subsequent actions.

[0083] According to some implementations of this disclosure, constraint 720, i.e., the constraints that should be followed during the execution of an action, can be determined. For example, it can be determined that the original posture of the coffee cup should be maintained during the movement of the coffee cup (e.g., keeping the rim upward and preventing it from tilting). Constraint 720 can be determined using a language model. For example, prompts can be generated such as, "Based on the following image, determine the constraints that should be followed during the movement of the coffee cup," or "Please determine the precautions during the movement of the coffee cup," etc. In this case, the language model can generate the corresponding constraints. The constraints can be input into the action model 530, and the series of actions output by the action model 530 will move the coffee cup to decouple it from the saucer while ensuring the constraints are met. Using some implementations of this disclosure, the safety of the robot device during operation can be ensured, thereby avoiding accidental damage to an object, etc.

[0084] According to some implementations of this disclosure, in order to decouple the first object 220 and the second object 230, the second object 230 can be removed first, and then the first object 220 can be retrieved. As shown in Figure 6B, according to some implementations of this disclosure, a target position can be determined, and the robot device can be instructed to move the second object 230 to the target position. At this time, the motion model will generate an action to control the robot device to move the second object 230 from position 760 to position 760'. In this way, the robot device can be supported in handling complex problems in complex environments, thereby performing user tasks in a more accurate manner.

[0085] It should be understood that although the foregoing description uses a Chinese language environment as an example to illustrate an exemplary implementation of this disclosure, alternatively and / or additionally, the technical solution of the exemplary implementation of this disclosure can be executed in multiple language environments. For example, the robot can be controlled in environments such as Chinese, English, Japanese, and French. Specifically, the multilingual capabilities provided by machine learning technology can be used to control the robot in application environments of different languages. Furthermore, although the foregoing description uses picking up a plate as an example to illustrate the process of using a robotic device to perform a user task, alternatively and / or additionally, the robotic device can be controlled to perform other user tasks, such as finding other items in a room, placing an item in a designated location, etc.

[0086] According to some implementations of this disclosure, various positioning algorithms can be used to determine the position of robotic devices and various objects in the physical environment. For example, a Global Positioning System (GPS) can be deployed at the robotic device, and satellite signals can be used to determine the precise position of the robotic device. Alternatively and / or additionally, a communication unit can be deployed at the robotic device, and the position of the robotic device can be determined by means of signals between the communication unit and a base station and by utilizing a communication network. Alternatively and / or additionally, a Wi-Fi access point can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspot to determine the position via Wi-Fi signal strength and the known location of the Wi-Fi access point. Alternatively and / or additionally, the communication unit at the robotic device can support Bluetooth functionality, in which case Bluetooth signals and the known locations of Bluetooth devices can be used to determine the position of nearby devices. An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, thereby determining the position of the robotic device.

[0087] Alternate and / or additional locations can be determined using a visual positioning system to pinpoint the location of the robotic device and / or individual objects. A map of the physical space can be pre-acquired, and the locations of each object can be marked on this map. The robotic device can utilize echo detection units to detect distances to surrounding objects and, by combining the acquired images with the physical space map, determine the precise location of each object. Specifically, computer-aided design (CAD) and geographic information systems (GIS) can be used, along with positioning algorithms to determine the location. Alternate and / or additional locations can also be used to deploy tracking units at important objects in the physical space; for example, tracking units can be added to remote controls for household appliances (e.g., television remotes, air conditioner remotes) so that the robotic device can promptly acquire the precise location of important objects, and so on.

[0088] According to some implementations of this disclosure, the robot's initial position and desired destination can be determined based on the methods described above. The robot can determine a path from its initial position to its destination. For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, continuously update the path, enabling the robot to move along the path to its destination.

[0089] According to some implementations of this disclosure, after reaching the destination location, the robotic device can perform a specified task. For example, it can acquire a specified object and move it to the appropriate location. Constraints, i.e., the constraints that should be followed during task execution, can be determined using a language model and / or a knowledge base. For example, an image and corresponding prompts can be acquired, and the image and prompts can be input into the language model, thereby receiving the constraints from the language model. For example, prompts can be determined as: "Based on the following image, determine the constraints that should be followed during the movement of object XXX," or "Please determine the precautions during the movement of object XXX," etc.

[0090] At this point, it can be determined that during the movement of an object (e.g., bottled water, plate, bowl, etc.), the object's original posture should be maintained (e.g., remaining vertical and not tilted). Furthermore, constraints can be input into the motion model, at which point the series of actions output by the motion model will perform the corresponding tasks while ensuring the constraints are met. Using some implementation methods of this disclosure, safety during the operation of robotic devices can be ensured, thereby preventing accidental damage to an object, and so on.

[0091] Alternatively and / or additionally, a user can interact with the robot device via the interaction unit 112, for example, the user inputs a task represented by text and / or images, and controls the robot device to perform the task. Alternatively and / or additionally, the user can specify the execution conditions of the task, for example, to execute the task immediately, to execute the task after a predetermined time, or to execute the task when predetermined conditions are determined to be met (e.g., after the user wakes up), etc.

[0092] Example process

[0093] Figure 7 illustrates a flowchart of a method 700 for performing a user task according to some implementations of the present disclosure. Specifically, at block 710, a user task is received from a user, instructing a robot device to acquire a first object; at block 720, the first object is located in the physical space based on a first image of the physical space where the robot device is located; at block 730, the robot device moves to the location of the first object to acquire a second image of the first object; and at block 740, in response to determining that the first object cannot be acquired based on the second image, the robot device manipulates the second object to acquire the first object.

[0094] According to some implementations of this disclosure, locating the first object includes: obtaining a first prompt word based on a first image and the first object, the first prompt word being used to locate the first object in the first image; and receiving a first response from a machine learning model to the first prompt word in order to locate the first object.

[0095] According to some implementations of this disclosure, locating the first object includes: identifying the first object from the first image.

[0096] According to some implementations of this disclosure, locating the first object includes: in response to determining that multiple first objects exist in the physical space, providing a first message to the user to indicate that multiple first objects exist in the physical space; and locating the first object based on a first response from the user to the first message.

[0097] According to some implementations of this disclosure, operating the second object includes: obtaining a second prompt word based on the second image and the first object, the second prompt word being used to determine the operation mode of the second object; determining the operation mode based on a second response to the second prompt word from a machine learning model; and the robot device operating the second object based on the operation mode.

[0098] According to some implementations of this disclosure, operating the second object further includes: determining safety constraints associated with the operation mode; determining the actions of the robot device based on the safety constraints and the operation mode; and the robot device operating the second object according to the actions.

[0099] According to some implementations of this disclosure, the method 700 further includes: providing a second message to a user, the second message indicating that the first object cannot be obtained; receiving a second response from the user to the second message; and the robotic device processing the first object based on the second response.

[0100] According to some implementations of this disclosure, the method 700 further includes at least one of the following: in response to determining that a second answer indicates operation on a second object, the robotic device operates the second object to obtain a first object; and in response to determining that a second answer indicates direct acquisition of the first object, the robotic device acquires the first object.

[0101] According to some implementations of this disclosure, the method 700 further includes: in response to determining that a first object cannot be obtained based on a second image, determining a candidate object to replace the first object; providing a third message to the user to indicate the candidate object; and receiving a third response from the user to the third message, and selecting the candidate object as the first object.

[0102] According to some implementations of this disclosure, the method 700 further includes: in response to determining that the first object is in an unavailable state, the robotic device acquires a third object for adjusting the unavailable state.

[0103] According to some implementations of this disclosure, the method 700 further includes: the robotic device moving the first object to the user's user location.

[0104] Example devices and equipment

[0105] Figure 8 shows a block diagram of an apparatus 800 for performing a user task according to some implementations of the present disclosure. The apparatus 800 includes: a receiving module 810 configured to receive a user task from a user, the user task instructing a robot device to acquire a first object; a positioning module 820 configured to locate the first object in a physical space based on a first image of the physical space where the robot device is located; an acquisition module 830 configured to move the robot device to the location of the first object to acquire a second image of the first object; and an execution module 840 configured to, in response to determining based on the second image that the first object cannot be acquired, cause the robot device to operate on the second object to acquire the first object.

[0106] According to some implementations of this disclosure, the positioning module 820 is further configured to: obtain a first prompt word based on a first image and a first object, the first prompt word being used to locate the first object in the first image; and receive a first response from a machine learning model to the first prompt word in order to locate the first object.

[0107] According to some implementations of this disclosure, the positioning module 820 is further configured to: identify a first object from a first image.

[0108] According to some implementations of this disclosure, the positioning module 820 is further configured to: provide a first message to a user in response to determining that multiple first objects exist in the physical space to indicate that multiple first objects exist in the physical space; and locate the first objects based on a first response from the user to the first message.

[0109] 5. The method of claim 1, wherein the execution module 840 is further configured to: obtain a second prompt word based on the second image and the first object, the second prompt word being used to determine the operation mode of the second object; determine the operation mode based on a second response of the second prompt word using a machine learning model; and cause the robot device to operate the second object based on the operation mode.

[0110] According to some implementations of this disclosure, the execution module 840 is further configured to: determine safety constraints associated with the operation mode; determine the actions of the robot device based on the safety constraints and the operation mode; and have the robot device operate the second object according to the actions.

[0111] According to some implementations of this disclosure, it further includes an interaction module configured to: provide a user with a second message indicating that the first object cannot be obtained; receive a second response from the user to the second message; and have the robotic device process the first object based on the second response.

[0112] According to some implementations of this disclosure, the execution module 840 is further configured to perform at least one of the following: in response to determining that a second answer indicates to operate on a second object, causing the robot device to operate on the second object in order to acquire the first object; or in response to determining that a second answer indicates to directly acquire the first object, causing the robot device to acquire the first object.

[0113] According to some implementations of this disclosure, the interaction module is further configured to: in response to determining that the first object cannot be obtained based on the second image, determine a candidate object to replace the first object; provide a third message to the user to indicate the candidate object; and receive a third response from the user to the third message, and select the candidate object as the first object.

[0114] According to some implementations of this disclosure, the execution module 840 is further configured to: in response to determining that the first object is in an unavailable state, cause the robot device to acquire a third object for adjusting the unavailable state.

[0115] According to some implementations of this disclosure, the execution module 840 is further configured to: move the first object to the user's user location using the robotic device.

[0116] Figure 9 shows a block diagram of a device 900 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 900 shown in Figure 9 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 900 shown in Figure 9 can be used to implement the methods described above.

[0117] As shown in Figure 9, the computing device 900 is in the form of a general-purpose computing device. Components of the computing device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage devices 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 900.

[0118] Computing device 900 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 920 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 930 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within computing device 900.

[0119] The computing device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 920 may include a computer program product 925 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.

[0120] The communication unit 940 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 900 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0121] Input device 950 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 960 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 900 can also communicate as needed with one or more external devices (not shown) via communication unit 940. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 900, or with any device (e.g., network card, modem, etc.) that enables computing device 900 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).

[0122] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.

[0123] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0124] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0125] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0127] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for performing a user task, comprising: receiving a user task from a user, the user task instructing a robotic device to acquire a first object; locating the first object in a physical space where the robotic device is located based on a first image of the physical space; the robotic device moving to a location of the first object in order to acquire a second image of the first object; and in response to determining that the first object cannot be acquired based on the second image, the robotic device operating a second object in order to acquire the first object.

2. The method of claim 1, wherein locating the first object comprises: acquiring a first cue for locating the first object in the first image based on the first image and the first object; and receiving a first response of a machine learning model to the first cue in order to locate the first object. identifying the first object from the first image.

4. The method of claim 1, wherein locating the first object comprises:

3. The method of claim 1, wherein positioning the first object comprises: in response to determining that a plurality of first objects exist in the physical space, providing a first message to the user to indicate that a plurality of first objects exist in the physical space; and locating the first object based on a first answer from the user to the first message.

5. The method of claim 1, wherein operating the second object comprises: acquiring a second cue for determining a manner of operation of the second object based on the second image and the first object; determining the manner of operation based on a second response of a machine learning model to the second cue; and the robotic device operating the second object based on the manner of operation.

6. The method of claim 5, wherein operating the second object further comprises: determining a safety constraint associated with the manner of operation; determining an action of the robotic device based on the safety constraint and the manner of operation; and the robotic device operating the second object in accordance with the action.

7. The method of claim 1, further comprising: providing a second message to the user, the second message indicating that the first object cannot be acquired; receiving a second answer of the user to the second message; and the robotic device processing the first object based on the second answer.

8. The method of claim 7, further comprising at least either of: in response to determining that the second answer indicates operating the second object, the robotic device operating the second object in order to acquire the first object; and in response to determining that the second answer indicates acquiring the first object directly, the robotic device acquiring the first object.

9. The method of claim 1, further comprising: in response to determining that the first object cannot be acquired based on the second image, determining a candidate object for replacing the first object; providing a third message to the user to indicate the candidate object; and ​ ​ ​ ​ ​ ​ receiving a third reply from the user for the third message, selecting the candidate object as the first object.

10. The method of claim 1, further comprising: in response to determining that the first object is in an unavailable state, the robotic device obtaining a third object for adjusting the unavailable state.

11. The method of claim 1, further comprising: the robotic device moving the first object to a user location of the user.

12. An apparatus for performing a user task, comprising: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to obtain a first object; a positioning module configured to position the first object in a physical space in which the robotic device is located based on a first image of the physical space; an obtaining module configured to move the robotic device to a location of the first object in order to obtain a second image of the first object; and an execution module configured to cause the robotic device to operate the second object in order to obtain the first object in response to determining that the first object cannot be obtained based on the second image.

13. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Service robot grasp guidance system and method thereof

    CN101559600A

  • Article holding system, robot, and method of controlling robot

    CN101890720A

  • Intelligent food material identifying and grabbing robot and control method thereof

    CN112318514A

  • Article grabbing planning method and system

    CN114102585A

  • Sundry cleaning robot system

    CN116709962A