Method and apparatus for executing user task, and device and medium
By acquiring and analyzing physical space images and using language and motion models to determine potential physical spaces and actions, robotic devices can find and obtain objects that meet user needs in complex environments. This solves the problem that existing robotic devices cannot perform multiple tasks and achieves more efficient task execution.
Patent Information
- Application Number
- PCT/CN2024/106244
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-22
AI Technical Summary
Existing robotic devices struggle to perform multiple tasks in complex environments according to user needs, especially when the target object cannot be directly located, making it difficult to determine the potential physical space and acquire the object.
The robot acquires an image of the first physical space using a robotic device. It then uses a language model and image recognition technology to determine whether the target object is present. If the target object is not present, the robot is directed to the second physical space. The robot then accesses this space to retrieve the object and controls its movements using a motion model to complete the task.
It improves the flexibility and accuracy of robotic devices in complex environments, enabling them to locate target objects in potential physical spaces and complete user tasks safely and efficiently.
Smart Images

Figure CN2024106244_22012026_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices, and media for performing user tasks Technical Field
[0001] The exemplary implementations of this disclosure generally relate to the field of robotics, and more particularly to methods, apparatus, devices, and computer-readable storage media for using robots to perform user tasks. Background Technology
[0002] Robotics technology has developed rapidly and is widely used in many technological fields. Various specialized robotic devices have been developed; for example, in industrial environments, robots can perform a variety of tasks such as processing, grasping, sorting, and packaging. In home environments, for instance, robotic vacuum cleaners and window cleaning robots have been developed. However, robots typically can only perform pre-set, fixed tasks and cannot perform different user-defined tasks according to user needs.
[0003] Summary of the Invention
[0004] In a first aspect of this disclosure, a method for performing a user task is provided. In this method, a user task is received from a user, which in turn tasks a robotic device to acquire a first object. A first image of a first physical space in which the robotic device is located is acquired. In response to determining that the first image indicates the first physical space does not contain the first object, a second physical space is determined. The robotic device accesses the second physical space to acquire the first object.
[0005] In a second aspect of this disclosure, an apparatus for performing a user task is provided. The apparatus includes: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to acquire a first object; an acquiring module configured to acquire a first image of a first physical space in which the robotic device is located; a determining module configured to determine a second physical space in response to determining that the first image indicates the first physical space does not contain the first object; and an executing module configured to cause the robotic device to access the second physical space in order to acquire the first object.
[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.
[0008] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] In the following detailed description, the above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent, taken in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;
[0012] Figure 2 shows a block diagram of some implementations of the present disclosure for performing user tasks;
[0013] Figure 3 shows a block diagram of an image acquisition process according to some implementations of this disclosure;
[0014] Figure 4 shows a flowchart of the process of invoking a language model according to some implementations of this disclosure;
[0015] Figure 5 shows a block diagram of a process for identifying objects from an image according to some implementations of this disclosure;
[0016] Figure 6 shows a block diagram of the process of invoking the action model according to some implementations of this disclosure;
[0017] Figures 7A and 7B respectively show block diagrams of the process of obtaining an object according to some implementations of this disclosure;
[0018] Figure 8 shows a flowchart of a method for performing user tasks according to some implementations of this disclosure;
[0019] Figure 9 shows a block diagram of an apparatus for performing user tasks according to some implementations of the present disclosure; and
[0020] Figure 10 shows a block diagram of a device capable of implementing various implementations of the present disclosure. Detailed Implementation
[0021] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0022] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.
[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0025] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0026] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0027] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0028] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
[0029] Example Environment
[0030] In recent years, robotics and machine learning technologies have been widely applied in various scenarios. However, robots typically can only perform pre-set, fixed tasks and cannot execute different user tasks according to user needs. In particular, in complex application environments, robotic devices struggle to determine user requirements and thus perform corresponding tasks.
[0031] Simple robotic devices have been developed to perform specific tasks. However, these simple robotic devices cannot understand complex user instructions, nor can they execute the desired tasks according to user instructions within a complex physical space. Therefore, it is desirable to control the robot's operation in an effective way to perform the desired tasks.
[0032] According to an exemplary implementation of this disclosure, a method for performing user tasks is proposed. Referring to Figure 1, which describes an application environment according to an exemplary implementation of this disclosure, Figure 1 shows a block diagram 100 of the application environment according to an exemplary implementation of this disclosure. As shown in Figure 1, a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform various tasks. The physical space 160 can include, but is not limited to, one or more rooms. For example, in a home environment, the physical space 160 can include, but is not limited to, a living room, bedroom, study, kitchen, toilet, etc., or a combination of one or more of the above.
[0033] As shown in Figure 1, the robot device 110 may include multiple parts. For example, the control unit 111 can serve as the control center of the robot device 110, and an application can be loaded into the control unit 111 to control the various parts of the robot device. The user 120 can use the interaction unit 112 to interact with the robot device 110, for example, by inputting control commands to the robot device 110 to perform desired tasks. The robot device 110 may include an arm 113 for performing actions such as grasping and releasing. For example, the arm 113 can grasp an object and move it to a desired position, and so on.
[0034] Alternatively and / or additionally, the robot device 110 may also include a data acquisition unit 114. Here, the data acquisition unit 114 may include various types, such as an image acquisition unit, a sound acquisition unit, etc. Alternatively and / or additionally, the robot device 110 may further include a sensing unit for detecting surrounding objects, for example, detecting the distance between the robot and surrounding objects based on laser light, etc. The robot device 110 may also include a drive unit 115; for example, the robot device 110 may be deployed on a movable base, and the drive unit 115 may drive the wheels of the base to move along a desired path.
[0035] Physical environment 160 may include one or more acquisition units 130, ..., and 132. For example, one or more image acquisition devices may be deployed in a room to acquire images of the room from various angles. Physical environment 160 may include control device 140, which can control one or more acquisition units 130, ..., and 132, etc., via a network (not shown). Alternatively and / or additionally, in a smart home environment, control device 140 can control various electrical devices in physical space 160.
[0036] Alternatively and / or additionally, a machine learning model (e.g., model 150) may be provided to manage physical space 160. It should be understood that although Figure 1 shows model 150 located inside physical space 160, alternatively and / or additionally, model 150 may be located at a remote device outside physical space 160, and control device 140, robotic device 110, or other device may access the remote model 150 via a network.
[0037] Model 150 may include one or more models. If model 150 includes multiple models, these multiple models may include multiple types of models. Model 150 may, for example, include at least a language model (LM) and an action model. The language model, by learning from a large corpus, is capable of question answering. The action model can control the robotic device 110 to perform various actions. Model 150 may also include, for example, an image recognition model, a text recognition model, and so on.
[0038] As shown in Figure 1, user 120 can instruct robot device 110 to manipulate various objects in physical space 110. Here, objects can be various items in the home environment. For example, user 120 can instruct robot device 110 to find a certain object in physical space 160; or user 120 can instruct robot device 110 to place the found object in a designated location, and so on.
[0039] Summary of the task to be performed
[0040] To at least partially address the shortcomings of the prior art, a method for performing user tasks is proposed according to an exemplary implementation of this disclosure. Referring to Figure 2, which describes an overview of an exemplary implementation of this disclosure, Figure 2 illustrates a block diagram 200 for performing user tasks according to some implementations of this disclosure.
[0041] As shown in Figure 2, the robot device 110 in physical space 160 (also referred to as the first physical space) can receive a user task 210 from user 120. In this case, user task 210 can instruct or control the robot device 110 to acquire a first object (which can be referred to as the target object for ease of description). For example, user 120 can say in natural language, "Get me a bottle of water," and in the example of Figure 2, the first object is "a bottle of water" (e.g., object 220). The robot device 110 can acquire a first image of the first physical space where the robot device is located. For example, the first image can be acquired via at least any one of acquisition units 114, 130, ..., and 132.
[0042] In response to determining that the first image represents a first physical space that does not include the first object, a second physical space can be determined. Here, the second physical space can be another physical space that may include the first object. For example, in a home environment, since bottled water might be placed in a refrigerator, kitchen, or other location, the second physical space can be determined to be the refrigerator and / or kitchen. As shown in Figure 2, assuming the room includes a refrigerator, the refrigerator can be determined as the second physical space (i.e., physical space 230). At this time, the robot device 110 accesses the second physical space to obtain the first object.
[0043] According to some implementations of this disclosure, the methods described above can be executed on any computing device with computing capabilities. For example, the methods described above can be executed using an application deployed on robot device 110. Alternatively and / or additionally, an application can be deployed on control device 140 to execute the methods described above. Specifically, the powerful processing capabilities of model 150 can be invoked to locate the physical space 230 that may include object 220. Robot device 110 can then proceed to the refrigerator, retrieve object 220, return along path 240 to the location of user 120, and provide the retrieved object 220 to user 120.
[0044] Using the exemplary implementations of this disclosure, robotic devices can perform user tasks in complex physical spaces. In this way, even if the target object cannot be directly found in the physical space, the robotic device can search for potential physical spaces that may contain the target object, thereby finding the desired object. This improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thus completing the intended user task.
[0045] Detailed process of executing the task
[0046] Having outlined some implementations according to this disclosure, further details regarding the execution of user tasks will be described below. For ease of description, the following example uses the control of robotic device 110 to retrieve bottled water as an illustration to illustrate further details of the execution of user tasks.
[0047] According to some implementations of this disclosure, the first image can come from at least one of the following: an acquisition device at the robot device, an acquisition device in a first physical space, or an acquisition device in a second physical space. Referring to Figure 3 for further details of image acquisition, Figure 3 shows a block diagram 300 of the image acquisition process according to some implementations of this disclosure. As shown in Figure 3, a first image (e.g., one or more images 310) of the physical space 160 can be acquired from the acquisition unit 114 at the robot device 110. Since the robot device 110 can move freely in the physical space 160, the acquisition unit 114 can acquire images from various locations in the physical space, thereby facilitating the search for the target object.
[0048] Alternatively and / or additionally, a first image of the physical space 160 can be acquired from acquisition units 130, ..., and 132. Here, acquisition units 130, ..., and 132 can be pre-deployed at designated locations within the physical space 160, such as a corner of the ceiling, etc. In this way, an image of the physical space 160 taken from a top-down angle can be obtained, facilitating an overall understanding of the layout of the physical space 160 and thus facilitating the location of target objects.
[0049] According to some implementations of this disclosure, whether a first image includes a first object can be determined in various ways. For example, the first object can be identified from the first image based on image recognition technology. Alternatively and / or additionally, a prompt word can be constructed and input into the model to invoke the model's processing capabilities to identify the first object from the first image. The prompt word can be, for example, represented as: "Please identify 'bottled water' from the following image," and the collected image and prompt word can be submitted to the model.
[0050] The model can process images, and if the image includes a first object, the model can output the location of the object (e.g., the coordinates of the object's region in the image, and / or directly output the image of the region where the object is located, etc.). If the image does not include the target object, the model can output a response such as "not found". Using some implementations of this disclosure, the presence or absence of a target object in an image can be detected in multiple ways, thereby improving the performance of robotic devices in acquiring target objects.
[0051] According to some implementations of this disclosure, in response to determining that the first image represents a first physical space that does not include the first object, a second physical space is determined. It should be understood that the second physical space here is a potential physical space that may include the first object. For example, in a home environment, since bottled water may be placed in a refrigerator or kitchen, the second physical space can be determined to be the refrigerator and / or the kitchen. Specifically, in determining the second physical space, based on the first image and the first object, cue words for locating the first object can be obtained; and the response of a machine learning model to the cue words can be received to determine the second physical space.
[0052] Referring to Figure 4 for further details regarding the determination of the second physical space, Figure 4 illustrates a block diagram 400 of the process of invoking a language model according to some implementations of this disclosure. As shown in Figure 4, the object 220 to be acquired can be determined from the user task 210 to be "bottled water". At this point, a corresponding prompt 410 can be generated based on the image 310 and the object 220. The prompt 410 can be expressed, for example, as: "Please determine the physical space that may include 'bottled water' from the following images", or as: "Where might 'bottled water' be placed in the following images", and so on. The prompt 410 and the image 310 can be input into the language model 420 so that the language model 420 can find the second physical space that may include bottled water from the image 310.
[0053] According to some implementations of this disclosure, the language model 420 is a trained and fine-tuned model with rich knowledge of performing tasks across multiple domains. The language model 420 can determine that image 310 includes a refrigerator and identify the refrigerator as a second physical space. Figure 4 is merely illustrative; the language model 420 can process one or more images from different acquisition devices and find one or more second physical spaces that may include bottled water. For example, suppose another image includes a locker, and the locker can be identified as a second physical space. Alternatively and / or additionally, suppose the user task is "find a fruit knife," the model can determine that the fruit knife may be placed in a drawer of a cabinet, in which case the drawer can be identified as a second physical space.
[0054] According to some implementations of this disclosure, fine-tuning operations can be performed on the model. For example, the robot device can be instructed to pre-collect images of various parts in the first physical space. For instance, the robot device 110 can open a refrigerator (or cabinet, drawer), etc., and collect images of various objects inside the refrigerator. The collected images can then be used to fine-tune the model. In this way, the model can determine the specific location of each object, thereby improving the accuracy of determining the second physical space.
[0055] By utilizing some implementation methods of this disclosure, the powerful processing capabilities and rich knowledge of the model can be leveraged to determine a second physical space that may include the first object. Subsequently, the robotic device can be instructed to proceed to the second physical space to continue searching for the target object. Compared to existing technologies that can only search for target objects in visible physical space through image recognition, the technical solution of this disclosure can search for target objects in more potential physical spaces, thereby finding hidden objects not directly exposed to the coverage of the acquisition device. In this way, the efficiency and accuracy of locating the first object can be improved, thereby increasing the overall efficiency of performing user tasks.
[0056] According to some implementations of this disclosure, a second physical space can be determined based on image recognition. Specifically, in determining the second physical space, multiple second objects can be identified from a first image, a second object can be selected from the multiple second objects based on a knowledge base, and the space associated with the second object can be used as the second physical space. According to some implementations of this disclosure, the associated space can include various cases. Specifically, the space of the second object itself can be used as the second physical space; for example, the second object can be a refrigerator, in which case the refrigerator can be used as the second physical space. Alternatively and / or additionally, a kitchen can be identified from the image; in this case, the second object is the kitchen, and the kitchen can be used as the second physical space. Alternatively and / or additionally, the second physical space can include other objects, and the spaces of other objects can be used as new second physical spaces. For example, a kitchen can include cabinets, in which case the cabinets can be used as the second space, and so on.
[0057] Here, the knowledge base can be predefined and includes the relationships between physical spaces and objects. For example, the knowledge base may include: (bottled water, refrigerator), (fruit knife, cabinet drawer), etc. The robot can be instructed to pre-collect images of various parts of the first physical space; for example, the robot can open the refrigerator (or cabinet, drawer), etc., and collect images related to various objects within the refrigerator, thereby determining the content of the knowledge base.
[0058] See Figure 5 for further details, which illustrates a block diagram 500 of a process for identifying objects from an image according to some implementations of this disclosure. As shown in Figure 5, objects 510 (refrigerator) and 520 (water dispenser) are identified from image 310. Assuming the knowledge base includes "(bottled water, refrigerator)," the physical space where object 510 is located can be considered as a second physical space. Using some implementations of this disclosure, image recognition technology can be used to determine various objects in the physical space, thereby determining potential physical spaces that may include the first object. In this way, the efficiency and accuracy of locating the first object can be improved, thereby improving the overall efficiency of performing user tasks.
[0059] According to some implementations of this disclosure, after determining the second physical space, a second image of the second physical space can be acquired. During the operation of the robot device 110, the second image can be continuously acquired from multiple devices. For example, the second image can be acquired from at least any one of acquisition devices 114, 130, ..., and 132. In response to determining that the second image indicates that the second physical space includes a first object, the robot device can be instructed to acquire the first object; and the robot device can be instructed to move the first object to the user's location. The robot device can be instructed to move to the second physical space to acquire the first object. Specifically, the robot device can determine how to move from its current location to the second physical space according to a pre-acquired map. For example, images of the surrounding environment can be continuously acquired, and a path to the second physical space can be determined while ensuring obstacle avoidance.
[0060] In some cases, the robotic device can directly enter the second physical space and retrieve the first object. For example, if the second physical space is a "kitchen," the robotic device can directly enter the kitchen and take bottled water from the kitchen counter. In other cases, the robotic device cannot directly enter the second physical space. For example, if the second physical space is a "refrigerator," the robotic device needs to determine the specific method for opening the refrigerator and accessing its internal space. According to some implementations of this disclosure, the robotic device can determine the access method for accessing the second physical space from a second image. Then, the robotic device can be instructed to access the second physical space according to the access method.
[0061] Specifically, the refrigerator handle can be identified from the second image, at which point it can be determined that the handle needs to be pulled to open the refrigerator. Alternatively and / or additionally, a corresponding prompt can be constructed and the model posed a question. The prompt could be, for example, "Determine how to open the refrigerator from the following images," and the prompt and the corresponding image can be sent to the model. The model can then return: Pull the handle. Subsequently, the robot can be instructed to pull the handle to open the refrigerator and locate the bottled water. Utilizing some implementations of this disclosure, the powerful processing capabilities of the model can be invoked to solve unknown problems in complex environments, thereby determining the actions that the robot needs to perform. In this way, the robot's ability to handle complex tasks can be improved, thus performing user tasks more accurately.
[0062] According to some implementations of this disclosure, a motion model can be used to determine the specific actions to be performed by the robot device. See Figure 6 for further details, which shows a block diagram 600 illustrating the process of invoking a motion model according to some implementations of this disclosure. As shown in Figure 6, a motion model 630 can be provided, which can determine the specific actions to be performed by the robot device based on the current state and instructions of the robot device. This motion model can be a pre-trained and fine-tuned model.
[0063] It should be understood that the current state may include data from multiple aspects, such as an image of the robot device, an image of the robot device's environment, pose data of the robot arm (e.g., the positions of the robot arm's joints (POS1, ...)), and the state of the tool (e.g., a gripper, a cutting tool, etc.) fixed to the end of the robot arm. For example, 0 can be used to represent the gripper's closed state, and 1 can be used to represent the gripper's open state. Instructions and the current state can be input into the motion model 630, which then uses the motion model to determine the action to be performed by the robot device based on the instructions and the current state. Here, the action can represent the difference between the robot device's current pose and the next pose, and the difference between the tool's current state and the next state, etc.
[0064] An instruction 610 (e.g., "open the refrigerator") can be input to the motion model 630. Here, the instruction 610 can be expressed in natural language, and the instruction 610 can be determined from the response of the language model. Furthermore, the current state of the robot device can be obtained, and the motion model 630 can determine the corresponding action 640 based on the input data. For example, the orientation, position, speed, acceleration, etc., of each joint in the arm, and / or the wheels and / or other movable devices of the robot device at the next time point can be determined. Furthermore, the determined action 640 can be used to control the state of the robot device at the next time point.
[0065] Using some implementation methods disclosed herein, a relationship can be established between the language model and the action model, and the user's initial input, expressed in natural language, can be converted into specific actions that can be performed by the robotic device. In this way, the actions of the robotic device can be precisely controlled, thereby executing the user task with higher efficiency.
[0066] According to some implementations of this disclosure, after entering the second physical space, an image of the second physical space can be acquired. If it is determined from the image that the refrigerator contains bottled water, and the bottled water is not obstructed by other objects, the robot device can be directly instructed to retrieve the bottled water. See Figures 7A and 7B for further details. Figure 7A shows a block diagram 700A illustrating the process of moving an object according to some implementations of this disclosure. As shown in Figure 7, in image 710, object 220 is not obstructed by other objects, and object 220 can be directly retrieved.
[0067] New instructions and the current state can be input into the motion model 630. The instruction can include "remove bottled water," and the current state can represent the state of the robotic device acquired after the refrigerator has been opened. Furthermore, the motion model 630 can generate new actions to control the robotic device to remove bottled water from the refrigerator. Using some implementations of this disclosure, new instructions and states can be continuously input into the motion model to determine subsequent actions.
[0068] According to some implementations of this disclosure, constraint 720, i.e., the constraints that should be followed during the execution of an action, can be determined. For example, it can be determined that the original posture of the bottled water should be maintained during the movement of the bottled water (e.g., kept vertical and not tilted). Constraint 720 can be determined using a language model. For example, prompts can be generated such as, "Please determine the constraints that should be followed during the movement of the bottled water based on the following image," or "Please determine the precautions during the movement of the bottled water," etc. In this case, the language model can generate the corresponding constraints. The constraints can be input into the action model 630, and the series of actions output by the action model 630 will remove the bottled water from the refrigerator while ensuring the constraints are met. Using some implementations of this disclosure, the safety of the robot device during operation can be ensured, thereby avoiding accidental damage to an object, etc.
[0069] According to some implementations of this disclosure, if bottled water is obstructed by other objects, those objects can be removed first, and then the bottled water can be retrieved. Specifically, in response to determining that a second image indicates that a first object is obstructed by a third object in a second physical space, the third object can be moved to retrieve the second object. See Figure 7B for further details, which shows a block diagram 700B of the process of moving an object according to some implementations of this disclosure. As shown in Figure 7B, in image 730, object 220 is the bottled water to be retrieved, and object 750 is located in front of and obstructs object 220. At this time, a robotic device can be instructed to move object 750 from position 760 to a position that does not obstruct the retrieval of object 220 (such as position 760' in image 740).
[0070] According to some implementations of this disclosure, a target position can be determined, and the robotic device can be instructed to move object 750 to the target position. At this point, the motion model will generate actions to control the robotic device to move object 750 from position 760 to position 760'. In this way, the robotic device can be supported in handling complex problems in complex environments, thereby performing user tasks in a more accurate manner.
[0071] According to some implementations of this disclosure, during the movement of a third object, constraints can be determined based on the pose of the third object, and the robot device can be instructed to move the third object under these constraints. Similar to the process described above, the object 750 can be moved if constraint 720 is met. In this way, it can be ensured that all actions of the robot device in complex environments comply with safety regulations.
[0072] According to some implementations of this disclosure, an exit method for exiting the second physical space can be determined, and the robot device can be instructed to exit the second physical space according to the exit method. Specifically, after the robot device has retrieved object 220, a new image can be acquired and a new prompt word can be constructed to query the language model for the next instruction. The prompt word can be represented, for example, as: "Please determine the next instruction based on the following image," "What to do next," etc. The language model can return "Close the refrigerator," at which point a corresponding action can be generated based on the instruction "Close the refrigerator" and the current state of the robot device to instruct the robot device to close the refrigerator.
[0073] According to some implementations of this disclosure, after the robot device closes the refrigerator, it can be instructed to retrieve a first object and proceed to the user's location. Specifically, images can be acquired in real time, and the user's location can be located within the images. Further, a corresponding instruction can be determined based on the robot device's current location (e.g., location A) and the user's location (e.g., location B). The instruction can then be expressed as: move from location A to location B. At this point, the motion model 630 will generate a corresponding action that controls the robot device 110 to move from location A to location B along a determined trajectory. In this way, the robot device 110 completes the task of "get me a bottle of water."
[0074] According to some implementations of this disclosure, in response to determining that the first image indicates that the first object is included in the first physical space, the robot device 110 can be instructed to retrieve the first object, or the robot device can be instructed to move the first object to the user's location. Specifically, if the first object is initially found to be included in the first physical environment 160 from the first image, the robot device 110 can be directly instructed to go to the location of the first object and retrieve the first object.
[0075] It should be understood that although the foregoing description uses a Chinese language environment as an example to illustrate an exemplary implementation of this disclosure, alternatively and / or additionally, the technical solution of the exemplary implementation of this disclosure can be executed in multiple language environments. For example, the robot can be controlled in environments such as Chinese, English, Japanese, and French. Specifically, the multilingual capabilities provided by machine learning technology can be used to control the robot in application environments in different languages. Furthermore, although the foregoing description uses retrieving bottled water as an example to illustrate the process of using a robotic device to perform user tasks, alternatively and / or additionally, the robotic device can be controlled to perform other user tasks, such as searching for other items in a room, placing an item in a designated location, etc.
[0076] According to some implementations of this disclosure, users can interact with the robotic device through language, actions, gestures, etc. For example, a user can state the user task they wish to perform, predefine a certain action to specify the user task, and so on. Specifically, a user can specify the action of holding a water bottle and drinking water, which can be used as the user task to trigger the robotic device to retrieve the bottled water. When the action is recognized from the acquired image sequence, the robotic device can automatically ask the user if they need bottled water, and if an affirmative answer is received, the robotic device can retrieve the bottled water.
[0077] Alternatively and / or additionally, a user can interact with the robot device via the interaction unit 112, for example, the user inputs a task represented by text and / or images, and controls the robot device to perform the task. Alternatively and / or additionally, the user can specify the execution conditions of the task, for example, to execute the task immediately, to execute the task after a predetermined time, or to execute the task when predetermined conditions are determined to be met (e.g., after the user wakes up), etc.
[0078] According to some implementations of this disclosure, the robotic device can provide users with various messages. For example, assuming the robotic device finds bottled water from multiple brands, it can ask the user which brand they want. Or, assuming the robotic device doesn't find bottled water but only bottled coffee, it can ask the user if they want coffee, and so on. Alternatively and / or additionally, the robotic device can ask the user where they can find the desired item and then go to the user-specified location to find it. Alternatively and / or additionally, if the desired item cannot be found, the robotic device can ask the user if they want to purchase it, and so on.
[0079] According to some implementations of this disclosure, various positioning algorithms can be used to determine the position of robotic devices and various objects in the physical environment. For example, a Global Positioning System (GPS) can be deployed at the robotic device, and satellite signals can be used to determine the precise position of the robotic device. Alternatively and / or additionally, a communication unit can be deployed at the robotic device, and the position of the robotic device can be determined by means of signals between the communication unit and a base station and by utilizing a communication network. Alternatively and / or additionally, a Wi-Fi access point can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspot to determine the position via Wi-Fi signal strength and the known location of the Wi-Fi access point. Alternatively and / or additionally, the communication unit at the robotic device can support Bluetooth functionality, in which case Bluetooth signals and the known locations of Bluetooth devices can be used to determine the position of nearby devices. An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, thereby determining the position of the robotic device.
[0080] Alternate and / or additional locations can be determined using a visual positioning system to pinpoint the location of the robotic device and / or individual objects. A map of the physical space can be pre-acquired, and the locations of each object can be marked on this map. The robotic device can utilize echo detection units to detect distances to surrounding objects and, by combining the acquired images with the physical space map, determine the precise location of each object. Specifically, computer-aided design (CAD) and geographic information systems (GIS) can be used, along with positioning algorithms to determine the location. Alternate and / or additional locations can also be used to deploy tracking units at important objects in the physical space; for example, tracking units can be added to remote controls for household appliances (e.g., television remotes, air conditioner remotes) so that the robotic device can promptly acquire the precise location of important objects, and so on.
[0081] According to some implementations of this disclosure, the robot's initial position and desired destination can be determined based on the methods described above. The robot can determine a path from its initial position to its destination. For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, continuously update the path, enabling the robot to move along the path to its destination.
[0082] According to some implementations of this disclosure, after reaching the destination location, the robotic device can perform a specified task. For example, it can acquire a specified object and move it to the appropriate location. Constraints, i.e., the constraints that should be followed during task execution, can be determined using a language model and / or a knowledge base. For example, an image and corresponding prompts can be acquired, and the image and prompts can be input into the language model, thereby receiving the constraints from the language model. For example, prompts can be determined as: "Based on the following image, determine the constraints that should be followed during the movement of object XXX," or "Please determine the precautions during the movement of object XXX," etc.
[0083] At this point, it can be determined that during the movement of an object (e.g., bottled water, plate, bowl, etc.), the object's original posture should be maintained (e.g., remaining vertical and not tilted). Furthermore, constraints can be input into the motion model, at which point the series of actions output by the motion model will perform the corresponding tasks while ensuring the constraints are met. Using some implementation methods of this disclosure, safety during the operation of robotic devices can be ensured, thereby preventing accidental damage to an object, and so on.
[0084] Using the exemplary implementations of this disclosure, robotic devices can perform user tasks in complex physical spaces. In this way, even if the desired object cannot be found directly in the physical space, the robotic device can search for potential physical spaces that may contain the object, thereby finding the desired object. This improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thus completing the intended user task.
[0085] Example process
[0086] Figure 8 illustrates a flowchart of a method 800 for performing a user task according to some implementations of this disclosure. At block 810, a user task is received from a user, instructing a robotic device to acquire a first object. At block 820, a first image of a first physical space in which the robotic device is located is acquired. At block 830, in response to determining that the first image indicates the first physical space does not contain the first object, a second physical space is determined. At block 840, the robotic device accesses the second physical space to acquire the first object.
[0087] According to some implementations of this disclosure, the robot device accessing the second physical space includes at least one of the following: in response to determining that the robot device can enter the second physical space, the robot device enters the second physical space; or in response to determining that the robot device cannot enter the second physical space, the robot device moves to a predetermined range around the second physical space.
[0088] According to some implementations of this disclosure, determining the second physical space includes: obtaining a prompt word for locating the first object based on the first image and the first object; and receiving a response from a machine learning model to the prompt word in order to determine the second physical space.
[0089] According to some implementations of this disclosure, determining the second physical space includes: identifying multiple second objects from a first image; selecting a second object from the multiple second objects based on a knowledge base; and using the space where the second object is located as the second physical space.
[0090] According to some implementations of this disclosure, the method 800 further includes: acquiring a second image of a second physical space; in response to determining that the second image represents a first object in the second physical space, the robotic device acquires the first object; and the robotic device moves the first object to the user's location.
[0091] According to some implementations of this disclosure, the method 800 further includes: acquiring a second image of a second physical space; and in response to determining that the second image indicates that a first object is occluded by a third object in the second physical space, the robotic device moves the third object to acquire the second object.
[0092] According to some implementations of this disclosure, moving a third object includes: determining constraints during the movement of the third object based on the pose of the third object; and the robotic device moving the third object under the constraints.
[0093] According to some implementations of this disclosure, the method 800 further includes: determining an access method for accessing the second physical space; and the robot device accessing the second physical space according to the access method.
[0094] According to some implementations of this disclosure, the method 800 further includes: determining an exit method for exiting the second physical space; and the robot device exiting the second physical space according to the exit method.
[0095] According to some implementations of this disclosure, the method 800 further includes: in response to determining that a first image represents a first object in a first physical space, a robotic device acquires the first object; and the robotic device moves the first object to the user's location.
[0096] According to some implementations of this disclosure, the first image comes from at least one of the following: an acquisition device at a robot device, an acquisition device in a first physical space, or an acquisition device in a second physical space.
[0097] Example devices and equipment
[0098] Figure 9 shows a block diagram of an apparatus 900 for performing a user task according to some implementations of the present disclosure. The apparatus 900 includes: a receiving module 910 configured to receive a user task from a user, the user task instructing a robot device to acquire a first object; an acquisition module 920 configured to acquire a first image of a first physical space in which the robot device is located; a determining module 930 configured to determine a second physical space in response to determining that the first image indicates the first physical space does not contain the first object; and an execution module 940 configured to cause the robot device to access the second physical space in order to acquire the first object.
[0099] According to some implementations of this disclosure, the execution module 940 is further configured to: in response to determining that the robot device can enter the second physical space, cause the robot device to enter the second physical space; or in response to determining that the robot device cannot enter the second physical space, cause the robot device to move to a predetermined range around the second physical space.
[0100] According to some implementations of this disclosure, the determining module 930 is further configured to: obtain a prompt word for locating the first object based on the first image and the first object; and receive a response from a machine learning model to the prompt word in order to determine a second physical space.
[0101] According to some implementations of this disclosure, the determining module 930 is further configured to: identify a plurality of second objects from a first image; select a second object from the plurality of second objects based on a knowledge base; and associate the space in which the second object is located as a second physical space.
[0102] According to some implementations of this disclosure, the acquisition module 920 is further configured to: acquire a second image of the second physical space; the execution module 940 is further configured to: in response to determining that the second image represents a first object in the second physical space, cause the robot device to acquire the first object; and cause the robot device to move the first object to the user's position.
[0103] According to some implementations of this disclosure, the acquisition module 920 is further configured to: acquire a second image of the second physical space; and the execution module 940 is further configured to: in response to determining that the second image indicates that the first object is occluded by a third object in the second physical space, cause the robot device to move the third object in order to acquire the second object.
[0104] According to some implementations of this disclosure, the execution module 940 is further configured to: determine constraints during the movement of the third object based on the pose of the third object; and cause the robot device to move the third object under the constraints.
[0105] According to some implementations of this disclosure, the execution module 940 is further configured to: determine the access method for accessing the second physical space; and cause the robot device to access the second physical space according to the access method.
[0106] According to some implementations of this disclosure, the execution module 940 is further configured to: determine an exit method for exiting the second physical space; and cause the robot device to exit the second physical space according to the exit method.
[0107] According to some implementations of this disclosure, the execution module 940 is further configured to: in response to determining that the first image represents that the first object is included in the first physical space, cause the robot device to acquire the first object; and cause the robot device to move the first object to the user's location.
[0108] According to some implementations of this disclosure, the first image comes from at least one of the following: an acquisition device at a robot device, an acquisition device in a first physical space, or an acquisition device in a second physical space.
[0109] Figure 10 shows a block diagram of a device 1000 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1000 shown in Figure 10 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 1000 shown in Figure 10 can be used to implement the methods described above.
[0110] As shown in Figure 10, the computing device 1000 is in the form of a general-purpose computing device. Components of the computing device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage devices 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. The processing unit 1010 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1020. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 1000.
[0111] Computing device 1000 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1030 can be removable or non-removable media and may include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within computing device 1000.
[0112] The computing device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 1020 may include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.
[0113] The communication unit 1040 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1000 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0114] Input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1060 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 1000 can also communicate with one or more external devices (not shown) via communication unit 1040 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 1000, or with any device (e.g., network card, modem, etc.) that enables computing device 1000 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interface (not shown).
[0115] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
[0116] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0117] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0118] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0120] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for performing a user task, comprising: receiving a user task from a user, the user task instructing a robotic device to retrieve a first object; retrieving a first image of a first physical space in which the robotic device is located; in response to determining that the first image represents that the first object is not included in the first physical space, determining a second physical space; and the robotic device accessing the second physical space in order to retrieve the first object.
2. The method of claim 1, wherein the robotic device accessing the second physical space comprises at least either of: in response to determining that the robotic device is able to enter the second physical space, the robotic device entering the second physical space; or in response to determining that the robotic device is unable to enter the second physical space, the robotic device moving to a predetermined range around the second physical space.
3. The method of claim 1, wherein determining the second physical space comprises: retrieving a cue for locating the first object based on the first image and the first object; and receiving a response of a machine learning model to the cue in order to determine the second physical space.
4. The method of claim 1, wherein determining the second physical space comprises: identifying a plurality of second objects from the first image; selecting a second object from the plurality of second objects based on a knowledge base; and identifying a space associated with the second object as the second physical space.
5. The method of claim 1, further comprising: retrieving a second image of the second physical space; in response to determining that the second image represents that the first object is included in the second physical space, the robotic device retrieving the first object; and the robotic device moving the first object to a location of the user.
6. The method of claim 1, further comprising: retrieving a second image of the second physical space; and in response to determining that the second image represents that the first object is occluded by a third object in the second physical space, the robotic device moving the third object in order to retrieve the second object.
7. The method of claim 6, wherein moving the third object comprises: determining a constraint condition during moving the third object based on a pose of the third object; and the robotic device moving the third object under the constraint condition.
8. The method of claim 3, further comprising: determining an access manner for accessing the second physical space; and the robotic device accessing the second physical space in accordance with the access manner.
9. The method of claim 8, further comprising: determining an exit manner for exiting the second physical space; and the robotic device exiting the second physical space in accordance with the exit manner.
10. The method of claim 1, further comprising: in response to determining that the first image represents that the first object is included in the first physical space, the robotic device retrieving the first object; and The robotic device moves the first object to the location of the user.
11. The method of claim 1, wherein the first image is from at least any of: a capture device at the robotic device, a capture device in a first physical space, a capture device in a second physical space.
12. An apparatus for performing a user task, comprising: a receiving module configured to receive a user task from a user, the user task robotic device to acquire a first object; an acquiring module configured to acquire a first image of a first physical space in which the robotic device is located; a determining module configured to determine a second physical space in response to determining that the first image represents that the first object is not included in the first physical space; and an executing module configured to cause the robotic device to access the second physical space in order to acquire the first object.
13. An electronic device, comprising: at least one processing unit; and at least one memory that is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1-11.
14. A computer readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to implement the method according to any one of claims 1-11.
15. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-11.
Citation Information
Patent Citations
Intelligent movable equipment capable of object searching and intelligent object searching method
CN107977625A
Sundry cleaning robot system
CN116709962A
Visual language navigation method combining image description and text generation image
CN117571014A
Method and device for generating video describing entity, equipment and medium
CN117793482A
Robot semantic mapping and navigation method based on life scene and robot
CN118149812A