Method and apparatus for executing user task, and device and medium
By receiving user tasks, acquiring images, and using machine learning models to determine the destination location, robotic devices can move objects flexibly and accurately in complex environments, solving the problem that existing robotic devices struggle to perform complex user tasks.
Patent Information
- Application Number
- PCT/CN2024/106255
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-22
Smart Images

Figure CN2024106255_22012026_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices, and media for performing user tasks Technical Field
[0001] The exemplary implementations of this disclosure generally relate to the field of robotics, and more particularly to methods, apparatus, devices, and computer-readable storage media for using robots to perform user tasks. Background Technology
[0002] Robotics technology has developed rapidly and is widely used in many technological fields. Various specialized robotic devices have been developed; for example, in industrial environments, robots can perform a variety of tasks such as processing, grasping, sorting, and packaging. In home environments, for instance, robotic vacuum cleaners and window cleaning robots have been developed. However, robots typically can only perform pre-set, fixed tasks and cannot perform different user-defined tasks according to user needs.
[0003] Summary of the Invention
[0004] In a first aspect of this disclosure, a method for performing a user task is provided. In this method, a user task is received from a user, the user task instructing a robotic device to classify multiple objects within a first range in a physical space; an image including the multiple objects is acquired; for a first object among the multiple objects, a first destination location in the physical space is determined based on the image of the first object; and the robotic device moves the first object to the first destination location.
[0005] In a second aspect of this disclosure, an apparatus for performing a user task is provided. The apparatus includes: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to classify a plurality of objects within a first range in a physical space; an acquiring module configured to acquire an image including the plurality of objects; a determining module configured to determine, based on the image, a first destination location in the physical space for a first object among the plurality of objects; and an executing module configured to cause the robotic device to move the first object to the first destination location.
[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.
[0008] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent in the following detailed description, taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;
[0012] Figure 2 shows a block diagram of some implementations of the present disclosure for performing user tasks;
[0013] Figure 3 shows a block diagram of an image acquisition process according to some implementations of this disclosure;
[0014] Figure 4 shows a block diagram of the process of moving an object according to some implementations of this disclosure;
[0015] Figure 5 shows a block diagram of the calling model process according to some implementations of this disclosure;
[0016] Figure 6 shows a block diagram of the process of invoking the action model according to some implementations of this disclosure;
[0017] Figure 7 shows a flowchart of a method for performing user tasks according to some implementations of this disclosure;
[0018] Figure 8 shows a block diagram of an apparatus for performing user tasks according to some implementations of the present disclosure; and
[0019] Figure 9 shows a block diagram of a device capable of implementing various implementations of the present disclosure. Detailed Implementation
[0020] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0021] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.
[0022] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0023] It is understood that before using the technical solutions disclosed in each implementation of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0024] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0025] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0026] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0027] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
[0028] Example Environment
[0029] In recent years, robotics and machine learning technologies have been widely applied in various scenarios. However, robots typically can only perform pre-set, fixed tasks and cannot execute different user tasks according to user needs. In particular, in complex application environments, robotic devices struggle to determine user requirements and thus perform corresponding tasks.
[0030] Simple robotic devices have been developed to perform specific tasks. However, these devices cannot understand complex user instructions, nor can they execute the desired tasks according to user commands in complex physical spaces. Therefore, it is desirable to control the robot's operation in an effective way to perform the desired tasks.
[0031] According to an exemplary implementation of this disclosure, a method for performing user tasks is proposed. Referring to Figure 1, which describes an application environment according to an exemplary implementation of this disclosure, Figure 1 shows a block diagram 100 of the application environment according to an exemplary implementation of this disclosure. As shown in Figure 1, a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform various tasks. The physical space 160 can include, but is not limited to, one or more rooms. For example, in a home environment, the physical space 160 can include, but is not limited to, a living room, bedroom, study, kitchen, toilet, etc., or a combination of one or more of the above.
[0032] As shown in Figure 1, the robot device 110 may include multiple parts. For example, the control unit 111 can serve as the control center of the robot device 110, and an application can be loaded into the control unit 111 to control the various parts of the robot device. The user 120 can use the interaction unit 112 to interact with the robot device 110, for example, by inputting control commands to the robot device 110 to perform desired tasks. The robot device 110 may include an arm 113 for performing actions such as grasping and releasing. For example, the arm 113 can grasp an object and move it to a desired position, and so on.
[0033] Alternatively and / or additionally, the robot device 110 may also include a data acquisition unit 114. Here, the data acquisition unit 114 may include various types, such as an image acquisition unit, a sound acquisition unit, etc. Alternatively and / or additionally, the robot device 110 may further include a sensing unit for detecting surrounding objects, for example, detecting the distance between the robot and surrounding objects based on laser light, etc. The robot device 110 may also include a drive unit 115; for example, the robot device 110 may be deployed on a movable base, and the drive unit 115 may drive the wheels of the base to move along a desired path.
[0034] Physical environment 160 may include one or more acquisition units 130, ..., and 132. For example, one or more image acquisition devices may be deployed in a room to acquire images of the room from various angles. Physical environment 160 may include control device 140, which can control one or more acquisition units 130, ..., and 132, etc., via a network (not shown). Alternatively and / or additionally, in a smart home environment, control device 140 can control various electrical devices in physical space 160.
[0035] Alternatively and / or additionally, a machine learning model (e.g., model 150) may be provided to manage physical space 160. It should be understood that although Figure 1 shows model 150 located inside physical space 160, alternatively and / or additionally, model 150 may be located at a remote device outside physical space 160, and control device 140, robotic device 110, or other device may access the remote model 150 via a network.
[0036] Model 150 may include one or more models. If model 150 includes multiple models, these multiple models may include multiple types of models. Model 150 may, for example, include at least a language model (LM) and an action model. The language model, by learning from a large corpus, is capable of question answering. The action model can control the robotic device 110 to perform various actions. Model 150 may also include, for example, an image recognition model, a text recognition model, and so on.
[0037] As shown in Figure 1, user 120 can instruct robot device 110 to manipulate various objects in physical space 110. Here, objects can be various items in the home environment. For example, user 120 can instruct robot device 110 to find a certain object in physical space 160; or user 120 can instruct robot device 110 to place the found object in a designated location, and so on.
[0038] Summary of the task to be performed
[0039] To at least partially address the shortcomings of the prior art, a method for performing user tasks is proposed according to an exemplary implementation of this disclosure. Referring to Figure 2, which describes an overview of an exemplary implementation of this disclosure, Figure 2 illustrates a block diagram 200 for performing user tasks according to some implementations of this disclosure.
[0040] As shown in Figure 2, the robot device 110 in the physical space 160 can receive a user task 210 from the user 120. In this case, the user task 210 can instruct the robot device 110 to categorize multiple objects within a specified area (e.g., area 230, also referred to as the first area) in the physical space 160. For example, the user 120 can say in natural language, "Organize the items on the desktop." In the example of Figure 2, the first area is the "desktop," and the multiple objects are the multiple objects placed on the desktop, including object 222 and object 224. The robot device 110 can acquire images including the multiple objects. For example, images can be acquired via at least one of acquisition units 114, 130, ..., and 132.
[0041] For a first object among multiple objects, a first destination location in physical space is determined based on an image. The first object can be any suitable object among the multiple objects. The first destination location can be the place where the first object should be placed. For example, in a home environment, a table includes a bottle of water (i.e., object 222) and a book (i.e., object 224). Object 222 should be placed in a refrigerator, kitchen, or other location; therefore, if object 222 is the first object, the first destination location could be a refrigerator (e.g., destination location 250) and / or a kitchen. The book should be placed on a bookshelf, in a study, or other location; therefore, if the book is the first object, the first destination location could be a bookshelf and / or a study. In this case, the robot device 110 can be instructed to move the first object to the first destination location. For example, the robot device 110 can be instructed to move object 222 to destination location 250.
[0042] According to some implementations of this disclosure, the methods described above can be executed at any computing device with computing capabilities. For example, the methods described above can be executed using an application deployed at robot device 110. Alternatively and / or additionally, an application can be deployed at control device 140 to execute the methods described above. Specifically, the powerful processing capabilities of model 150 can be invoked to find a first range that may include multiple objects and the destination location in physical space for each of the multiple objects. Subsequently, robot device 110 can move along path 242 to range 230 and along path 244 move object 222 of the multiple objects within range 230 to its destination location 250 in physical space 160.
[0043] Using the exemplary implementation of this disclosure, a robotic device can be controlled to move multiple objects within a first range to their corresponding destination locations. This can improve the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby completing the intended user task.
[0044] Detailed process of executing the task
[0045] Having described an outline of some implementations according to this disclosure, further details regarding the execution of user tasks will be described below. For ease of description, the following example uses the control of robot device 110 to organize items on a desktop as an illustration to further illustrate the execution of user tasks.
[0046] According to some implementations of this disclosure, the image can come from at least one of the following: an acquisition device at the robot device, an acquisition device in a first physical space, or an acquisition device in a second physical space. Referring to Figure 3 for further details of image acquisition, Figure 3 shows a block diagram 300 of the image acquisition process according to some implementations of this disclosure. As shown in Figure 3, images (e.g., one or more images 310) of the physical space 160 can be acquired from the acquisition unit 114 at the robot device 110. Since the robot device 110 can move freely within the physical space 160, the acquisition unit 114 can acquire images from various locations within the physical space, thereby facilitating the location of a first range and multiple objects.
[0047] Alternatively and / or additionally, images of the physical space 160 can be acquired from acquisition units 130, ..., and 132. Here, acquisition units 130, ..., and 132 can be pre-deployed at designated locations within the physical space 160, such as a corner of the ceiling, etc. In this way, images of the physical space 160 taken from a top-down angle can be obtained, facilitating an overall understanding of the layout of the physical space 160. This can help determine a first range and the location of a first destination corresponding to a first object within the first range.
[0048] According to some implementations of this disclosure, an image of the physical space 160 where the robot device 110 is located can be acquired, and a first range can be located based on the image. For example, if user 120 says "tidy up the items on the desktop" in natural language, then the user task 210 instructs the robot device 110 to sort multiple objects on the desktop. The range 230 where the desktop is located can be determined based on the image of the physical space 160. For example, a prompt for locating the first range can be acquired: "Please determine the location of the desktop in the following image," and the range 230 can be located using a model. Then, the robot device 110 can be instructed to move to the range 230.
[0049] According to some implementations of this disclosure, images of multiple objects within a first range can be acquired simultaneously with images of the physical space 160 where the robot device 110 is located. For example, images of multiple objects within the first range can be determined based on an image of the physical space 160. For instance, images of multiple objects within the first range can be cropped from an image of the physical space 160. Alternatively or additionally, according to some implementations of this disclosure, images of multiple objects within the first range can also be acquired in response to the robot device 110 moving to the first range. For example, images of multiple objects within the first range can be acquired using the acquisition unit 114 at the robot device 110 in response to the robot device 110 moving to the first range.
[0050] According to some implementations of this disclosure, the images of multiple objects may include one or more images. If the images of multiple objects include multiple images, each image may correspond to one object. For example, the images of multiple objects may include multiple images corresponding to each of the multiple objects. For a first object, the position of the first object can be adjusted to acquire the image of the first object. At this time, the position of the acquisition device used to acquire the image can be adjusted (e.g., an angle that avoids occlusion can be found) to acquire the image of the first object. For example, for multiple objects within range 230, the position of object 222 and / or the position of the mobile robot device 110 can be adjusted so that the robot device 110 can acquire an image containing only object 222 using the acquisition unit 114. Thus, the image corresponding to each of the multiple objects can be acquired separately, which can improve the accuracy of the image corresponding to each object, and thus improve the accuracy of subsequent image-based task execution.
[0051] If the images of multiple objects consist of only one image, it can be determined whether an occlusion relationship exists between the multiple objects. Any suitable method can be used to determine whether an occlusion relationship exists between the multiple objects. For example, model 150 can be used to determine whether an occlusion relationship exists between the multiple objects. Alternatively, it can be based on any suitable rule or algorithm. Alternatively or additionally, prompts can be provided to user 120 to allow user 120 to determine whether an occlusion relationship exists between the multiple objects. The determination result provided by user 120 can be received, and the existence of an occlusion relationship between the multiple objects can be determined based on that determination result. Alternatively and / or additionally, if an occlusion relationship exists, the user can be instructed to eliminate the occlusion relationship.
[0052] If there is no occlusion relationship between multiple objects, an image including multiple objects can be directly acquired. If there is an occlusion relationship between multiple objects, in response to determining that there is an occlusion relationship between multiple objects, the robot device 110 can be instructed to move at least one of the multiple objects, thereby acquiring an image including multiple objects. See Figure 4 for further details, which shows a block diagram 400 of the process of moving objects according to some implementations of this disclosure. As shown in Figure 4, object 402 is located in front of object 401 and occludes object 401. In this case, the robot device can be instructed to move object 402 from position 430 to a position that does not occlude object 401 (such as position 430' in image 420).
[0053] According to some implementations of this disclosure, a target position can be determined, and the robotic device can be instructed to move object 402 to the target position. For example, actions can be generated using model 150 to control the robotic device to move object 402 from position 430 to position 430'. In this way, the robotic device can be supported in handling complex problems in complex environments, thereby performing user tasks in a more accurate manner.
[0054] After obtaining images corresponding to multiple objects, the first destination location of the first object in physical space can be determined based on the images. According to some implementations of this disclosure, the first type of the first object can be determined based on the images. Regarding the specific method for determining the first type of the first object, any suitable method can be used. For example, the first type of the first object can be determined based on a pre-stored database. The pre-stored database can store multiple objects and their respective types. The first object can be retrieved from this pre-stored database to determine its first type.
[0055] For example, multiple text items (e.g., labels on objects) associated with multiple objects can be identified from an image (e.g., using model 150 to identify the image). For a first object, a first type can be determined based on the first text item associated with the first object. Taking object 222 as an example, if object 222 is bottled water, the text item associated with object 222 could be, for example, the text item on the bottled water packaging. Object 222 can be determined to be bottled water based on the text item on the bottled water packaging.
[0056] According to some implementations of this disclosure, the location of a second object of the first type in physical space can also be determined so as to determine the location of a first destination based on the location of the second object. It is understood that any suitable method can be used to determine the location of the second object. For example, the location of the second object can be determined based on an image of the physical space 160 where the robot device 110 is located. Taking bottled water as an example, an image of the physical space 160 where the robot device 110 is located can be acquired, and the locations of other bottled water items in the physical space 160 can be determined based on this image. Taking other bottled water items as being placed in the kitchen as an example, the location of the second object can be determined as the kitchen. As another example, the location of the second object can also be determined based on predetermined rules or algorithms. Again, taking bottled water as an example, if a predetermined rule indicates that the bottled water is placed in the refrigerator, the location of the second object can be determined as the refrigerator based on the predetermined rule. It is understood that the location of the second object can also be determined by means of a model (e.g., model 150), or manually by means of a user 120, etc. For example, the robot device 110 can ask the user, "Where should I put the bottled water?" and place the bottled water in the location specified by the user.
[0057] To determine the location of the second object using a model, according to some implementations of this disclosure, a prompt word can be constructed based on the image of the physical space 160 and the first type of the first object. This prompt word can be used to determine the location of the first type of object in the image of the physical space 160. Taking the determination that the first type of the first object is bottled water as an example, the prompt word could be expressed as: "Please identify the location where 'bottled water' is placed from the following image."
[0058] The cue word can be provided to a machine learning model (e.g., model 150) to determine the location of a second object of the first type. The machine learning model's response to the cue word, i.e., the model output of the machine learning model to the cue word, can indicate the location of the second object (e.g., a refrigerator). Thus, the location of the second object, i.e., the first destination location of the first object, can be determined based on the machine learning model's response to the cue word.
[0059] Specifically, the machine learning model can process images, and if the image includes a second object, the model can output the location of the second object (e.g., the region coordinates of the second object in the image, and / or directly output the image of the region where the second object is located, etc.). If the image does not include the second object, the model can output a response such as "not found". Using some implementations of this disclosure, the presence or absence of a second object in an image can be detected in multiple ways, thereby improving the performance of robotic devices in detecting the location of a second object in an image.
[0060] Alternatively or additionally, according to some implementations of this disclosure, in response to determining that the physical space 160 in which the image represents the robot device 110 is located does not contain the second object, another physical space associated with the physical space 160 may be determined. In this case, the physical space 160 in which the robot device 110 is located may be referred to as the first physical space, and the other physical space associated with the physical space 160 may be referred to as the second physical space.
[0061] It should be understood that the second physical space here is a potential physical space that may include the first object. For example, in a home environment, since bottled water may be placed in the refrigerator, the second physical space can be identified as the refrigerator. Specifically, in determining the second physical space associated with the first physical space, based on the image of physical space 160 and the first type, cue words for locating the second object can be obtained, and the response of the machine learning model to the cue words can be received to determine the second physical space. After determining the second physical space that may include the second object, the location of the second physical space in the first physical space can be determined as the location of the second object, that is, the first destination location of the first object.
[0062] Referring to Figure 5 for further details regarding the determination of the second physical space, Figure 5 illustrates a block diagram 500 of the process of invoking the model according to some implementations of this disclosure. As shown in Figure 5, multiple objects on the desktop (e.g., bottled water and books) can be moved to corresponding locations based on user task 210. For bottled water, a corresponding prompt 510 can be obtained based on image 310. Prompt 510 can be expressed, for example, as: "Please determine the physical space that may include 'bottled water' from the following images," or, for example, as: "Where might 'bottled water' be placed in the following images," and so on. Prompt 510 and image 310 can be input to language model 520 so that language model 520 can find the second physical space that may include bottled water from image 310. Language model 520 can be, for example, a model included in model 150.
[0063] According to some implementations of this disclosure, after determining the first destination location, a message associated with the first object can be provided to the user 120. This message can be used to prompt the user 120 to confirm whether to move the first object to the first destination location. Regarding the specific method of providing the message, it can be provided to the user 120 via the display screen of the robot device 110, a voice playback device, etc. Alternatively, in response to receiving a reply from the user 120 regarding the message, the decision to move the first object to the first destination location can be determined based on the reply. If the reply indicates that the first object should be moved to the first destination location, the robot device 110 can be instructed to move the first object to the first destination location. Therefore, the object can be moved only upon obtaining user confirmation, and the accuracy of the robot device performing tasks can be improved by following user instructions.
[0064] According to some implementations of this disclosure, the motion trajectory from the position of the first object to the first destination position can also be determined based on an image of the physical space 160. It is understood that the motion trajectory can also be determined by any suitable method, such as using model 150, manual intervention, or based on predetermined rules or algorithms. This motion trajectory can, for example, instruct the robotic device 110 to avoid obstacles in the physical space 160 and move from the position of the first object to the first destination position along a shorter path. After determining the motion trajectory, the robotic device 110 can be instructed to move the first object to the first destination position according to the motion trajectory.
[0065] According to some implementations of this disclosure, the actions to be performed by the robot device 110 can be determined using model 150, and the robot device 110 can be controlled to perform these actions using model 150. For example, if the first destination location is a refrigerator, the specific method of moving the first object to the refrigerator (e.g., opening the refrigerator, placing the first object, etc.) can be determined using model 150. Exemplarily, a corresponding prompt can be constructed to ask model 150 how to open the refrigerator. The prompt could be, for example, "Determine how to open the refrigerator from the following images," and the prompt and the corresponding image (e.g., an image of the physical environment 160 including the refrigerator) can be sent to model 150.
[0066] Model 150 could, for example, return: pull the handle to open the refrigerator door. Then, it can instruct robot device 110 to pull the handle to open the refrigerator. Thus, by utilizing some implementations of this disclosure, the powerful processing capabilities of the model can be invoked to solve unknown problems in complex environments, thereby determining the actions that the robot device needs to perform. In this way, the robot device's ability to handle complex tasks can be improved, thereby executing user tasks in a more accurate manner.
[0067] According to some implementations of this disclosure, a motion model can be used to determine the specific actions to be performed by the robot device. See Figure 6 for further details, which shows a block diagram 600 illustrating the process of invoking a motion model according to some implementations of this disclosure. As shown in Figure 6, a motion model 630 can be provided, which can determine the specific actions to be performed by the robot device based on the current state and instructions of the robot device. This motion model 630 can be a pre-trained and fine-tuned model. The motion model 630 can also be one of the models included in model 150.
[0068] The current state 620 may include data from various aspects, such as an image of the robot device 110, an image of the robot device 110's environment, pose data of the robot arm (e.g., the positions of the robot arm's joints (POS1, ...)), and the state of the tool (e.g., a gripper, a cutting tool, etc.) fixed to the end of the robot arm. For example, 0 can be used to represent the gripper's closed state, and 1 can be used to represent the gripper's open state. Instructions and the current state can be input into the motion model 630, which then uses the motion model to determine the action to be performed by the robot device 110 based on the instructions and the current state. Here, the action can represent the difference between the robot device 110's current pose and the next pose, and the difference between the tool's current state and the next state, etc.
[0069] An instruction 610 (e.g., "open the refrigerator") can be input to the motion model 630. Here, the instruction 610 can be expressed in natural language, and the instruction 610 can be determined from the response of the language model 630. Furthermore, the current state of the robot device 110 can be obtained, and the motion model 630 can determine the corresponding action 640 based on the input data. For example, the orientation, position, speed, acceleration, etc., of each joint in the arm, and / or the wheels and / or other movable devices of the robot device at the next time point can be determined. Furthermore, the determined action 640 can be used to control the state of the robot device 110 at the next time point.
[0070] In this way, the robot's movements can be controlled by using a model. The robot can be instructed to move a first object based on the model's output, and the robot's movements can be precisely controlled, thereby executing user tasks more efficiently.
[0071] It is understandable that multiple objects within a first range can be identified as having multiple destination locations. A similar approach can be used to control the robot device 110 to move the multiple objects sequentially to their respective destination locations. For example, the robot device 110 can be controlled to move bottled water to the refrigerator, then move a book to the bookshelf, and so on. It is important to note that if at least two of the multiple objects are of the same type, these two objects can be moved together. For example, if the multiple objects include a group of objects of the first type, and this group includes at least two objects, the robot device 110 can be instructed to move this group of objects at once. Therefore, by moving a group of objects at a time based on their different types, the efficiency of the robot device 110 in moving objects can be improved.
[0072] Alternatively or additionally, according to some implementations of this disclosure, if the number of a group of objects of a first type among a plurality of objects is determined to meet a threshold condition, the robot device 110 may be instructed to acquire a third object. The threshold condition may, for example, indicate a threshold number of objects in a group. For example, if the threshold number is 4, the robot device 110 may be instructed to acquire a third object in response to the number of objects included in the group of objects of the first type reaching 4. The third object may, for example, be another object used to move the group of objects. The third object may be a pre-determined object.
[0073] For example, user 120 can pre-configure robot device 110 to move multiple objects using specified objects. For instance, taking a set of objects as a group of bottled water, the third object could be any object that helps move this group of bottled water, such as a tray, basket, bag, or trailer. Furthermore, robot device 110 can be instructed to move a group of objects to a first destination location via the third object. For example, if the third object is a tray, robot device 110 can be instructed to move a group of objects (e.g., a group of bottled water) to the first destination location using the tray.
[0074] According to some implementations of this disclosure, after moving at least one object of the first type to the first destination location, the destination location of at least one object of the second type can be determined, and the at least one object of the second type can be moved to the second destination location of the second type of object in physical space. For example, model 150 can be used to determine the way to leave the first destination location, the movement trajectory to return to the first range, the movement trajectory to move from the first range to the second destination, and so on. New images can be acquired and new prompts can be constructed to query model 150 (e.g., a language model) for the next instruction. Prompts can be represented, for example, as: "Please determine the next instruction based on the following image," "What to do next," etc. For example, after the robot device has placed bottled water in the refrigerator, the language model can return "Close the refrigerator" based on the currently received image. At this time, based on the instruction "Close the refrigerator" and the current state of the robot device, a corresponding action can be generated to instruct the robot device to close the refrigerator.
[0075] According to some implementations of this disclosure, the priorities of multiple objects within a first range can also be determined, and the multiple objects can be moved sequentially based on their priorities. The priorities of the multiple objects can be determined in any way; for example, they can be determined using model 150 or predetermined rules. For instance, objects that need to be frozen (e.g., ice cream, frozen meat, etc.) will have a higher priority than objects that can be stored at room temperature. In this way, objects with higher priorities can be moved first, improving the quality of task execution by the robotic device 110.
[0076] It should be understood that although the foregoing description uses a Chinese language environment as an example to illustrate an exemplary implementation of this disclosure, alternatively and / or additionally, the technical solution of the exemplary implementation of this disclosure can be executed in multiple language environments. For example, the robot can be controlled in environments such as Chinese, English, Japanese, and French. Specifically, the multilingual capabilities provided by machine learning technology can be used to control the robot in application environments of different languages. Furthermore, although the foregoing description uses retrieving bottled water as an example to illustrate the process of using a robotic device to perform user tasks, alternatively and / or additionally, the robotic device can be controlled to perform other user tasks, such as searching for other items in a room, placing an item in a designated location, etc.
[0077] According to some implementations of this disclosure, users can interact with the robot device through language, actions, gestures, etc. For example, users can state the user task they wish to perform, predefine a certain action to specify the user task, and so on. Specifically, users can make the action of placing items, and this action can be used as a trigger for the robot device to perform the user task of sorting and placing items. When the action is recognized from the acquired image sequence, the robot device can automatically ask the user whether items need to be sorted and place and ask the user about the scope of the task to be processed. If an affirmative answer is received, the robot device can perform the user task.
[0078] Alternatively and / or additionally, a user can interact with the robot device via the interaction unit 112, for example, the user inputs a task represented by text and / or images, and controls the robot device to perform the task. Alternatively and / or additionally, the user can specify the execution conditions of the task, for example, to execute the task immediately, to execute the task after a predetermined time, or to execute the task when predetermined conditions are determined to be met (e.g., after the user has eaten), etc.
[0079] According to some implementations of this disclosure, the robotic device can provide users with various messages. For example, regarding a first object (e.g., bottled water), if the physical space includes multiple first destination locations (e.g., a refrigerator and a kitchen), the device can ask the user whether to move the bottled water to the refrigerator or the kitchen. Or, for example, assuming the robotic device cannot find a first destination location, it can ask the user where to move the first object, and so on.
[0080] According to some implementations of this disclosure, various positioning algorithms can be used to determine the position of robotic devices and various objects in the physical environment. For example, a Global Positioning System (GPS) can be deployed at the robotic device, and satellite signals can be used to determine the precise position of the robotic device. Alternatively and / or additionally, a communication unit can be deployed at the robotic device, and the position of the robotic device can be determined by means of signals between the communication unit and a base station and by utilizing a communication network. Alternatively and / or additionally, a Wi-Fi access point can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspot to determine the position via Wi-Fi signal strength and the known location of the Wi-Fi access point. Alternatively and / or additionally, the communication unit at the robotic device can support Bluetooth functionality, in which case Bluetooth signals and the known locations of Bluetooth devices can be used to determine the position of nearby devices. An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, thereby determining the position of the robotic device.
[0081] Alternate and / or additional locations can be determined using a visual positioning system to pinpoint the location of the robotic device and / or individual objects. A map of the physical space can be pre-acquired, and the locations of each object can be marked on this map. The robotic device can utilize echo detection units to detect distances to surrounding objects and, by combining the acquired images with the physical space map, determine the precise location of each object. Specifically, computer-aided design (CAD) and geographic information systems (GIS) can be used, along with positioning algorithms to determine the location. Alternate and / or additional locations can also be used to deploy tracking units at important objects in the physical space; for example, tracking units can be added to remote controls for household appliances (e.g., television remotes, air conditioner remotes) so that the robotic device can promptly acquire the precise location of important objects, and so on.
[0082] According to some implementations of this disclosure, the robot's initial position and desired destination can be determined based on the methods described above. The robot can determine a path from its initial position to its destination. For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, continuously update the path, enabling the robot to move along the path to its destination.
[0083] According to some implementations of this disclosure, after reaching the destination location, the robotic device can perform a specified task. For example, it can acquire a specified object and move it to the appropriate location. Constraints, i.e., the constraints that should be followed during task execution, can be determined using a language model and / or a knowledge base. For example, an image and corresponding prompts can be acquired, and the image and prompts can be input into the language model, thereby receiving the constraints from the language model. For example, prompts can be determined as: "Based on the following image, determine the constraints that should be followed during the movement of object XXX," or "Please determine the precautions during the movement of object XXX," etc.
[0084] At this point, it can be determined that during the movement of an object (e.g., bottled water, plate, bowl, etc.), the object's original posture should be maintained (e.g., remaining vertical and not tilted). Furthermore, constraints can be input into the motion model, at which point the series of actions output by the motion model will perform the corresponding tasks while ensuring the constraints are met. Using some implementation methods of this disclosure, safety during the operation of robotic devices can be ensured, thereby preventing accidental damage to an object, and so on.
[0085] Using the exemplary implementation of this disclosure, a robotic device can perform user tasks in complex physical spaces. In this way, the robotic device can autonomously move multiple objects within a first range to their corresponding destination locations. This improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby completing the intended user task.
[0086] Example process
[0087] Figure 7 illustrates a flowchart of a method 700 for performing a user task according to some implementations of this disclosure. At block 710, a user task is received from a user, instructing a robotic device to classify multiple objects within a first range in physical space. At block 720, an image including the multiple objects is acquired. At block 730, for a first object among the multiple objects, a first destination location in physical space is determined based on the image. At block 740, the robotic device moves the first object to the first destination location.
[0088] According to some implementations of this disclosure, acquiring an image includes: in response to determining that an occlusion relationship exists between multiple objects, a robotic device moves at least one of the multiple objects; and acquiring an image including the multiple objects.
[0089] According to some implementations of this disclosure, the image of the first object is determined based on: adjusting the position of the first object to acquire an image of the first object; and adjusting the position of the acquisition device used to acquire the image to acquire an image of the first object.
[0090] According to some implementations of this disclosure, method 700 further includes: acquiring an image of the physical space where the robot device is located; locating a first range based on the image; and moving the robot device to the first range.
[0091] According to some implementations of this disclosure, determining the location of the first destination includes: determining a first type of a first object based on an image; determining the location of a second object of the first type in physical space; and determining the location of the first destination based on the location of the second object.
[0092] According to some implementations of this disclosure, determining the first type of the first object further includes: identifying multiple text items from the image that are associated with multiple objects respectively; and, for the first object among the multiple objects, determining the first type based on the first text item associated with the first object among the multiple text items.
[0093] According to some implementations of this disclosure, determining the location of the second object includes at least one of the following: identifying the second object in an image in physical space to determine the location of the second object.
[0094] According to some implementations of this disclosure, determining the location of the second object includes at least one of the following: constructing a cue word in an image in physical space and a first type of the first object, the cue word being used to determine the location of the first type of object in the image in physical space; and determining the location of the second object based on the response of a machine learning model to the cue word.
[0095] According to some implementations of this disclosure, the robotic device moves a first object to a first destination location by: determining a motion trajectory from the location of the first object to the first destination location based on an image of the physical space; and moving the first object to the first destination location according to the motion trajectory.
[0096] According to some implementations of this disclosure, moving a first object to a first destination location by a robotic device includes: providing a user with a message associated with the first object; and moving the first object to the first destination location in response to receiving a response from the user to the message.
[0097] According to some implementations of this disclosure, the robot device moves a first object to a first destination location by: in response to determining that the number of a group of objects of a first type among a plurality of objects meets a threshold condition, the robot device acquires a third object; the robot device moves a group of objects to the first destination location via the third object.
[0098] Example devices and equipment
[0099] Figure 8 shows a block diagram of an apparatus 800 for performing a user task according to some implementations of the present disclosure. The apparatus 800 includes: a receiving module 810 configured to receive a user task from a user, the user task instructing a robotic device to classify a plurality of objects within a first range in physical space; an acquiring module 820 configured to acquire an image including the plurality of objects; a determining module 830 configured to determine a first destination location in physical space for a first object among the plurality of objects based on the image; and an executing module 840 configured to cause the robotic device to move the first object to the first destination location.
[0100] According to some implementations of this disclosure, the acquisition module 820 is further configured to: in response to determining that there is an occlusion relationship between multiple objects, cause the robot device to move at least one of the multiple objects; and acquire an image including the multiple objects.
[0101] According to some implementations of this disclosure, the image of the first object is determined based on: adjusting the position of the first object to acquire an image of the first object; and adjusting the position of the acquisition device used to acquire the image to acquire an image of the first object.
[0102] According to some implementations of this disclosure, the acquisition module 820 is further configured to: acquire an image of the physical space where the robot device is located; locate a first range based on the image; and move the robot device to the first range.
[0103] According to some implementations of this disclosure, the determining module 830 is further configured to: determine a first type of a first object based on an image; determine the location of a second object of the first type in physical space; and determine the location of a first destination based on the location of the second object.
[0104] According to some implementations of this disclosure, the determining module 830 is further configured to: identify multiple text items from an image that are associated with multiple objects respectively; and, for a first object among the multiple objects, determine a first type based on a first text item associated with the first object among the multiple text items.
[0105] According to some implementations of this disclosure, the determining module 830 is further configured to: identify a second object in an image in physical space to determine the location of the second object.
[0106] According to some implementations of this disclosure, the determining module 830 is further configured to: construct a prompt word in the image of the first object in the physical space and the first type of the first object, the prompt word being used to determine the location of the object of the first type in the image of the physical space; and determine the location of the second object based on the response of the prompt word by a machine learning model.
[0107] According to some implementations of this disclosure, the execution module 840 is further configured to: determine a motion trajectory from the position of the first object to the first destination position based on an image of the physical space; and cause the robot device to move the first object to the first destination position according to the motion trajectory.
[0108] According to some implementations of this disclosure, the execution module 840 is further configured to: provide a message associated with the first object to a user; and in response to receiving a response from the user to the message, cause the robotic device to move the first object to a first destination location.
[0109] According to some implementations of this disclosure, the execution module 840 is further configured to: in response to determining that the number of a group of objects of a first type among a plurality of objects meets a threshold condition, cause the robot device to acquire a third object; cause the robot device to move a group of objects to a first destination location via the third object.
[0110] Figure 9 shows a block diagram of a device 900 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 900 shown in Figure 9 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 900 shown in Figure 9 can be used to implement the methods described above.
[0111] As shown in Figure 9, the computing device 900 is in the form of a general-purpose computing device. Components of the computing device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage devices 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 900.
[0112] Computing device 900 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 920 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 930 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within computing device 900.
[0113] The computing device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 920 may include a computer program product 925 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0114] The communication unit 940 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 900 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0115] Input device 950 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 960 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 900 can also communicate as needed with one or more external devices (not shown) via communication unit 940. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 900, or with any device (e.g., network card, modem, etc.) that enables computing device 900 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).
[0116] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
[0117] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0118] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0119] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0121] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for performing a user task, comprising: receiving a user task from a user, the user task instructing a robotic device to sort a plurality of objects within a first range in a physical space; acquiring an image of the plurality of objects; determining, for a first object of the plurality of objects, a first destination location of the first object in the physical space based on the image; and moving, by the robotic device, the first object to the first destination location.
2. The method of claim 1, wherein acquiring the image comprises: moving, by the robotic device, at least one object of the plurality of objects in response to determining that there is an occlusion relationship between the plurality of objects; and acquiring the image of the plurality of objects.
3. The method of claim 1, wherein the image of the first object is determined based on: adjusting a position of the first object so as to acquire the image of the first object; and adjusting a position of an acquisition device used to acquire the image so as to acquire the image of the first object.
4. The method of claim 1, further comprising: acquiring an image of a physical space in which the robotic device is located; locating the first range based on the image; and moving, by the robotic device, to the first range.
5. The method of claim 1, wherein determining the first destination location comprises: determining, based on the image, a first type of the first object; determining a location of a second object of the first type in the physical space; and determining the first destination location based on the location of the second object.
6. The method of claim 5, wherein determining the first type of the first object further comprises: identifying, from the image, a plurality of textual items respectively associated with the plurality of objects; determining, for a first object of the plurality of objects, the first type based on a first textual item of the plurality of textual items that is associated with the first object.
7. The method of claim 5, wherein determining the location of the second object comprises at least one of: identifying the second object in an image of the physical space to determine the location of the second object.
8. The method of claim 5, wherein determining the location of the second object comprises at least one of: constructing, based on the image of the physical space and the first type of the first object, a prompt for determining a location of an object of the first type in the image of the physical space; and determining the location of the second object based on a response to the prompt by a machine learning model.
9. The method of claim 1, wherein moving, by the robotic device, the first object to the first destination location comprises: determining, based on an image of the physical space, a motion trajectory from a location of the first object to the first destination location; and moving, by the robotic device, the first object to the first destination location in accordance with the motion trajectory. 10. The method of claim 1, wherein the robotic device moving the first object to the first destination location comprises: providing a message associated with the first object to the user; and in response to receiving a response to the message from the user, the robotic device moving the first object to the first destination location.
11. The method of claim 5, wherein the robotic device moving the first object to the first destination location comprises: in response to determining that a quantity of a group of objects of the first type among the plurality of objects satisfies a threshold condition, the robotic device retrieving a third object; the robotic device moving the group of objects to the first destination location via the third object.
12. An apparatus for performing a user task, comprising: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to sort a plurality of objects within a first range in a physical space; a retrieving module configured to retrieve an image of the plurality of objects; a determining module configured to determine, for a first object among the plurality of objects, a first destination location of the first object in the physical space based on the image; and an executing module configured to cause the robotic device to move the first object to the first destination location.
13. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions which when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1-11.
14. A computer-readable storage medium having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to any one of claims 1-11.
15. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-11.
Citation Information
Patent Citations
Automatic food material classification scheduling robot and control method thereof
CN112476443A
Restaurant tableware cleaning-cleaning intelligent system and control method thereof
CN115500755A
Sundry cleaning robot system
CN116709962A
Household intelligent arrangement robot
CN117140476A
Method and device for identifying object from image, equipment and medium
CN117992629A