Method and apparatus for executing user task, and device and medium
By receiving user tasks, acquiring images, identifying candidate objects, and utilizing interactive and machine learning models, the robot system can accurately execute user needs in complex environments, solving the problem that robots cannot understand complex instructions in existing technologies, and achieving higher task execution accuracy and personalized services.
Patent Information
- Application Number
- PCT/CN2024/106258
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-22
AI Technical Summary
Existing robot systems are unable to perform different tasks according to user needs in complex application scenarios, and have difficulty understanding and executing complex user instructions.
By receiving user tasks, acquiring images of the physical space, identifying multiple candidate objects, eliminating ambiguity using interactive and machine learning models, and dynamically selecting the target object by combining user historical data and current state.
It improves the flexibility and accuracy of robots in performing tasks in complex environments, provides personalized and context-aware services, and enhances the intelligence and adaptability of robotic devices.
Smart Images

Figure CN2024106258_22012026_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices, and media for performing user tasks Technical Field
[0001] The exemplary implementations of this disclosure generally relate to the field of robotics, and more particularly to methods, apparatus, devices, and computer-readable storage media for using robots to perform user tasks. Background Technology
[0002] Robotics technology has developed rapidly and is widely used in various technological fields. Currently, a variety of specialized robotic devices have been developed. For example, in industrial environments, robots can be used to perform various tasks such as processing, grasping, sorting, and packaging. In home environments, for instance, robotic vacuum cleaners and window cleaning robots have been developed. However, existing robotic systems have some limitations. In complex application scenarios, robots can typically only perform pre-set fixed tasks and cannot perform different user tasks according to user needs, which greatly reduces the practicality of robots.
[0003] Summary of the Invention
[0004] In a first aspect of this disclosure, a method for performing a user task is provided. In this method, a user task is received from a user, instructing a robotic device to acquire a first object. A first image of a first physical space in which the robotic device is located is acquired. In response to determining that the first image indicates that the first physical space includes a plurality of candidate objects corresponding to the first object, the first object is determined from the plurality of candidate objects. The robotic device acquires the first object.
[0005] In a second aspect of this disclosure, an apparatus for performing a user task is provided. The apparatus includes: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to acquire a first object; an acquiring module configured to acquire a first image of a first physical space in which the robotic device is located; a determining module configured to determine the first object from a plurality of candidate objects corresponding to the first object in response to determining that the first image indicates the first physical space includes a plurality of candidate objects corresponding to the first object; and an executing module configured to cause the robotic device to acquire the first object.
[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.
[0008] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent in the following detailed description, taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;
[0012] Figure 2 shows a block diagram of some implementations of the present disclosure for performing user tasks;
[0013] Figure 3 shows a block diagram of an image acquisition process according to some implementations of this disclosure;
[0014] Figure 4 shows a flowchart of the process of invoking a language model according to some implementations of this disclosure;
[0015] Figure 5 shows a flowchart of a method for performing user tasks according to some implementations of this disclosure;
[0016] Figure 6 shows a block diagram of an apparatus for performing user tasks according to some implementations of the present disclosure; and
[0017] Figure 7 shows a block diagram of a device capable of implementing various implementations of the present disclosure. Detailed Implementation
[0018] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0019] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.
[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0022] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0023] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0025] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
[0026] Example Environment
[0027] In recent years, robotics and machine learning technologies have been widely applied in various scenarios. However, robots typically can only perform pre-set, fixed tasks and cannot execute different user tasks according to user needs. In particular, in complex application environments, robotic devices struggle to determine user requirements and thus perform corresponding tasks.
[0028] Simple robotic devices have been developed to perform specific tasks. However, these simple robotic devices cannot understand complex user instructions, nor can they execute the desired tasks according to user instructions within a complex physical space. Therefore, it is desirable to control the robot's operation in an effective way to perform the desired tasks.
[0029] According to an exemplary implementation of this disclosure, a method for performing user tasks is proposed. Referring to Figure 1, which describes an application environment according to an exemplary implementation of this disclosure, Figure 1 shows a block diagram 100 of an application environment according to an exemplary implementation of this disclosure. As shown in Figure 1, a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform various tasks. The physical space 160 can include, but is not limited to, one or more rooms. For example, in a home environment, the physical space 160 can include, but is not limited to, a living room, bedroom, study, kitchen, toilet, etc., or a combination of one or more of the above.
[0030] As shown in Figure 1, the robot device 110 may include multiple parts. For example, the control unit 111 can serve as the control center of the robot device 110, and an application can be loaded into the control unit 111 to control the various parts of the robot device. The user 120 can use the interaction unit 112 to interact with the robot device 110, for example, by inputting control commands to the robot device 110 to perform desired tasks. The robot device 110 may include an arm 113 for performing actions such as grasping and releasing. For example, the arm 113 can grasp an object and move it to a desired position, and so on.
[0031] Alternatively and / or additionally, the robot device 110 may also include a data acquisition unit 114. Here, the data acquisition unit 114 may include various types, such as an image acquisition unit, a sound acquisition unit, etc. Alternatively and / or additionally, the robot device 110 may further include a sensing unit for detecting surrounding objects, for example, detecting the distance between the robot and surrounding objects based on laser light, etc. The robot device 110 may also include a drive unit 115; for example, the robot device 110 may be deployed on a movable base, and the drive unit 115 may drive the wheels of the base to move along a desired path.
[0032] Physical environment 160 may include one or more acquisition units 130, ..., and 132. For example, one or more image acquisition devices may be deployed in a room to acquire images of the room from various angles. Physical environment 160 may include control device 140, which can control one or more acquisition units 130, ..., and 132, etc., via a network (not shown). Alternatively and / or additionally, in a smart home environment, control device 140 can control various electrical devices in physical space 160.
[0033] Alternatively and / or additionally, a machine learning model (e.g., model 150) may be provided to manage physical space 160. It should be understood that although Figure 1 shows model 150 located inside physical space 160, alternatively and / or additionally, model 150 may be located at a remote device outside physical space 160, and control device 140, robotic device 110, or other device may access the remote model 150 via a network.
[0034] As shown in Figure 1, user 120 can instruct robot device 110 to manipulate various objects in physical space 110. Here, objects can be various items in the home environment. For example, user 120 can instruct robot device 110 to find a certain object in physical space 160; or user 120 can instruct robot device 110 to place the found object in a designated location, and so on.
[0035] Summary of the task to be performed
[0036] To at least partially address the shortcomings of the prior art, a method for performing user task 210 is proposed according to an exemplary implementation of the present disclosure. Referring to FIG2, which describes an outline of an exemplary implementation of the present disclosure, FIG2 illustrates a block diagram 200 for performing user task 210 according to some implementations of the present disclosure.
[0037] As shown in Figure 2, the robot device 110 in physical space 160 (also referred to as the first physical space) can receive a user task 210 from user 120. In this case, user task 210 can instruct the robot device 110 to acquire a first object (for ease of description, the first object can be referred to as the target object). For example, user 120 can say "Give me an apple" in natural language; in the example of Figure 2, the first object is "apple" (e.g., object 220). The robot device 110 can acquire a first image of the first physical space where the robot device 110 is located. For example, the first image can be acquired via at least any one of acquisition units 114, 130, ..., and 132.
[0038] Referring to Figure 3 for further details of image acquisition, Figure 3 illustrates a block diagram 300 of an image acquisition process according to some implementations of this disclosure. As shown in Figure 3, a first image (e.g., one or more images 310) of the physical space 160 can be acquired from the acquisition unit 114 at the robot device 110. Since the robot device 110 can move freely in the physical space 160, the acquisition unit 114 can acquire images of various locations in the physical space, thereby facilitating the search for target objects.
[0039] Alternatively and / or additionally, a first image of the physical space 160 can be acquired from acquisition units 130, ..., and 132. Here, acquisition units 130, ..., and 132 can be pre-deployed at designated locations within the physical space 160, such as a corner of the ceiling, etc. In this way, an image of the physical space 160 taken from a top-down angle can be obtained, facilitating an overall understanding of the layout of the physical space 160 and thus facilitating the location of target objects.
[0040] Furthermore, in response to determining that the first image indicates that the first physical space includes multiple candidate objects corresponding to the first object, the robot device 110 can determine the first object from the multiple candidate objects. In this example, the first image may show two candidate objects, red apples and yellow apples, in the first physical space (e.g., a room). To accurately perform the user task 210, the robot device 110 may further interact or analyze. For example, the robot device 110 may ask the user 120, "I see red apples and yellow apples, which color apple do you want?", which will be described in detail below. At this time, based on the user 120's answer, the robot device 110 can obtain the specific first object. For example, if the user 120 specifies a red apple, the robot device 110 will go to the location of the red apple, grab a red apple, and then return to the user 120's location along path 240 and hand the apple to the user 120.
[0041] According to some implementations of this disclosure, the methods described above can be executed on any computing device with computing capabilities. For example, the methods described above can be executed using an application deployed on robot device 110. Alternatively and / or additionally, an application can be deployed on control device 140 to execute the methods described above. Specifically, the powerful processing capabilities of model 150 can be invoked to analyze images, identify candidate objects, and find a first object based on user input. Subsequently, robot device 110 can travel to the location of the red apple, retrieve object 220, return to the location of user 120 along path 240, and provide the retrieved object 220 to user 120.
[0042] Using the exemplary implementation of this disclosure, the robot device 110 can perform user task 210 in complex physical spaces. In this way, the robot device 110 can accurately perform user task 210 in complex scenarios with multiple candidate objects. Through environmental perception, object recognition, and intelligent interaction, the robot device 110 can make correct choices under uncertain conditions. In this way, the flexibility and accuracy of the robot device 110 in performing tasks in complex environments can be improved, thereby completing the expected user task 210.
[0043] Detailed process of executing the task
[0044] Having described an overview of some implementations according to this disclosure, further details regarding the execution of user task 210 will be described below. For ease of description, the following description uses only the example of controlling robot device 110 to retrieve an apple to illustrate further details of the execution of user task 210. Further details regarding the determination of the first physical space will be described below with reference to Figure 4, which shows a block diagram 400 of the process of invoking the language model according to some implementations of this disclosure.
[0045] According to some implementations of this disclosure, the process of determining a first object from multiple candidate objects may include interaction with user 120. This interaction can not only improve the accuracy of task execution but also continuously refine the needs of user 120. Specifically, a first message may be provided to user 120, indicating that a first physical space (e.g., a room) includes multiple candidate objects. For example, when robot device 110 finds red and yellow apples in the room, robot device 110 may ask user 120, "I see red and yellow apples in the room, which kind of apple do you want?", thereby allowing user 120 to clearly know the available options and make a specific choice.
[0046] Subsequently, after user 120 provides a first response to the first message, robot device 110 can determine the first object based on the first response. For example, if user 120 answers "I want a red apple", robot device 110 will determine the red apple as the first object.
[0047] By utilizing the exemplary implementation of this disclosure, this method ensures that the robot device 110 can accurately understand and execute the specific needs of the user 120, avoiding potential misunderstandings and erroneous operations, and improving the accuracy of task execution.
[0048] According to some implementations of this disclosure, in order to provide more detailed and accurate information to user 120, the first message can be generated based on more in-depth information gathering. This process may include the following steps: First, instructing robot device 110 to move to each of the multiple candidate objects to allow robot device 110 to observe each candidate object up close and obtain more detailed information.
[0049] Secondly, the robotic device 110 can collect detailed information about the candidate object. Furthermore, this detailed information may include close-up photographs or the use of various sensors to detect the object's characteristics. For example, for an apple, the robotic device 110 can detect its color, size, shape, ripeness, sweetness, and so on.
[0050] During the process of collecting detailed information, the robot device 110 can perform the following operations: rotate the apple to take photos from all angles; use a spectral analyzer to detect the apple's sugar content; and use a pressure sensor to test the apple's firmness. This detailed information can help the user 120 make a more informed choice.
[0051] Furthermore, based on this collected detailed information, the robot device 110 can generate a more informative first message. For example, the robot device 110 can generate a prompt 410 to tell the user 120: “I see two kinds of apples. The red apple is about 8 cm in diameter, weighs about 200 grams, has a sugar content of 14°Bx, a glossy surface, and looks very fresh and ripe; the yellow apple is slightly smaller, about 7 cm in diameter, weighs about 180 grams, has a sugar content of 12°Bx, and looks slightly firm. Which kind of apple do you want?” Utilizing the exemplary implementation of this disclosure, this detailed description can help the user 120 make a more informed choice.
[0052] According to some implementations of this disclosure, during the process of receiving user task 210, if it is determined that user task 210 is ambiguous, measures can be taken to eliminate such ambiguity. This process involves the complexity of language understanding and interaction, which can be assisted by language model 420. Specifically, when robot device 110 detects ambiguity in user task 210, it can provide user 120 with a second message to eliminate the ambiguity.
[0053] For example, if user 120 says "I want to eat fruit," the task becomes ambiguous because "fruit" is a broad category. In this case, robot device 110 can generate a message such as, "There are many kinds of fruit. What kind of fruit would you like to eat? For example, apples, bananas, oranges, etc." This follow-up question helps clarify user 120's specific needs.
[0054] After user 120 provides a second response to the second message, robot device 110 can update user task 210 based on this response. For example, if user 120 answers "I want to eat an apple," then the originally vague "eat fruit" task is updated to the specific "eat an apple" task.
[0055] By utilizing the exemplary implementation of this disclosure, ambiguity in the task can be effectively eliminated, ensuring that the robot device 110 can accurately understand and execute the intentions of the user 120.
[0056] In this process, language model 420 can help robotic device 110 understand the context of user 120's input, identify potential ambiguities, and generate appropriate follow-up questions. Simultaneously, it can also understand user 120's responses and integrate this information into the updated task.
[0057] According to other implementations of this disclosure, during the process of receiving user task 210, if it is determined that user task 210 is ambiguous and / or it is believed that user task can be further optimized, the robot device 110 can not only simply ask user 120, but also adopt a context-aware method.
[0058] Specifically, the robot device 110 analyzes the user 120's historical behavior patterns, takes into account environmental factors such as the current time and weather, and combines the user 120's health data (if authorized to access it) to perform contextual understanding and reasoning using the language model 420.
[0059] For example, if user 120 says "I want to eat fruit," robot device 110 can generate a prompt 410, such as: "Considering it's summer now and you've had a lot of exercise today, we suggest you choose some fruits with high water content, such as watermelon or grapes. Also, according to your dietary records, you haven't eaten citrus fruits in the past week, so supplementing with vitamin C will be beneficial for you. Which type of fruit do you prefer?"
[0060] By utilizing the exemplary implementation of this disclosure, not only is ambiguity eliminated, but valuable suggestions are also provided to user 120, enabling robot device 110 to become an intelligent assistant.
[0061] According to some implementations of this disclosure, with user permission, historical data of user 120 can also be utilized in the process of determining the first object from multiple candidate objects. This method can improve the efficiency of the robot device 110 in performing tasks, reduce unnecessary interactions, and also make the behavior of the robot device 110 more intelligent.
[0062] Specifically, robot device 110 can access and analyze user 120's historical data, such as past selection records and interest settings. For example, if historical data shows that user 120 more often chooses red apples, or has explicitly stated that they like red apples, then when faced with a choice between red and yellow apples, robot device 110 can prioritize red apples.
[0063] This method not only allows for a simple review of user 120's past preferences but also employs a dynamic learning and prediction model. Specifically, the robotic device 110 can learn the patterns of user 120's taste changes by analyzing user 120's long-term and short-term interest trends, taking into account seasonal factors and special occasions (such as holidays), and predicting new varieties that user 120 might want to try.
[0064] For example, if historical data shows that user 120 usually prefers red apples, but robot device 110 notices that user 120 has recently started trying more varieties of fruit, it can generate a prompt 410 to tell user 120, such as: "I've noticed you've been trying more varieties of fruit lately. While you usually like red apples, considering your recent increased interest in sweet and sour flavors, you might like to try the yellow apples here. They have a unique honey aroma and a crisp texture. Would you like to try them?"
[0065] By utilizing the exemplary implementation of this disclosure, this approach not only considers the user 120's historical interests but also encourages the user 120 to try new things, thereby providing a more personalized and dynamic service. In this way, the robotic device 110 becomes not only a tool for executing commands but also an intelligent robotic device capable of understanding and predicting the user 120's needs.
[0066] According to some implementations of this disclosure, the process of determining the first object from multiple candidate objects can be based on the current state of user 120, the time information of the task, and a more complex understanding of the scene. This method not only considers the immediate needs of user 120, but also combines time and environmental factors, enabling the robotic device 110 to perform tasks more intelligently.
[0067] Specifically, when determining the first object based on the current state of user 120, robot device 110 can determine the activity being performed by user 120 through visual recognition and behavior analysis. For example, it can capture images of user 120 through a camera on robot device 110, or acquire activity data of user 120 through other sensors. If the image shows user 120 preparing noodles, robot device 110 can recognize this state; if user 120 is preparing to drink soup, robot device 110 can also recognize this state. For example, when user 120 asks for utensils, robot device 110 can observe user 120's actions and the surrounding environment. If it recognizes that user 120 is preparing to eat noodles, robot device 110 can select chopsticks or a fork as the most suitable utensils; if it recognizes that user 120 is preparing to drink soup, robot device 110 can select a spoon.
[0068] Furthermore, the robot device 110 performs corresponding actions to obtain the most suitable tableware as determined, by locating the position of the tableware, planning the path 240, and performing a series of steps such as grasping.
[0069] To improve the accuracy of this process, multimodal perception technologies can be used in combination. For example, in addition to visual recognition, sound recognition can be used to determine the user's activity (such as hearing the sound of boiling water), or a thermal sensor can be used to detect the temperature of food, thereby more accurately inferring the type of tableware needed by the user. This multimodal perception method can greatly improve the robot device 110's judgment ability in complex environments.
[0070] Furthermore, when determining the first object based on the time information of user task 210, the robot device 110 can consider not only the current time, but also the user 120's schedule information and habit patterns. The current time information can be obtained through a built-in clock system or by synchronizing with an external time server.
[0071] For example, when user 120 requests a beverage, robot device 110 can consider the following factors: a) Current time: users tend to choose coffee in the morning and water in the evening. b) User 120's schedule: if user 120 has an important meeting coming up, a refreshing beverage can be chosen. c) User 120's habits: if user 120 usually drinks tea in the afternoon, then when a request for a beverage is received in the afternoon, robot device 110 can choose tea.
[0072] Therefore, the robot device 110 can generate a prompt word 410, such as "Please determine what beverage the user wants to drink now," and input this prompt word 410 into a language model 150. The model 150 can comprehensively consider factors such as time information and user habits to give a reasonable judgment.
[0073] Then, based on the output of model 150, robot device 110 can determine the most suitable beverage and perform the corresponding beverage grabbing action.
[0074] Therefore, a dynamic model 150 can be constructed, which continuously learns and updates the habits and interests of user 120. This model 150 may include time series analysis algorithms to predict the needs of user 120 at different points in time. Simultaneously, reinforcement learning techniques can be incorporated to enable the robotic device 110 to continuously optimize its decision-making process based on feedback from user 120.
[0075] Furthermore, in more complex scenarios, the robot device 110 can simultaneously analyze the environmental image (first image) and the user image (second image), combining them with the user task 210 to generate more comprehensive prompts 410. The first image can be a photograph of the surrounding environment taken by the robot device 110, while the second image is a photograph of the user 120's current state. This process can utilize advanced computer vision and natural language processing technologies.
[0076] For example: a) Scene understanding: Deep learning models can be used to identify objects in the environment and the actions of user 120. b) Activity recognition: Temporal models can be used to understand the sequence of activities that user 120 is currently performing. c) Contextual reasoning: Graph neural networks can be used to build a graph of relationships between objects, user 120, and activities, thereby enabling deeper contextual reasoning.
[0077] Therefore, the robot device 110 can generate a comprehensive prompt 410 based on the first image, the second image, and the user task 210. For example, if user 120 requests "get me some drinking utensils," and the first image shows that there are various drinking utensils in the room (such as teacups, coffee cups, wine glasses, etc.), and the second image shows that user 120 is making tea, then the prompt 410 could be "User 120 is making tea. There are teacups, coffee cups, and wine glasses in the room. Please determine which cup user 120 needs most."
[0078] This prompt word 410 can be input into a specifically trained model 150. This model 150 can be a language model, a model specifically trained for this scenario, or a model 150 fine-tuned for home settings and user 120 behavior. The model 150 can return a response like: "Based on the current scenario, user 120 is preparing tea, and the most suitable choice is a teacup. It is recommended to choose a teacup that matches the style of the teapot to ensure overall harmony."
[0079] Furthermore, to further enhance the adaptability and personalization of the robotic device 110, a feedback mechanism can be introduced. After each task is performed, the robotic device 110 can record the user 120's reaction. For example, whether the user 120 accepted the robotic device 110's choice, or whether they requested a different item. This feedback can be used to fine-tune the decision-making model, making it better suited to the preferences and habits of a specific user 120.
[0080] Using the exemplary implementation of this disclosure, the robot device 110 can more accurately understand the needs of the user 120 and make more appropriate choices. This not only improves the accuracy of task execution but also enables the robot device 110 to better adapt to complex daily life scenarios. For example, even if the user 120 does not explicitly request a teacup, the robot device 110 can make the correct judgment based on the context.
[0081] Furthermore, multi-agent collaborative systems can be introduced. In complex home environments, multiple robotic devices 110 or intelligent devices exist. By establishing a distributed decision-making network, these devices can share information and coordinate their actions. For example, when a robotic device 110 is selecting a beverage, it can query other devices for more contextual information, such as the inventory in the refrigerator or the temperature of other rooms, thus making a more comprehensive decision.
[0082] This decision-making method, based on multi-dimensional information (user status, time information, and scene understanding), combined with advanced artificial intelligence technologies (including computer vision, natural language processing, and machine learning), enables the robotic device 110 to perform user tasks 210 more intelligently and flexibly when facing complex and dynamic home environments. This not only improves the accuracy and efficiency of task execution but also provides more personalized and context-aware services.
[0083] It should be understood that although the foregoing description uses a Chinese language environment as an example to illustrate an exemplary implementation of this disclosure, alternatively and / or additionally, the technical solution of the exemplary implementation of this disclosure can be executed in multiple language environments. For example, the robot device 110 can be controlled in environments such as Chinese, English, Japanese, and French. Specifically, the multilingual capabilities provided by machine learning technology can be used to control the robot device 110 in application environments of different languages. Furthermore, although the foregoing description uses retrieving an apple as an example to illustrate the process of using the robot device 110 to perform user task 210, alternatively and / or additionally, the robot device 110 can be controlled to perform other user tasks 210, such as searching for other items in a room, placing an item in a designated location, etc.
[0084] According to some implementations of this disclosure, user 120 can interact with robot device 110 via language, actions, gestures, etc. For example, user 120 can verbally state the user task 210 to be performed, predefine an action to specify user task 210, and so on. Specifically, user 120 can specify to make the action of holding an apple and eating it, and can use this action as the user task 210 to trigger robot device 110 to retrieve the apple. When the action is recognized from the acquired image sequence, robot device 110 can automatically ask user 120 if they need the apple, and if an affirmative answer is received, robot device 110 can retrieve the apple.
[0085] Alternatively and / or additionally, user 120 may interact with robot device 110 via interaction unit 112. For example, user 120 may input a task represented by text and / or images and control robot device 110 to perform the task. Alternatively and / or additionally, user 120 may specify the execution conditions of the task, such as executing the task immediately, executing the task after a predetermined time, or executing the task when predetermined conditions are determined to be met (e.g., after user 120 wakes up), etc.
[0086] According to some implementations of this disclosure, the robot device 110 can provide various messages to the user 120. For example, assuming the robot device 110 finds multiple varieties of apples, it can ask the user 120 which variety they need. Or, assuming the robot device 110 doesn't find apples but only pears, it can ask the user 120 if they need pears, and so on. Alternatively and / or additionally, the robot device 110 can ask the user 120 where the desired object can be found and then go to the location specified by the user 120 to find the desired object. Alternatively and / or additionally, if the desired object cannot be found, the robot device 110 can ask the user 120 if they want to purchase it, and so on.
[0087] According to some implementations of this disclosure, various positioning algorithms can be used to determine the position of robotic devices and various objects in the physical environment. For example, a Global Positioning System (GPS) can be deployed at the robotic device, and satellite signals can be used to determine the precise position of the robotic device. Alternatively and / or additionally, a communication unit can be deployed at the robotic device, and the position of the robotic device can be determined by means of signals between the communication unit and a base station and by utilizing a communication network. Alternatively and / or additionally, a Wi-Fi access point can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspot to determine the position via Wi-Fi signal strength and the known location of the Wi-Fi access point. Alternatively and / or additionally, the communication unit at the robotic device can support Bluetooth functionality, in which case Bluetooth signals and the known locations of Bluetooth devices can be used to determine the position of nearby devices. An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, thereby determining the position of the robotic device.
[0088] Alternate and / or additional locations can be determined using a visual positioning system to pinpoint the location of the robotic device and / or individual objects. A map of the physical space can be pre-acquired, and the locations of each object can be marked on this map. The robotic device can utilize echo detection units to detect distances to surrounding objects and, by combining the acquired images with the physical space map, determine the precise location of each object. Specifically, computer-aided design (CAD) and geographic information systems (GIS) can be used, along with positioning algorithms to determine the location. Alternate and / or additional locations can also be used to deploy tracking units at important objects in the physical space; for example, tracking units can be added to remote controls for household appliances (e.g., television remotes, air conditioner remotes) so that the robotic device can promptly acquire the precise location of important objects, and so on.
[0089] According to some implementations of this disclosure, the robot's initial position and desired destination can be determined based on the methods described above. The robot can determine a path from its initial position to its destination. For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, continuously update the path, enabling the robot to move along the path to its destination.
[0090] According to some implementations of this disclosure, after reaching the destination location, the robotic device can perform a specified task. For example, it can acquire a specified object and move it to the appropriate location. Constraints, i.e., the constraints that should be followed during task execution, can be determined using a language model and / or a knowledge base. For example, an image and corresponding prompts can be acquired, and the image and prompts can be input into the language model, thereby receiving the constraints from the language model. For example, prompts can be determined as: "Based on the following image, determine the constraints that should be followed during the movement of object XXX," or "Please determine the precautions during the movement of object XXX," etc.
[0091] At this point, it can be determined that during the movement of an object (e.g., bottled water, plate, bowl, etc.), the object's original posture should be maintained (e.g., remaining vertical and not tilted). Furthermore, constraints can be input into the motion model, at which point the series of actions output by the motion model will perform the corresponding tasks while ensuring the constraints are met. Using some implementation methods of this disclosure, safety during the operation of robotic devices can be ensured, thereby preventing accidental damage to an object, and so on.
[0092] Using these implementations disclosed herein, the robotic device 110 can perform user tasks 210 more intelligently and flexibly. Through interaction with the user 120, in-depth observation of candidate objects, elimination of ambiguity in the task, and utilization of the user 120's historical data, the robotic device 110 can make more accurate and personalized choices when faced with multiple candidate objects. This not only improves the accuracy and efficiency of task execution but also enables the robotic device 110 to better adapt to complex daily life scenarios, providing more considerate services to the user 120, thereby completing the expected user task 210.
[0093] Example process
[0094] Figure 5 illustrates a flowchart of a method 500 for performing a user task according to some implementations of this disclosure. At block 510, a user task is received from a user, instructing a robotic device to acquire a first object. At block 520, a first image of the first physical space where the robotic device is located is acquired. At block 530, in response to determining that the first image indicates the first physical space includes a plurality of candidate objects corresponding to the first object, the first object is determined from the plurality of candidate objects. At block 540, the robotic device acquires the first object.
[0095] According to some implementations of this disclosure, determining a first object from multiple candidate objects includes: providing a first message to a user, the first message indicating that a first physical space includes multiple candidate objects; and determining the first object based on the user's first response to the first message.
[0096] According to some implementations of this disclosure, the first message is generated based on the following steps: the robot moves to a candidate object among multiple candidate objects; collects detailed information about the candidate object; and generates the first message based on the detailed information.
[0097] According to some implementations of this disclosure, receiving a user task from a user includes: in response to determining that the user task is ambiguous, providing the user with a second message to dispel the ambiguity; and updating the user task based on the user's second response to the second message.
[0098] According to some implementations of this disclosure, determining the first object from multiple candidate objects includes: determining the first object based on the user's historical data.
[0099] According to some implementations of this disclosure, determining the first object from multiple candidate objects includes: determining the first object based on the user's current state.
[0100] According to some implementations of this disclosure, determining the first object from multiple candidate objects includes: determining the first object based on the time information of the user task.
[0101] According to some implementations of this disclosure, determining a first object from multiple candidate objects includes: acquiring a second image of the user; generating a prompt word based on the first image, the second image, and the user task; and determining the first object based on the response of a machine learning model to the prompt word.
[0102] Example devices and equipment
[0103] Figure 6 shows a block diagram of an apparatus 600 for performing a user task according to some implementations of the present disclosure. The apparatus 600 includes: a receiving module 610 configured to receive a user task from a user, the user task instructing a robot device to acquire a first object; an acquisition module 620 configured to acquire a first image of a first physical space in which the robot device is located; a determining module 630 configured to determine the first object from a plurality of candidate objects corresponding to the first object in response to the determination that the first image indicates the first physical space includes a plurality of candidate objects corresponding to the first object; and an execution module 640 configured to cause the robot device to acquire the first object.
[0104] According to some implementations of this disclosure, the determining module 630 is further configured to: provide a user with a first message indicating that a first physical space includes multiple candidate objects; and determine a first object based on the user's first response to the first message.
[0105] According to some implementations of this disclosure, the determining module 630 is further configured to: move the robot device to a candidate among a plurality of candidate objects; collect detailed information about the candidate object; and generate a first message based on the detailed information.
[0106] According to some implementations of this disclosure, the receiving module 610 is further configured to: provide the user with a second message to dispel ambiguity in response to determining that the user task is ambiguous; and update the user task based on the user's second response to the second message.
[0107] According to some implementations of this disclosure, the determining module 630 is further configured to: determine the first object based on the user's historical data.
[0108] According to some implementations of this disclosure, the determining module 630 is further configured to: determine the first object based on the user's current state.
[0109] According to some implementations of this disclosure, the determining module 630 is further configured to: determine the first object based on the time information of the user task.
[0110] According to some implementations of this disclosure, the determining module 630 is further configured to: acquire a second image of the user; generate prompt words based on the first image, the second image, and the user task; and determine a first object based on the response of a machine learning model to the prompt words.
[0111] Figure 7 shows a block diagram of a device 700 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 700 shown in Figure 7 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 700 shown in Figure 7 can be used to implement the methods described above.
[0112] As shown in Figure 7, the computing device 700 is in the form of a general-purpose computing device. Components of the computing device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage devices 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 700.
[0113] Computing device 700 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within computing device 700.
[0114] The computing device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0115] The communication unit 740 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 700 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 700 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.
[0116] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 700 can also communicate as needed with one or more external devices (not shown) via communication unit 740. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 700, or with any device (e.g., network card, modem, etc.) that enables computing device 700 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interface (not shown).
[0117] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
[0118] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0119] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0120] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0122] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for performing a user task, comprising: receiving a user task from a user, the user task instructing a robotic device to retrieve a first object; retrieving a first image of a first physical space in which the robotic device is located; in response to determining that the first image indicates that the first physical space includes a plurality of candidate objects corresponding to the first object, determining the first object from the plurality of candidate objects; and causing the robotic device to retrieve the first object. 2.The method of claim 1, wherein determining the first object from the plurality of candidate objects comprises: providing a first message to the user, the first message indicating that the first physical space includes the plurality of candidate objects; and determining the first object based on a first answer of the user to the first message. 3.The method of claim 2, wherein the first message is generated based on: moving, by the robotic device, to a candidate object of the plurality of candidate objects; collecting detailed information of the candidate object; and generating the first message based on the detailed information. 4.The method of claim 1, wherein receiving the user task from the user comprises: in response to determining that the user task is ambiguous, providing a second message to the user for resolving the ambiguity; and updating the user task based on a second answer of the user to the second message. determining the first object based on historical data of the user. determining the first object based on a current state of the user.
5. The method of claim 1, wherein determining the first object from the plurality of candidate objects comprises: determining the first object based on temporal information of the user task.
6. The method of claim 1, wherein determining the first object from the plurality of candidate objects comprises: 8.The method of claim 1, wherein determining the first object from the plurality of candidate objects comprises:
7. The method of claim 1, wherein determining the first object from the plurality of candidate objects comprises: retrieving a second image of the user; generating a cue word based on the first image, the second image, and the user task; and determining the first object based on a response of a machine learning model to the cue word. 9.An apparatus for performing a user task, comprising: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to retrieve a first object; a retrieving module configured to retrieve a first image of a first physical space in which the robotic device is located; a determining module configured to determine the first object from a plurality of candidate objects corresponding to the first object in response to determining that the first image indicates that the first physical space includes the plurality of candidate objects; and an executing module configured to cause the robotic device to retrieve the first object. 10.An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method of any one of claims 1-8. 11. A computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, causes the processor to carry out the method according to any one of claims 1 to 8.
12. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, carries out the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Object grabbing method and device, computer equipment and storage medium
CN115713514A
Data processing method and device, equipment and computer medium
CN117484502A
Target object recognition method, object recognition model training method, target object processing method and information processing method
CN117809121A
Method and device for identifying object from image, equipment and medium
CN117992629A
Method and System for Training a Robot
US20230330847A1