Method and apparatus for executing user task, device, and medium

By receiving user tasks and utilizing machine learning models and image recognition technology, robotic devices can identify and acquire related objects in complex environments, solving the problem that existing robotic devices struggle to perform multiple user tasks and achieving more efficient and intelligent task execution.

WO2026016136A1PCT designated stage Publication Date: 2026-01-22BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/106250
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-01-22

Smart Images

  • Figure CN2024106250_22012026_PF_FP_ABST
    Figure CN2024106250_22012026_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for executing a user task, a device, and a medium. The method for executing a user task comprises the following steps: receiving a user task from a user, the user task instructing a robotic device to acquire a first object; determining a second object associated with the first object; and in response to determining that a first image of a physical space where the robotic device is located indicates that a first physical space contains the first object and the second object, instructing the robotic device to acquire the first object and the second object. The method allows a robotic device to execute user tasks in complex physical spaces, giving the robotic device greater flexibility and accuracy in executing tasks in complex environments, thereby enabling the completion of anticipated user tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices, and media for performing user tasks Technical Field

[0001] The exemplary implementations of this disclosure generally relate to the field of robotics, and more particularly to methods, apparatus, devices, and computer-readable storage media for using robots to perform user tasks. Background Technology

[0002] Robotics technology has developed rapidly and is widely used in many technological fields. Various specialized robotic devices have been developed; for example, in industrial environments, robots can perform a variety of tasks such as processing, grasping, sorting, and packaging. In home environments, for instance, robotic vacuum cleaners and window cleaning robots have been developed. However, robots typically can only perform pre-set, fixed tasks and cannot perform different user-defined tasks according to user needs.

[0003] Summary of the Invention

[0004] In a first aspect of this disclosure, a method for performing a user task is provided. In this method, a user task is received from a user, the user task instructing a robotic device to acquire a first object; a second object associated with the first object is determined; and in response to a first image indicating that the first physical space contains both the first and second objects, the robotic device is instructed to acquire the first and second objects.

[0005] In a second aspect of this disclosure, an apparatus for performing a user task is provided. The apparatus includes: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to acquire a first object; a determining module configured to determine a second object associated with the first object; and an acquiring module configured to, in response to a first image indicating that a first physical space in which the robotic device is located includes the first object and the second object, instruct the robotic device to acquire the first object and the second object.

[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.

[0008] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] In the following detailed description, the above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent, taken in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;

[0012] Figure 2 shows a block diagram of some implementations of the present disclosure for performing user tasks;

[0013] Figure 3 shows a block diagram of an image acquisition process according to some implementations of this disclosure;

[0014] Figure 4 shows a flowchart of the process of invoking a language model according to some implementations of this disclosure;

[0015] Figure 5 shows a block diagram of the process of invoking the language model according to some other implementations of this disclosure;

[0016] Figure 6 shows a block diagram of the process of invoking the action model according to some implementations of this disclosure;

[0017] Figure 7 shows a block diagram of the process of obtaining an object according to some implementations of this disclosure;

[0018] Figure 8 shows a flowchart of a method for performing user tasks according to some implementations of this disclosure;

[0019] Figure 9 shows a block diagram of an apparatus for performing user tasks according to some implementations of the present disclosure; and

[0020] Figure 10 shows a block diagram of a device that can implement various implementations of the present disclosure. Detailed Implementation

[0021] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0022] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.

[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0025] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0026] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0027] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0028] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.

[0029] Example Environment

[0030] In recent years, robotics and machine learning technologies have been widely applied in various scenarios. However, robots typically can only perform pre-set, fixed tasks and cannot execute different user tasks according to user needs. In particular, in complex application environments, robotic devices struggle to determine user requirements and thus perform corresponding tasks.

[0031] Simple robotic devices have been developed to perform specific tasks. However, these devices cannot understand complex user instructions, nor can they execute the desired tasks according to user commands in complex physical spaces. Therefore, it is desirable to control the robot's operation in an effective way to perform the desired tasks.

[0032] According to an exemplary implementation of this disclosure, a method for performing user tasks is proposed. Referring to Figure 1, which describes an application environment according to an exemplary implementation of this disclosure, Figure 1 shows a block diagram 100 of an application environment according to an exemplary implementation of this disclosure. As shown in Figure 1, a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform various tasks. The physical space 160 can include, but is not limited to, one or more rooms. For example, in a home environment, the physical space 160 can include, but is not limited to, a living room, bedroom, study, etc., or a combination of one or more of the above. As another example, in a teaching environment, the physical space 160 can include, but is not limited to, a classroom, laboratory, library, etc.

[0033] As shown in Figure 1, the robot device 110 may include multiple parts. For example, the control unit 111 can serve as the control center of the robot device 110, and an application can be loaded into the control unit 111 to control the various parts of the robot device. The user 120 can use the interaction unit 112 to interact with the robot device 110, for example, by inputting control commands to the robot device 110 to perform desired tasks. The robot device 110 may include an arm 113 for performing actions such as grasping and releasing. For example, the arm 113 can grasp an object and move it to a desired position, and so on.

[0034] Alternatively and / or additionally, the robot device 110 may also include a data acquisition unit 114. Here, the data acquisition unit 114 may include various types, such as an image acquisition unit, a sound acquisition unit, etc. Alternatively and / or additionally, the robot device 110 may further include a sensing unit for detecting surrounding objects, for example, detecting the distance between the robot and surrounding objects based on laser light, etc. The robot device 110 may also include a drive unit 115; for example, the robot device 110 may be deployed on a movable base, and the drive unit 115 may drive the wheels of the base to move along a desired path.

[0035] Physical space 160 may include one or more acquisition units 130, ..., and 132. For example, one or more image acquisition devices may be deployed in a room to acquire images of the room from various angles. Physical space 160 may include control device 140, which can control one or more acquisition units 130, ..., and 132, etc., via a network (not shown). Alternatively and / or additionally, in a smart home environment, control device 140 can control various electrical devices in physical space 160.

[0036] Alternatively and / or additionally, a machine learning model (e.g., model 150) may be provided to manage physical space 160. It should be understood that although Figure 1 shows model 150 located inside physical space 160, alternatively and / or additionally, model 150 may be located at a remote device outside physical space 160, and control device 140, robotic device 110, or other device may access the remote model 150 via a network.

[0037] Model 150 may include one or more models. If model 150 includes multiple models, these multiple models may include multiple types of models. Model 150 may, for example, include at least a language model (LM) and an action model. The language model, by learning from a large corpus, can possess question-answering capabilities. The action model can control the robotic device 110 to perform various actions. Model 150 may also include, for example, an image recognition model, a text recognition model, and so on.

[0038] As shown in Figure 1, user 120 can instruct robot device 110 to manipulate various objects in physical space 160. Here, objects can be various items in the home environment. For example, user 120 can instruct robot device 110 to find a certain object in physical space 160; or user 120 can instruct robot device 110 to place the found object in a designated location, and so on.

[0039] Summary of the task to be performed

[0040] To at least partially address the shortcomings of the prior art, a method for performing user tasks is proposed according to an exemplary implementation of this disclosure. Referring to Figure 2, which describes an overview of an exemplary implementation of this disclosure, Figure 2 illustrates a block diagram 200 for performing user tasks according to some implementations of this disclosure.

[0041] As shown in Figure 2, the robot device 110 in the physical space 160 (also referred to as the first physical space) can receive a user task 210 from the user 120. In this case, the user task 210 can instruct the robot device 110 to retrieve a first object 220. For example, in the example of Figure 2, the first object 220 can be an "English book," and the user 120 can say "Give me the English book" in natural language to the interaction unit 112. The speech recognition module built into the robot device 110 receives and parses this speech instruction, converting it into an executable user task 210.

[0042] After receiving the user task 210, the robot device 110 can directly give the first object 220 to the user 120, or it can identify a second object 230 (e.g., an English exercise book) that is closely related to the first object 220 mentioned in the user task 210 (e.g., an English book). To understand the task more intelligently, the robot device 110 can call the machine learning model 150 to analyze the user 120's historical behavior. For example, whenever the user 120 picks up the English book, they often pick up the English exercise book to practice. Based on historical data, the model 150 predicts that the user 120 may want to obtain both the "English book" and the "English exercise book" at the same time, thus confirming the English exercise book as the second object.

[0043] Alternatively and / or additionally, when user 120 issues the command "Get an English book" to robot device 110, interaction unit 112 can capture this voice command. The natural language processing (NLP) algorithm built into control unit 111 begins parsing the text to understand the meaning of the command. Model 150 can also access an object relation database that stores information about the relationships between various items. For example, the database might record the high frequency of co-occurrence between English books and English exercise sets, indicating a close connection between them in usage scenarios. Upon receiving the command "English book," model 150 can query the database and find that the English exercise set, as an object that frequently appears alongside the English book, can be obtained simultaneously in this task, thus identifying the English exercise set as the second object 230.

[0044] Alternatively and / or additionally, Model 150 can also analyze the names and attributes of items to identify potential associations. For example, the name "English Exercises" itself implies that it is learning material related to "English books." Through the keywords "English" and "exercises" in the name, the model can identify that the two objects belong to the same learning domain, thus determining that there is an association between them.

[0045] Next, the control unit 111 of the robot device 110 can invoke an image acquisition device, such as a camera, from at least one of the acquisition units 114, 130, ..., and 132 to acquire a first image 240 of the physical space 160. For example, the robot device 110 can scan the first physical space 160 using a camera on its head or body. With the support of the model 150, the robot device 110 can identify specific object shapes, colors, and text from the complex first image 240, thereby accurately locating the positions of the English book and the English exercise set.

[0046] In response to a first image 240 indicating that the physical space where the robot device 110 is located includes a first object 220 and a second object 230 in the first physical space 160, the robot device 110 can be instructed to acquire the first object 220 and the second object 230. For example, after determining that an English book and an English exercise book are present in the first image 240, the control unit 111 of the robot device 110 can plan a path to reduce the travel distance and time consumption. The drive unit 115 of the robot device 110 then activates and moves along the planned path to the vicinity of the English book. Upon reaching the destination, the arm 113 of the robot device 110 extends and uses the gripping device at its end to pick up the English book. Then, the robot device 110 repeats the above process, moving to the location of the English exercise book and picking it up.

[0047] Finally, the robot device 110 can carry the English book and English exercise book to the location of the user 120, place the first and second objects on the desk, or deliver the first object 220 and the second object 230 to the user 120. Alternatively and / or additionally, after the robot device 110 has completed its actions, it can report to the user 120 through the interaction unit 112 that the task has been completed and awaits further instructions.

[0048] According to some implementations of this disclosure, the methods described above can be executed on any computing device with computing capabilities. For example, the methods described above can be executed using an application deployed on robot device 110. Alternatively and / or additionally, an application can be deployed on control device 140 to execute the methods described above. Specifically, the powerful processing capabilities of model 150 can be invoked to determine a second object 230 closely associated with the first object 220, and to locate the first object 220 and the second object 230 from the first image 240.

[0049] Using the exemplary implementations of this disclosure, robotic devices can perform user tasks in complex physical spaces. In this way, by receiving user tasks, analyzing the tasks, locating objects, planning paths, and executing actions, the robotic device can intelligently complete the task of acquiring a first object and an associated second object. The robotic device can anticipate potential user needs and provide more intelligent and human-like services. This approach improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby fulfilling the intended user tasks.

[0050] Detailed process of executing the task

[0051] Having outlined some implementations according to this disclosure, further details regarding the execution of user tasks will be described below. For ease of description, the following example uses the control of robot device 110 to obtain an English book and an English exercise set as an illustration to further illustrate the execution of user tasks.

[0052] According to some implementations of this disclosure, the first image can come from at least one of the following: an acquisition unit at the robot device 110, or an acquisition unit in the first physical space. Referring to FIG3 for further details of image acquisition, FIG3 shows a block diagram 300 of the image acquisition process according to some implementations of this disclosure. As shown in FIG3, a first image (e.g., one or more images 310) of the physical space 160 can be acquired from the acquisition unit 114 at the robot device 110. Since the robot device 110 can move freely in the physical space 160, the acquisition unit 114 can acquire images from various locations in the physical space, thereby facilitating the search for the target object.

[0053] Alternatively and / or additionally, a first image of the physical space 160 can be acquired from acquisition units 130, ..., and 132. Here, acquisition units 130, ..., and 132 can be pre-deployed at designated locations within the physical space 160, such as in a study, etc. In this way, richer data can be acquired from multiple perspectives.

[0054] According to some implementations of this disclosure, when determining a second object 230 associated with the first object 220, a first prompt word 410 can be obtained based on user task 210 and user information 120-1 of user 120. This first prompt word 410 is used to determine the second object 230. Next, the robot device 110 can receive a first response from the machine learning model 150 to the first prompt word 410 to determine the second object 230. The prompt word could be, for example, "Please determine other books closely associated with 'English book'", and the prompt word 410 can be submitted to the model 150.

[0055] As shown in Figure 4, when user 120 issues user task 210, "Give me the English book," through interaction unit 112, the control unit 111 of robot device 110 can parse the literal meaning of user task 210, namely, locate and retrieve the English book. Simultaneously, robot device 110 can also access the user information database to analyze user 120's historical behavior, interests, and usage habits. For example, the database records that user 120 has a high probability (e.g., a preset probability threshold) of simultaneously using an English book and an English exercise book when learning English. Based on user task 210 (e.g., "Give me the English book") and user information 120-1 (e.g., Grade 7 - Alice), robot device 110 can obtain the first prompt word 410. "English book" as the core vocabulary, combined with the highly relevant "English exercise book" derived from user behavior analysis, constitutes the first prompt word 410. This prompt word 410 not only contains the direct instruction of user task 210 but also implicitly includes possible subsequent needs of user 120.

[0056] Next, the robot device 110 sends the first prompt word 410 to the machine learning model (e.g., language learning model 420), requesting the language model 420 to make a first response based on this first prompt word 410. After receiving the first prompt word 410, the language model 420 analyzes and reasons, and returns a response (i.e., the first response), confirming that "English exercise set" and "English book" have a strong correlation in the current context and should be regarded as the second object 230. The robot device 110 receives and parses the first response of the language model 420, confirming that "English exercise set" is the second object 230 related to "English book" in user task 210.

[0057] Using some implementations of this disclosure, the robot device 110 can acquire and analyze prompt words and interact with the language model 420, providing the ability to handle complex tasks. Furthermore, by determining the second object 230 based on user information 120-1 and user task 210, the robot device 110's task execution capability in complex environments is enhanced, and more considerate and efficient services are provided to the user 120.

[0058] According to some implementations of this disclosure, the first prompt word can also be obtained based on the first image. As shown in Figure 5, when the robot device 110 receives the user task 210 "Give me the English book," the robot device 110 first moves to the study and uses its acquisition unit 114, for example, through a camera or vision sensor, to capture a first image 310 of the physical space 160. Alternatively and / or additionally, the robot device 110 or its connected image processing module can preprocess the first image 310, including adjusting brightness, contrast, etc., to improve image quality. Next, an object detection algorithm is used to identify various objects in the first image 310, including but not limited to the English book.

[0059] After determining the location of the English book, the robotic device 110 can also construct a first cue word 510, such as, "Please identify the object in the following image that is associated with the English book." Here, the first cue word 510 not only includes a text description but also a description of the first image, which can guide the model 520 to identify other items in the image that are closely related to the English book.

[0060] Next, the robot device 110 sends the first image 310 and the first cue word 510 to the model 520, requesting the language model 520 to analyze the image for a second object 230 related to an English book. Upon receiving the first cue word 510, the model 520 uses its deep learning algorithm to analyze the first image 310 and identify items highly associated with English books, such as an English workbook. Then, the model returns a response confirming that "English workbook" is the second object 230. Finally, based on the response from the model 520, the robot device 110 confirms that the English workbook is the second object 230.

[0061] Using some implementations of this disclosure, robot device 110 can construct a first prompt word 510 using image information, and then determine a second object associated with the first object, which helps to improve the task understanding and execution capabilities of robot device 110.

[0062] According to some implementations of this disclosure, in response to the first image indicating that the physical space includes a plurality of second objects associated with the first object, a second object can be selected from the plurality of second objects. For example, there may be a plurality of other objects associated with the first object (e.g., an English book), such as an English exercise book, an English notebook, a vocabulary book, etc., in the first image.

[0063] After determining the existence of multiple second objects, the robot device 110 needs to further determine which second objects are most relevant or most likely to be needed by the user 120. In some implementations, the robot device 110 can determine specific second objects based on user interests. For example, if a user's historical behavior indicates a tendency to use an English exercise book while acquiring an English book, then the robot device 110 will prioritize the English exercise book as a second object. In other implementations, by analyzing the relative positions and states of the objects in the first image, the robot device 110 can infer which second objects are most likely to be used together with the English book. For example, the second object closest to the first object may have been used together previously.

[0064] Alternatively and / or additionally, machine learning models can be used to analyze the first image and user information 120-1 to predict the user's possible subsequent needs, thereby making the best choice. For example, if the user is a primary school student, then English books and workbooks for primary school level will be provided to the user first. As another example, if the user is a secondary school student, then English books and workbooks for secondary school level will be provided to the user first.

[0065] Using some implementations of this disclosure, the robot device can not only identify multiple second objects associated with the first object in the physical space 160, but also select the second object that best meets the user's needs from the multiple second objects, which helps to improve the efficiency and accuracy of the robot device 110 in performing its tasks.

[0066] According to some implementations of this disclosure, in response to the determination that a first image of physical space 160 indicates that the first physical space includes a first object but does not include a second object, the robot device can be instructed to acquire the first object. For example, through analysis of the first image, the robot device 110 can identify the presence of a first object (e.g., an English book) in physical space 160, but does not detect the presence of a second object (e.g., an English exercise book). In this case, the robot device 110 can be instructed to perform the initial task, i.e., acquire only the first object (the English book).

[0067] In this way, when the robot device 110 confirms that the first object exists but the second object does not exist, the robot device 110 will directly focus on acquiring the first object without performing an additional search for the second object. This allows the robot device 110 to accurately execute the user task 210 and effectively utilize resources.

[0068] According to some implementations of this disclosure, when the first image of physical space 160 indicates that the first physical space includes the first object but not the second object, a message can also be provided to user 120 to inquire about the location of the second object in physical space 160. Next, in response to receiving a response from user 120 to the message, the robot device 110 can be instructed to retrieve the second object based on the response.

[0069] For example, robot device 110 confirms the presence of a first object (e.g., an English book) by analyzing a first image in physical space 160, but fails to identify a related second object (e.g., an English workbook) in the image. Robot device 110 can further interact intelligently with user 120, supplementing the lack of environmental perception through user 120's responses.

[0070] Next, after user 120 receives the message, it can provide information about the location of the second object. For example, user 120's response to the message could be "The English exercise book is on the table in the study." Here, user 120's response contains key information needed for robot device 110 to complete its task. After receiving and parsing user 120's response, robot device 110 can incorporate this new information into its task planning, replan its path, and locate and retrieve the second object, such as the English exercise book.

[0071] Using some implementations of this disclosure, the robot device 110 can dynamically adjust its ability to perform tasks based on feedback from the user 120 during task execution.

[0072] According to some implementations of this disclosure, user 120 can also control robot device 110 to read text aloud to assist user 120 in reading. For example, if the first object is a book, user 120 can request robot device 110 to convert the text information in the book into speech. In some implementations, the request from user 120 can be completed via voice command, touchscreen operation, or any other user interface.

[0073] Next, upon receiving a request from user 120, robot device 110 begins to recognize the text content in the book. For example, robot device 110 scans the book pages using optical character recognition technology and converts the printed text into digital text format, thus forming first text data. Robot device 110 then converts the first text data into first audio data. For example, robot device 110 uses text-to-speech (TTS) technology to synthesize the text information into human-readable speech, allowing user 120 to hear the content of the book without having to read it directly.

[0074] Using some of the implementation methods disclosed herein, the robot device 110 can complete the conversion from book text to voice output after receiving instructions from the user 120. This can provide a way for visually impaired users 120 to obtain written information, and can also allow busy users or those who prefer listening to books to access book content while doing other things, thus expanding the ways to access book content and the usage scenarios.

[0075] According to some implementations of this disclosure, when requesting the robot device 110 to convert text information in a book into speech, it can also be requested to convert text data at a specific location in the book into audio data. For example, user 120 can issue a specific request to robot device 110, and robot device 110 determines the location of the first text data in the book based on the request. Next, the first text data at the corresponding location is converted into first audio data. Here, the first text data at the corresponding location can be, for example, a specific chapter, paragraph, page, or even several lines of text. User 120 can issue this request through voice commands, touchscreen input, or other interactive methods, and the request can include the book's title, author's name, page number, chapter title, or keywords, etc.

[0076] Next, after receiving the request from user 120, robot device 110 converts the text data at the corresponding location into voice output. This method can enhance the application prospects of robot device 110 in fields such as reading assistance, information retrieval, and entertainment, and can provide users 120 with more convenient, efficient, and personalized reading experiences for different needs.

[0077] According to some implementations of this disclosure, the robot device 110 can also assist the user 120 in consulting the content of the second object, thereby improving the user 120's learning efficiency. For example, the robot device 110 can identify second text data associated with the first text data in the second object. Then, the second text data is provided to the user 120.

[0078] As an example, during the process of robot device 110 assisting user 120 in reading the first text data, robot device 110 can further play its auxiliary role by supplementing and deepening user 120's learning experience by consulting the content of a second object. For example, robot device 110 uses its information retrieval and analysis capabilities to find second text data related to the first text data from the second object. Next, robot device 110 can present the second text data to user 120, for example, through voice reading, screen display, or other suitable output methods.

[0079] By utilizing some implementation methods disclosed herein, the robotic device 110 can not only increase the depth and breadth of user 120's learning, but also improve user 120's learning efficiency. User 120 does not need to interrupt reading to search for relevant information on their own, thereby improving their focus on learning knowledge.

[0080] According to some implementations of this disclosure, the second text data may be, for example, exercises for a portion of a chapter, which the user 120 can practice. For instance, when the user 120 is learning a particular topic or chapter, the robotic device 110 can provide a series of exercises related to the current learning content, which can help the user 120 consolidate and test their understanding and mastery of the learned knowledge.

[0081] After user 120 completes the exercises and submits their answers, robot device 110 analyzes the received responses, such as checking the correctness, completeness, and rationality of the solution approach. Robot device 110 can evaluate each exercise answer submitted by user 120, such as determining whether the answer is right or wrong, etc.

[0082] By utilizing some implementation methods of this disclosure, the robot device 110 can help the user 120 to instantly check their learning results by providing exercises closely related to the learning content, while the automatic grading function ensures that the user 120 can obtain timely and accurate feedback, thereby promoting the improvement of learning effectiveness and self-correction.

[0083] According to some implementations of this disclosure, a motion model can be used to determine the specific actions to be performed by the robot device. See Figure 6 for further details, which shows a block diagram 600 illustrating the process of invoking a motion model according to some implementations of this disclosure. As shown in Figure 6, a motion model 630 can be provided, which can determine the specific actions to be performed by the robot device based on the current state and instructions of the robot device. This motion model can be a pre-trained and fine-tuned model.

[0084] It should be understood that the current state may include data from multiple aspects, such as an image of the robot device, an image of the robot device's environment, pose data of the robot arm (e.g., the positions of the robot arm's joints (POS1, ...)), and the state of the tool (e.g., a gripper, a cutting tool, etc.) fixed to the end of the robot arm. For example, 0 can be used to represent the gripper's closed state, and 1 can be used to represent the gripper's open state. Instructions and the current state can be input into the motion model 630, which then uses the motion model to determine the action to be performed by the robot device based on the instructions and the current state. Here, the action can represent the difference between the robot device's current pose and the next pose, and the difference between the tool's current state and the next state, etc.

[0085] An instruction 610 (e.g., "get an English book") can be input to the motion model 630. Here, the instruction 610 can be expressed in natural language, and the instruction 610 can be determined from the response of the language model. Furthermore, the current state of the robot device can be acquired, and the motion model 630 can determine the corresponding action 640 based on the input data. For example, the orientation, position, speed, acceleration, etc., of each joint in the arm, and / or the wheels and / or other movable devices of the robot device at the next time point can be determined. Furthermore, the determined action 640 can be used to control the state of the robot device at the next time point.

[0086] Using some implementation methods disclosed herein, a relationship can be established between the language model and the action model, and the user's initial input, expressed in natural language, can be converted into specific actions that can be performed by the robotic device. In this way, the actions of the robotic device can be precisely controlled, thereby executing the user task with higher efficiency.

[0087] According to some implementations of this disclosure, if the first object is obscured by other objects, those objects can be removed first, and then the first object can be retrieved. Specifically, in response to determining that the first image indicates the first object is obscured by a third object in the first physical space, the third object can be moved to retrieve the first object. See Figure 7 for further details, which shows a block diagram 700 of the process of moving an object according to some implementations of this disclosure. As shown in Figure 7, in image 710, the first object 720 is an English book to be retrieved, and the third object 740 is located to the left of the first object 720 and obscures it. At this time, the robot device 110 can be instructed to move the third object 740 from its current position to a position that does not obstruct the retrieval of the first object 720.

[0088] According to some implementations of this disclosure, a target location can be determined, and the robotic device can be instructed to move the third object 740 to the target location. At this point, a motion model will generate actions to control the robotic device to move the third object 740 from its current location to the target location. In this way, the robotic device can handle complex problems in complex environments, thereby performing user tasks more accurately.

[0089] According to some implementations of this disclosure, during the movement of a third object, constraints can be determined based on the pose of the third object, and the robot can be instructed to move the third object under these constraints. In this way, it can be ensured that all actions of the robot in complex environments comply with safety regulations.

[0090] According to some implementations of this disclosure, the method of organizing the bookshelf can also be determined, and the robot device 110 can be instructed to organize the remaining objects in the bookshelf according to the above-described method. Specifically, after the robot device has retrieved the first object 720, a new image can be acquired and a new prompt word can be constructed to query the language model for the next instruction. The prompt word can be represented, for example, as: "Please determine the next instruction based on the following image", "What should be done next", etc. The language model can return "Arrange the remaining books neatly", at which point a corresponding action can be generated based on the instruction "Arrange the remaining books neatly" and the current state of the robot device to instruct the robot device to organize the remaining books.

[0091] According to some implementations of this disclosure, after the robot device 110 has finished organizing the remaining books, it can be instructed to retrieve the first and second objects and proceed to the location of the user 120. Specifically, images can be acquired in real time, and the user's position can be located within the images. Furthermore, corresponding instructions can be determined based on the robot device's current position (e.g., position A) and the user's position (e.g., position B). The instruction can then be expressed as: move from position A to position B. At this point, the motion model 630 will generate a corresponding action that controls the robot device 110 to move from position A to position B along a determined trajectory. In this way, the robot device 110 completes the task of "give me the English book."

[0092] It should be understood that although the foregoing description uses a Chinese language environment as an example to illustrate one implementation of this disclosure, alternatively and / or additionally, the technical solution of one implementation of this disclosure can be executed in multiple language environments. For example, the robot can be controlled in environments such as Chinese, English, Japanese, and French. Specifically, the multilingual capabilities provided by machine learning technology can be used to control the robot in application environments in different languages. Furthermore, although the foregoing description of obtaining English books and English exercise sets as examples illustrates the process of using a robotic device to perform user tasks, alternatively and / or additionally, the robotic device can be controlled to perform other user tasks, such as finding other items in a room, placing an item in a designated location, etc.

[0093] According to some implementations of this disclosure, users can interact with the robot device through language, actions, gestures, etc. For example, users can state the user task they wish to perform, predefine a certain action to specify the user task, and so on. When the action is recognized from the acquired image sequence, the robot device can automatically ask the user if they need to obtain learning materials. If a positive response is received, the robot device can retrieve the learning materials.

[0094] Alternatively and / or additionally, a user can interact with the robot device via the interaction unit 112, for example, the user inputs a task represented by text and / or images, and controls the robot device to perform the task. Alternatively and / or additionally, the user can specify the execution conditions of the task, for example, to execute the task immediately, to execute the task after a predetermined time, or to execute the task when predetermined conditions are determined to be met (e.g., after the user leaves school), etc.

[0095] According to some implementations of this disclosure, the robot device can provide users with various messages. For example, assuming the robot device finds multiple types of English books, it can ask the user which type they need. Or, assuming the robot device doesn't find any English books and only finds math books, it can ask the user if they need a math book, and so on. Alternatively and / or additionally, the robot device can ask the user where the desired object can be found and then go to the user-specified location to find the desired object. Alternatively and / or additionally, if the desired object cannot be found, the robot device can ask the user if they want to purchase it, and so on.

[0096] According to some implementations of this disclosure, various positioning algorithms can be used to determine the position of robotic devices and various objects in the physical environment. For example, a Global Positioning System (GPS) can be deployed at the robotic device, and satellite signals can be used to determine the precise position of the robotic device. Alternatively and / or additionally, a communication unit can be deployed at the robotic device, and the position of the robotic device can be determined by means of signals between the communication unit and a base station and by utilizing a communication network. Alternatively and / or additionally, a Wi-Fi access point can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspot to determine the position via Wi-Fi signal strength and the known location of the Wi-Fi access point. Alternatively and / or additionally, the communication unit at the robotic device can support Bluetooth functionality, in which case Bluetooth signals and the known locations of Bluetooth devices can be used to determine the position of nearby devices. An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, thereby determining the position of the robotic device.

[0097] Alternate and / or additional locations can be determined using a visual positioning system to pinpoint the location of the robotic device and / or individual objects. A map of the physical space can be pre-acquired, and the locations of each object can be marked on this map. The robotic device can utilize echo detection units to detect distances to surrounding objects and, by combining the acquired images with the physical space map, determine the precise location of each object. Specifically, computer-aided design (CAD) and geographic information systems (GIS) can be used, along with positioning algorithms to determine the location. Alternate and / or additional locations can also be used to deploy tracking units at important objects in the physical space; for example, tracking units can be added to remote controls for household appliances (e.g., television remotes, air conditioner remotes) so that the robotic device can promptly acquire the precise location of important objects, and so on.

[0098] According to some implementations of this disclosure, the robot's initial position and desired destination can be determined based on the methods described above. The robot can determine a path from its initial position to its destination. For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, continuously update the path, enabling the robot to move along the path to its destination.

[0099] According to some implementations of this disclosure, after reaching the destination location, the robotic device can perform a specified task. For example, it can acquire a specified object and move it to the appropriate location. Constraints, i.e., the constraints that should be followed during task execution, can be determined using a language model and / or a knowledge base. For example, an image and corresponding prompts can be acquired, and the image and prompts can be input into the language model, thereby receiving the constraints from the language model. For example, prompts can be determined as: "Based on the following image, determine the constraints that should be followed during the movement of object XXX," or "Please determine the precautions during the movement of object XXX," etc.

[0100] At this point, it can be determined that during the movement of an object (e.g., bottled water, plate, bowl, etc.), the object's original posture should be maintained (e.g., remaining vertical and not tilted). Furthermore, constraints can be input into the motion model, at which point the series of actions output by the motion model will perform the corresponding tasks while ensuring the constraints are met. Using some implementation methods of this disclosure, safety during the operation of robotic devices can be ensured, thereby preventing accidental damage to an object, and so on.

[0101] Using the exemplary implementations of this disclosure, robotic devices can perform user tasks in complex physical spaces. In this way, by receiving user tasks, analyzing the tasks, locating objects, planning paths, and executing actions, the robotic device can intelligently complete the task of acquiring a first object and an associated second object. The robotic device can anticipate potential user needs and provide more intelligent and human-like services. This approach improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby fulfilling the intended user tasks.

[0102] Example process

[0103] Figure 8 illustrates a flowchart of a method 800 for performing a user task according to some implementations of this disclosure. At block 810, a user task is received from a user, instructing a robotic device to acquire a first object. At block 820, a second object associated with the first object is determined. At block 830, in response to a first image indicating that the physical space where the robotic device is located includes both the first and second objects, the robotic device acquires both the first and second objects.

[0104] According to some implementations of this disclosure, determining the second object includes: obtaining a first prompt word based on the user task and the user's user information, the first prompt word being used to determine the second object; and receiving a first response from a machine learning model to the first prompt word in order to determine the second object.

[0105] According to some implementations of this disclosure, obtaining the first prompt word further includes: obtaining the first prompt word based on the first image.

[0106] According to some implementations of this disclosure, the method 800 further includes: in response to a first image indicating that the physical space includes a plurality of second objects associated with the first object, selecting a second object from the plurality of second objects.

[0107] According to some implementations of this disclosure, the method 800 further includes: in response to a first image indicating that a first physical space includes a first object but does not include a second object, a robotic device acquires the first object.

[0108] According to some implementations of this disclosure, the method 800 further includes: providing a message to a user, the message being used to inquire with the user about the location of the second object in physical space; and in response to receiving a response from the user to the message, the robotic device acquiring the second object based on the response.

[0109] According to some implementations of this disclosure, the first object is a book, and the method 800 further includes: receiving a request from a user to provide voice data associated with the book, the robotic device recognizing first text data in the book; and converting the first text data into first audio data.

[0110] According to some implementations of this disclosure, the method 800 further includes: determining the position of the first text data in the book based on a request; and converting the first text data at the position into first audio data.

[0111] According to some implementations of this disclosure, the second object is a book, and the method 800 further includes: determining second text data associated with the first text data in the second object; and providing the second text data to the user.

[0112] According to some implementations of this disclosure, the method 800 further includes: in response to receiving a user's response to the second text data, determining an evaluation of the response.

[0113] Example devices and equipment

[0114] Figure 9 shows a block diagram of an apparatus 900 for performing a user task according to some implementations of the present disclosure. The apparatus 900 includes: a receiving module 910 configured to receive a user task from a user, the user task instructing a robot device to acquire a first object; a determining module 920 configured to determine a second object associated with the first object; and an acquiring module 930 configured to, in response to a first image indicating that the first physical space in which the robot device is located includes the first object and the second object, cause the robot device to acquire the first object and the second object.

[0115] According to some implementations of this disclosure, the determining module 920 is further configured to: obtain a first prompt word based on the user task and the user's user information, the first prompt word being used to determine a second object; and receive a first response from a machine learning model to the first prompt word in order to determine the second object.

[0116] According to some implementations of this disclosure, the determining module 920 is further configured to: obtain a first prompt word based on the first image.

[0117] According to some implementations of this disclosure, the acquisition module 930 is further configured to: select a second object from the plurality of second objects in response to a first image indicating that the physical space includes a plurality of second objects associated with the first object.

[0118] According to some implementations of this disclosure, the acquisition module 930 is further configured to: in response to a first image indicating that the first physical space includes a first object but does not include a second object, cause the robotic device to acquire the first object.

[0119] According to some implementations of this disclosure, the acquisition module 930 is further configured to: provide a message to a user, the message being used to inquire of the location of the second object in physical space; and in response to receiving a response from the user to the message, cause the robotic device to acquire the second object based on the response.

[0120] According to some implementations of this disclosure, the first object is a book, and the receiving module 910 is further configured to: receive a request from a user to provide voice data associated with the book, enabling the robotic device to recognize first text data in the book; and convert the first text data into first audio data.

[0121] According to some implementations of this disclosure, the receiving module 910 is further configured to: determine the position of the first text data in the book based on a request; and convert the first text data at the position into first audio data.

[0122] According to some implementations of this disclosure, the second object is a book, and the receiving module 910 is further configured to: determine second text data associated with the first text data in the second object; and provide the second text data to the user.

[0123] According to some implementations of this disclosure, the receiving module 910 is further configured to: determine the evaluation of the response in response to receiving a user's response to the second text data.

[0124] Figure 10 shows a block diagram of a device 1000 that can implement various implementations of the present disclosure. It should be understood that the computing device 1000 shown in Figure 10 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 1000 shown in Figure 10 can be used to implement the methods described above.

[0125] As shown in Figure 10, the computing device 1000 is in the form of a general-purpose computing device. Components of the computing device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage devices 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. The processing unit 1010 may be a physical or virtual processor and can perform various processes according to programs stored in memory 1020. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 1000.

[0126] Computing device 1000 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1030 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 1000.

[0127] The computing device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 1020 may include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.

[0128] The communication unit 1040 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1000 can be implemented as a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0129] Input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1060 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 1000 can also communicate with one or more external devices (not shown) via communication unit 1040 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 1000, or with any device (e.g., network card, modem, etc.) that enables computing device 1000 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interface (not shown).

[0130] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.

[0131] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0132] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0133] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0135] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for performing a user task, comprising: receiving a user task from a user, the user task instructing a robotic device to retrieve a first object; determining a second object associated with the first object; and in response to determining that a first image of a physical space in which the robotic device is located indicates that the physical space includes the first object and the second object, the robotic device retrieving the first object and the second object.

2. The method of claim 1, wherein determining the second object comprises: based on the user task and user information of the user, retrieving a first prompt word, the first prompt word used to determine the second object; and receiving a first response of a machine learning model to the first prompt word in order to determine the second object.

3. The method of claim 2, wherein obtaining the first prompt word further comprises: retrieving the first prompt word based on the first image.

4. The method of claim 2, further comprising: in response to the first image indicating that the physical space includes a plurality of second objects associated with the first object, selecting the second object from the plurality of second objects.

5. The method of claim 1, further comprising: in response to determining that a first image of the physical space indicates that the physical space includes the first object but does not include the second object, the robotic device retrieving the first object.

6. The method of claim 5, further comprising: providing a message to the user, the message used to ask the user for a location of the second object in the physical space; and in response to receiving a response of the user to the message, the robotic device retrieving the second object based on the response.

7. The method of claim 1, wherein the first object is a book, and the method further comprises: receiving a request from the user to provide voice data associated with the book, the robotic device identifying first textual data in the book; and converting the first textual data to first audio data.

8. The method of claim 7, further comprising: determining a location of the first textual data in the book based on the request; and converting the first textual data at the location to the first audio data.

9. The method of claim 8, wherein the second object is a book, and the method further comprises: determining second textual data associated with the first textual data in the second object; and providing the second textual data to the user.

10. The method of claim 9, further comprising: in response to receiving a response of the user to the second textual data, determining an evaluation of the response.

11. An apparatus for performing a user task, comprising: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to retrieve a first object; a determining module configured to determine a second object associated with the first object; and a retrieving module configured to, in response to determining that a first image of a physical space in which the robotic device is located indicates that the physical space includes the first object and the second object, the robotic device retrieving the first object and the second object.

12. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1 to 10.

13. A computer readable storage medium having stored thereon a computer program, the computer program causing a processor to implement the method according to any one of claims 1 to 10 when executed by the processor.

14. A computer program product comprising a computer program, wherein the computer program implements the method according to any one of claims 1 to 10 when executed by a processor.

Citation Information

Patent Citations

  • Article recommendation method and device based on semantic recognition, computer equipment and medium

    CN112070586A

  • Article matching determination method and device, equipment and storage medium

    CN112750000A

  • Robot shopping guide method and device, electronic equipment and storage medium

    CN113771048A

  • Article recommendation method and device, electronic equipment and storage medium

    CN117216373A

  • Data processing method and device, equipment and computer medium

    CN117484502A