Method and apparatus for executing user task, and device and medium

By acquiring and identifying media items from electronic devices using robotic equipment, and utilizing language and motion models to identify key information, this technology solves the problem that robotic equipment struggles to perform user tasks in complex environments, enabling personalized and flexible user task execution.

WO2026016141A1PCT designated stage Publication Date: 2026-01-22BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/106256
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing robotic devices struggle to perform multiple tasks in complex environments according to user needs, especially in home environments where they have difficulty understanding complex user instructions and executing corresponding tasks.

Method used

The robot acquires and identifies media items provided by electronic devices, utilizes language and motion models to identify key information and provide relevant data, and dynamically adjusts the content to meet user needs. This includes identifying information such as people, subtitles, scenes, and providers, and acquiring rich data through image acquisition and environmental acquisition units.

Benefits of technology

It improves the flexibility and accuracy of robotic equipment in complex environments, can dynamically adjust content according to user needs, provides personalized and immersive services, and enhances the flexibility and accuracy of user task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024106256_22012026_PF_FP_ABST
    Figure CN2024106256_22012026_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for executing a user task, and a device and a medium The method for executing a user task comprises the following steps: acquiring a media item provided by an electronic device; in response to having received a query request from a user, identifying key information from the media item, wherein the key information comprises at least one of the following: a character, a subtitle, a scene and a sound in the media item, and a provider of the media item; on the basis of the key information, determining related data of the media item; and a robot device providing the related data of the media item to the user. By means of the method, a robot device can execute in a complex physical space a user task involving a complex query, thereby improving the flexibility and accuracy of the robot device in executing tasks in a complex environment, and thus completing the expected user task.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatuses, devices, and media for performing user tasks TECHNICAL FIELD

[0001] Exemplary implementations of the present disclosure generally relate to the field of robotics, and in particular, to methods, apparatuses, devices, and computer-readable storage media for performing user tasks using robots. BACKGROUND

[0002] Robotics technology has been rapidly developed and has been widely used in multiple technical fields. Currently, various special-purpose robotic devices have been developed, for example, in an industrial environment, robots can be used to perform various tasks such as processing, grabbing, sorting, packaging, etc. For another example, in a home environment, a sweeping robot, a glass wiping robot, etc. have been developed. However, robots can usually only perform pre-set fixed tasks and cannot perform different user tasks according to user needs.

[0003] SUMMARY

[0004] In a first aspect of the present disclosure, a method for performing a user task is provided. In the method, a media item provided by an electronic device is obtained. In response to receiving a query request from a user, key information is identified from the media item, the key information including at least any of a character in the media item, a subtitle, a scene, a sound, and a provider of the media item. Relevant data of the media item is determined based on the key information. A robotic device provides the relevant data of the media item to the user.

[0005] In a second aspect of the present disclosure, an apparatus for performing a user task is provided. The apparatus includes an obtaining module configured to obtain a media item provided by an electronic device; an identifying module configured to identify, in response to receiving a query request from a user, key information from the media item, the key information including at least any of a character in the media item, a subtitle, a scene, a sound, and a provider of the media item; a determining module configured to determine relevant data of the media item based on the key information; and a providing module configured to cause a robotic device to provide the relevant data of the media item to the user.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, causes the processor to implement the method according to the first aspect of the present disclosure.

[0008] In a fifth aspect of the disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method according to the first aspect of the disclosure when executed by a processor.

[0009] It is to be understood that the details set forth herein are not intended to limit the key or critical features of the implementations of the present disclosure or to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description, which is given by way of example only. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, aspects and advantages of various implementations of the present disclosure will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which like reference numerals refer to like elements, in which:

[0011] FIG. 1 shows a block diagram of an application environment according to one example implementation of the present disclosure;

[0012] FIGS. 2A and 2B show block diagrams for performing a user task according to some implementations of the present disclosure, respectively;

[0013] FIG. 3 shows a block diagram of a process for determining relevant data from an image according to some implementations of the present disclosure;

[0014] FIG. 4 shows a block diagram of a process for invoking a language model according to some implementations of the present disclosure;

[0015] FIG. 5 shows a block diagram of a process for recognizing an object from an image according to some implementations of the present disclosure;

[0016] FIG. 6 shows a block diagram of a process for invoking an action model according to some implementations of the present disclosure;

[0017] FIG. 7 shows a block diagram of a process for obtaining an object according to some implementations of the present disclosure;

[0018] FIG. 8 shows a flowchart of a method for performing a user task according to some implementations of the present disclosure;

[0019] FIG. 9 shows a block diagram of an apparatus for performing a user task according to some implementations of the present disclosure; and

[0020] FIG. 10 shows a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION

[0021] Implementations of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several implementations of the present disclosure are described, it should be understood that the present disclosure can be embodied in many other forms and should not be construed as limited to the implementations set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It should be understood that the drawings and implementations described are for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0022] In the description of implementations of the present disclosure, the term "includes" and its derivatives mean "including but not limited to". The term "based on" means "based at least in part on". The term "one implementation" or "the implementation" means "at least one implementation". The term "some implementations" means "at least some implementations". Other explicit or implicit definitions can also be included below. As used herein, the term "model" can represent the relationship between various data. For example, the above-mentioned relationship can be obtained based on various technical solutions known at present and / or to be developed in the future.

[0023] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0024] It can be understood that before using the technical solutions disclosed by the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scene of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0025] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be executed will require the acquisition and use of the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that executes the operation of the technical solutions of the present disclosure according to the prompt information.

[0026] As an optional but non-limiting implementation, in response to receiving the active request of the user, the way of sending prompt information to the user, for example, can be the way of pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0027] It can be understood that the above-mentioned notification and user authorization process is only illustrative, and does not limit the implementations of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementations of the present disclosure.

[0028] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of the execution of a subsequent action performed in response to the event or condition is not necessarily strongly correlated with the time at which the event occurs or the condition is satisfied. For example, in some cases, the subsequent action can be performed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action can be performed after a period of time has elapsed since the occurrence of the event or the satisfaction of the condition.

[0029] Example Environment

[0030] In recent years, robotic technology and machine learning technology have been widely applied to multiple application scenarios. However, robots are generally only capable of performing pre-set fixed tasks and are not able to perform different user tasks according to user needs. In particular, in complex application environments, it is difficult for a robotic device to determine user needs and then perform corresponding tasks.

[0031] Simple robotic devices that perform specific tasks have been developed, however, such simple robotic devices are not able to understand complex user instructions and are not able to perform desired tasks according to user instructions in complex physical spaces. At this time, it is desirable to control the operation of the robot in an effective manner and then perform the desired task.

[0032] According to one example implementation of the present disclosure, a method for performing a user task is proposed. Referring to FIG. 1, an application environment according to one example implementation of the present disclosure is described, which shows a block diagram 100 of an application environment according to one example implementation of the present disclosure. As shown in FIG. 1, a robotic device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robotic device 110 to perform multiple tasks. The physical space 160 can include, but is not limited to, one or more rooms. For example, in a home environment, the physical space 160 can include, but is not limited to, a living room, a bedroom, a study, etc., or a combination of one or more of the above.

[0033] As shown in FIG. 1, the robotic device 110 can include multiple parts. For example, a control unit 111 can serve as the control center of the robotic device 110, and an application program can be loaded into the control unit 111 in order to control various parts of the robotic device. The user 120 can use an interaction unit 112 to interact with the robotic device 110, for example, to input control instructions to the robotic device 110 in order to perform a desired task using the robotic device 110. The robotic device 110 can include an arm 113 for performing actions such as grasping, releasing, etc. For example, the arm 113 can grasp an object and move the object to a desired location, etc.

[0034] Alternatively and / or additionally, the robotic device 110 can further include a collection unit 114. Here, the collection unit 114 can include various types, e.g., an image collection unit, a sound collection unit, an environment collection unit, etc. Alternatively and / or additionally, the robotic device 110 can further include a sensing unit for detecting surrounding objects, e.g., can detect the distance between the robotic device and surrounding objects based on laser, etc. The robotic device 110 can further include a driving unit 115, e.g., the robotic device 110 can be deployed on a movable base, and the driving unit 115 can drive the wheels of the base to move along a desired path.

[0035] The physical space 160 can include one or more collection units 130, …, and 132, e.g., one or more image collection devices can be deployed in a room to collect images of the room from various angles. The physical space 160 can include a control device 140, which can control the one or more collection units 130, …, and 132, etc. via a network (not shown). Alternatively and / or additionally, in a smart home environment, the control device 140 can control various electrical appliances in the physical space 160.

[0036] Alternatively and / or additionally, a machine learning model (e.g., the model 150) can be provided to manage the physical space 160. It should be appreciated that although FIG. 1 shows the model 150 located inside the physical space 160, alternatively and / or additionally, the model 150 can be located at a remote device outside the physical space 160, and the control device 140, the robotic device 110, or other devices can access the remote model 150 via a network.

[0037] The model 150 can include one or more models. If the model 150 includes multiple models, the multiple models can include multiple types of models. The model 150 can include at least a language model (LM) and an action model, for example. The language model can have the ability to answer questions by learning from a large amount of corpus. The action model can control the robotic device 110 to perform various actions. The model 150 can further include an image recognition model, a text recognition model, etc., for example.

[0038] As shown in FIG. 1, the user 120 can instruct the robotic device 110 to operate various objects in the physical space 160. Here, the objects can be various items in a home environment, e.g., the user 120 can instruct the robotic device 110 to turn on or off a certain home appliance in the physical space 160; or the user 120 can instruct the robotic device 110 to find a target object and place the found target object to a designated location, etc.

[0039] Summary of performing a task

[0040] To at least partially address the deficiencies in the prior art, according to one example implementation of the present disclosure, a method for performing a user task is proposed. An overview of one example implementation of the present disclosure is described with reference to FIG. 2A and FIG. 2B, which show a block diagram 200A and a block diagram 200B for performing a user task according to some implementations of the present disclosure.

[0041] As shown in FIG. 2A, the robotic device 110 in the physical space 160 (also referred to as, a first physical space) can receive a start request 210 from the user 120. At this time, the start request 210 can instruct the robotic device 110 to start an electronic device 170 in the physical space 160. For example, in the example of FIG. 2A, the electronic device 170 is a television, and the user 120 can speak "turn on the television" in natural language. The robotic device 110, in response to receiving the start request 210 from the user 120, can start the electronic device 170 in the physical space 160.

[0042] As an example, the robotic device 110 can start the television through a control device of the television. For example, the location of the control device or control panel of the television can be identified and located via at least any of the acquisition units 114, 130, …, and 132. Next, the robotic device 110 moves to the vicinity of the control device or control panel so that the arm 113 of the robotic device 110 can reach the control device or control panel. In this process, navigation and path planning need to be performed to avoid colliding with other objects or furniture. Then, when the robotic device 110 reaches the appropriate position, the robotic device 110 uses its arm 113 or fingers to simulate finger movements to press the start button. After the operation is completed, the robotic device 110 can also confirm whether the television has been started. For example, by detecting whether the television screen is lit or whether the television plays a start sound, etc. through the acquisition unit 114. Further, if it is detected that the television cannot be started normally, the robotic device 110 can take emergency measures, such as multiple attempts or broadcast failure information to the user 120, etc.

[0043] As another example, the robotic device 110 can directly start the television. For example, the robotic device 110 can communicate with the television or its control device. For example, the television is connected through infrared (IR) signals, Bluetooth, or Wi-Fi. The robotic device 110 can turn on the television in response to the voice command of the user 120 (i.e., the start request 210).

[0044] As shown in FIG. 2B, after the electronic device 170 is normally started, the robotic device 110 can acquire the media item provided by the electronic device 170. For example, the robotic device 110 can acquire the media item provided by the electronic device 170 via at least any one of the acquisition units 114, 130, …, and 132.

[0045] As shown in FIG. 2B, in response to receiving the query request 220 from the user 120, the robotic device 110 can provide the user 120 with the related data of the media item. For example, the user 120 can speak in natural language “what program is being played”, and the robotic device 110 can provide the user 120 with the acquired related data of the media item. For example, what program is being played on the TV, the provider of the media item, etc.

[0046] According to some implementations of the present disclosure, the above-described method can be executed at any computing device with computing capability. For example, the above-described method can be executed by utilizing an application deployed at the robotic device 110. Alternatively and / or additionally, the application can be deployed at the control device 140 so as to execute the above-described method. Specifically, the powerful processing capability of the model 150 can be invoked to acquire and identify the related data of the media item provided by the electronic device 170 after the electronic device 170 is started. In turn, the robotic device 110 can provide the user 120 with the related data of the media item.

[0047] With the exemplary implementations of the present disclosure, the robotic device can perform the user task in a complex physical space. In this way, the robotic device can start the electronic device within the physical space according to the start request of the user, and acquire the media item provided by the electronic device after the electronic device is started, and in turn provide the user with the related data of the media item upon receiving the query request issued by the user. In this way, the flexibility and accuracy of the robotic device in performing the task in a complex environment can be improved, and in turn the expected user task can be completed.

[0048] Detailed process of performing the task

[0049] Having described the outline according to some implementations of the present disclosure, in the following, more details about performing the user task will be described. For ease of description, in the following, only taking an example of controlling the robotic device 110 to assist the user 120 in operating the TV as an example, more details about performing the user task will be described.

[0050] According to some implementations of the present disclosure, the relevant data includes at least any of audio data, video data, image data, text data. As shown in FIG. 2B, the media item provided by the electronic device 170 can include any content that can be played or displayed through the television, such as a movie, a television program, music, a game, a picture, etc. The media item can be stored on a hard disk inside the television or streamed from an online service provider through the Internet.

[0051] During this process, the user 120 can issue a query request to the robotic device 110 through voice commands or other input methods. For example, during the process of watching a television program, the user 120 can want to view the relevant data of the current program, such as the user 120 can speak in natural language "what program is being played".

[0052] When the robotic device 110 receives the query request 220 from the user, the relevant data of the media item displayed by the television can be provided to the user 120. For example, the robotic device 110 can play audio data through the built-in speaker to provide the user 120 with audio comments, background music, audio track options, etc. related to the media item. For another example, the robotic device 110 is equipped with a display screen and can play video content such as relevant educational videos, entertainment clips, tutorials, etc. to provide the user 120 with video data including trailers, behind-the-scenes, etc. For another example, the robotic device 110 can display text data related to the media item including titles, subtitles, scripts, comments, news articles, etc. through the display screen or convey the text data through text-to-speech technology. In this way, the robotic device 110 can provide information in a visual, auditory, or both ways. Through these different data types and display methods, the robotic device 110 can adapt to various use scenarios to meet the needs of different users.

[0053] According to some implementations of the present disclosure, the robotic device 110 can identify key information 311 from the media item and determine the relevant data based on the key information 311. More details of determining the relevant data are described with reference to FIG. 3, which shows a block diagram 300 of a process of determining the relevant data according to some implementations of the present disclosure. As shown in FIG. 3, for example, the first image (e.g., one or more images 310) of the electronic device 170 can be obtained from the acquisition unit 114 of the robotic device 110. Since the robotic device 110 can move freely in the physical space 160, the acquisition unit 114 can be aligned with the position of the electronic device 170 so that the image 310 of the media item can be obtained.

[0054] Alternatively and / or additionally, the first image of the physical space 160 can be acquired from the acquisition units 130, 132, and so on. Here, the acquisition units 130, 132, and so on can be pre-deployed at designated locations within the physical space 160, e.g., at the ceiling of the living room, and so on. In this way, richer data can be acquired from multiple angles.

[0055] As shown in FIG. 3, the key information 311 of the media item includes at least any of the following: a person 311-1 in the media item, a caption 311-2, a scene 311-3, a sound, and a provider 311-4 of the media item. As an example, the robotic device 110 can identify the key information 311 from the image 310 based on image recognition techniques, and then search the key information 311 through the network to determine the relevant data of the media item. For example, the key information 311 includes a football player and a sports channel X, the robotic device 110 can search the network based on the key information 311 to determine detailed event information.

[0056] Alternatively and / or additionally, the processing capability of the model can be invoked to identify the key information 311 from the image 310. For example, the person 311-1 in the media item can refer to identifying actors, athletes, or people in the image in a video. For a video or movie that provides captions 311-2, the robotic device 110 can read and analyze the caption text to understand the content or dialogue of the media item. The scene 311-3 refers to the background or environment in the media item, such as a football field and a match, a cityscape, an indoor setting, or other specific types of events. The sound refers to identifying specific sounds in the media item, such as background music, sound effects, or voices. The provider 311-4 of the media item refers to the creator, publisher, and so on of the media content, which helps to confirm the source of the media and related information.

[0057] With some implementations of the present disclosure, the relevant data can be determined based on the key information 311 in the media item. In this way, the accuracy of the relevant data provided to the user 120 can be ensured. Meanwhile, the robotic device 110 can also recommend similar or related media content to the user 120 based on the identified key information 311. In this way, the robotic device 110 can extract deep information from the media item in order to better understand and utilize these data, thereby providing richer and more personalized services to the user 120.

[0058] According to some implementations of the present disclosure, the relevant data of the media item can also be determined by a machine learning model and the key information. More details of the machine learning model determining the relevant data of the media item are described with reference to FIG. 4, which shows a block diagram 400 of a process of determining the relevant data of the media item according to some implementations of the present disclosure. For example, a prompt 410 can be constructed and input to a machine learning model (e.g., a language model 420) to invoke the processing capability of the model to determine the relevant data of the media item. The prompt 410 can be obtained according to the image 310, the key information 311 and the query request 220, for example. The prompt 410 can be expressed as “please identify the program content of the media item from the following image and key information”, for example, and the collected image and the prompt are submitted to the model. The model can process the image and identify the intention of the prompt, and the model can output a first response to the image and the prompt, and then determine the relevant data of the media item according to the first response.

[0059] With some implementations of the present disclosure, the relevant data of the media item can be determined in various ways, thereby improving the accuracy and comprehensiveness of the robot device in determining the relevant data.

[0060] According to some implementations of the present disclosure, the robot device 110 can further provide information about the media item or the relevant data when assisting the user 120 in operating the electronic device 170. For example, the robot device 110 can provide a response to the acquisition request of the user 120 in response to receiving the acquisition request of the user for the relevant data or the media item. For example, the user 120 can ask the robot device 110 about the score of a sports match and the season data of a certain player, etc.

[0061] Alternatively and / or additionally, a prompt can be constructed based on the acquisition request and input to a machine learning model to invoke the processing capability of the model to determine the intention of the user 120, thereby providing detailed response content to the user 120. For example, the user 120 asks in natural language “how well did the XXX player perform in this game”, and the constructed prompt can be “the average score of the XXX player in the season, the score in this game and whether the performance meets the average level”, etc. Alternatively and / or additionally, the robot device 110 can access the relevant webpage of the game and present the webpage at the interaction unit 112. In this way, the user 120 can communicate with other users in the webpage during the process of watching the game.

[0062] According to some implementations of the present disclosure, the robotic device 110 can further provide additional information of the media item according to the level of attention of the user 120 to the media item. For example, the state data of the user 120 can be acquired from the acquisition unit 114 of the robotic device 110, and the level of attention of the user 120 to the media item can be determined according to the state data of the user 120. The robotic device 110 can provide additional data about the media item to the user 120 in response to determining that the level of attention of the user 120 is higher than a predetermined threshold. For example, the robotic device 110 can determine and evaluate the state data of the user through the acquisition unit 114. The state data can be, for example, the length of time of watching or listening, for example, the longer the user 120 stays when watching a video or listening to audio, the higher the interest can be. The state data can also be the frequency and type of interaction of the user 120 with the device, such as asking questions or facial reactions, which can reflect the level of engagement.

[0063] Alternatively and / or additionally, the state data can be physiological indicators of the user 120, for example, the heart rate, pupil dilation, and other physiological responses of the user 120 determined by the acquisition unit to more accurately determine the attention and emotional state of the user 120.

[0064] When the robotic device 110 determines that the level of attention of the user 120 exceeds the predetermined threshold, the robotic device 110 will automatically or at the request of the user 120 provide more additional data related to the media item. The additional data can include in-depth content beyond the surface information of the media item, which can enrich the information sources and understanding of the user 120. For example, when watching a sports event or an interview of an athlete, the additional data can include background information of the athlete to help the user 120 better evaluate the career stage and future potential of the athlete. The additional data can also include personal growth experiences, share the growth story and training process of the athlete, which can increase emotional resonance and enable the audience to better understand the background and efforts of the athlete.

[0065] With some implementations of the present disclosure, the robotic device 110 not only provides basic media playback services, but also dynamically adjusts the content according to the specific needs and interests of the user 120, thereby providing more personalized and immersive services.

[0066] According to some implementations of the present disclosure, in response to determining that the attention level of the user 120 is lower than a predetermined threshold, a message can be provided to the user 120, which asks the user 120 whether the user 120 wishes to switch the media item. Next, in response to receiving a switching request from the user 120 for switching the media item, a data channel that matches the switching request is determined; and the robotic device is instructed to switch the data source of the electronic device to the data channel. For example, the robotic device 110 continuously determines the attention level of the user 120 to the currently played media item by collecting units. For example, the eye movement track, facial expression, body posture, interaction frequency with the electronic device 170, and even physiological characteristics (such as heart rate variation) of the user 120 are collected.

[0067] Next, when it is evaluated that the interest of the user 120 to the current content is reduced, i.e., the attention level is lower than a preset threshold, the robotic device 110 actively asks the user 120 whether the user 120 is willing to change the media content being played. The message can be displayed to the user 120 in the form of a pop-up window, a voice prompt, or a prompt information on the screen, etc., to provide the user 120 with an explicit opportunity to choose. If the user 120 is not interested in the current content, the user 120 can choose to confirm the switching request, for example, by clicking a confirmation button, issuing a voice instruction, or performing other pre-set actions.

[0068] Further, when the user 120 expresses the intention to switch the media item, the robotic device 110 needs to determine a new data source, for example, a data channel that matches the switching request can be determined according to the use history of the user 120, a recommendation algorithm, or an instant selection specified by the user 120.

[0069] Finally, the robotic device 110 performs the actual switching operation to change the data source of the television to the newly determined data channel. In this way, the playing of the current media item can be stopped, and the new media item selected by the user 120 can be started to be loaded and played.

[0070] With some implementations of the present disclosure, the robotic device 110 can flexibly respond to the changes in the needs of the user 120, thereby providing more personalized and timely services.

[0071] Alternatively and / or additionally, the robotic device 110 can also provide a control device of the electronic device 170 to the user 120, so that the user 120 changes the data source of the electronic device 170 to the newly determined data channel. Further, the user 120 can be facilitated to switch the data source at any time by using the control device.

[0072] According to some implementations of the present disclosure, the robotic device 110 can also determine a plurality of types of a plurality of media items provided by a plurality of data sources of the electronic device 170, respectively. Next, a data source matching the switching request is selected from the plurality of data sources. For example, the robotic device 110 can identify and classify various media content played by different data sources connected to the electronic device 170 (e.g., a television). The data sources can be satellite television service, cable television network, internet streaming service, local media library, etc., which can provide different types of media items, such as music, sports events, movies, news, educational programs, etc., respectively. The robotic device 110 classifies these media items into a plurality of types by analyzing metadata, channel lists, content tags, etc., for subsequent filtering and selection. When the user 120 makes a request to switch media items, the robotic device 110 can select a data source matching the request from the plurality of data sources according to the user’s interest, current context, or explicit user instruction.

[0073] Alternatively and / or additionally, the processing capability of a machine learning model can be utilized to select a data source matching the switching request. For example, a cue of content the user wants to watch or listen to is constructed and input to the model. The model can output a response to the cue based on the cue, and then the response can be used to find a data source that can provide media items of the corresponding type. Next, after finding the matching data source, the robotic device 110 switches the data source of the electronic device 170 to the data source and starts playing the media content the user desires.

[0074] For example, when the user wants to switch from a current news program to a live sports event, the user 120 only needs to express this intention to the robotic device 110, and the robotic device 110 can quickly locate a television station providing sports content from a plurality of data sources and complete the switching to seamlessly access the media content of interest to the user.

[0075] With some implementations of the present disclosure, the robotic device 110 can intelligently perform media content management and personalized services, enhancing the user’s ability to obtain information, so that the user can quickly find the content he or she likes in a rich media resource.

[0076] According to some implementations of the present disclosure, when instructing the robotic device 110 to switch the data source of the electronic device 170 to the data channel, the robotic device 110 can be instructed to obtain a control device for switching the data source of the electronic device, and then instruct the robotic device 110 to provide the control device to the user 120 for inputting a switching instruction for switching the data source by the user 120. For example, when the user 120 wants to directly control the switching of the data source, the robotic device 110 can be instructed to obtain a control device, which can be a remote controller, a control panel, or other types of user interface devices, for example, for the user to manually input the switching instruction.

[0077] After the robotic device 110 obtains the control device, the control device can be provided to the user 120. For example, the robotic device 110 delivers a remote controller to the user 120, or the robotic device 110 displays a virtual control panel through its own display screen, or connects to the smart terminal (e.g., a mobile phone or a tablet computer) of the user 120 wirelessly, so that the user 120 can operate on his / her own terminal device. After receiving the control device, the user 120 can directly input the switching instruction on the control device according to his / her own needs and preferences. For example, selecting a specific channel, browsing a program list, searching for a specific type of content, etc.

[0078] With some implementations of the present disclosure, the robotic device 110 can not only automatically perform tasks, but also provide more personalized and customized services according to the needs and interests of the user 120. By introducing the control device, the robotic device 110 enhances the interaction between the user 120 and the electronic device 170, making the entire switching process more intuitive and friendly.

[0079] According to some implementations of the present disclosure, when instructing the robotic device 110 to obtain the control device, a second image of the physical space 160 can be obtained, and then the position of the control device is located based on the second image, and the robotic device 110 is instructed to obtain the control device from the determined position. More details of obtaining the control device are described with reference to FIG. 5, which shows a block diagram 500 of a process of obtaining a control device according to some implementations of the present disclosure. As shown in FIG. 5, for example, the second image (e.g., one or more images 510) of the physical space 160 can be obtained from the acquisition unit 114 of the robotic device 110. Since the robotic device 110 can move freely in the physical space 160, the acquisition unit 114 can acquire the second image 510 of each position in the physical space 160, thereby facilitating the search for the control device 530.

[0080] Alternatively and / or additionally, a second image of the physical space 160 can be acquired from the acquisition units 130, …, and 132. Here, the acquisition units 130, …, and 132 can be pre-deployed at designated locations within the physical space 160, e.g., at the ceiling of a living room, etc. In this way, the second image of the physical space 160 taken from a top-down perspective can be obtained, which facilitates understanding the layout of the physical space 160 as a whole and thus locating the control device.

[0081] According to some implementations of the present disclosure, the control device 530 in the second image can be determined based on various manners. For example, the control device 530 can be recognized from the second image based on image recognition techniques. Alternatively and / or additionally, a prompt can be constructed and input to the model so as to invoke the processing capability of the model to recognize the control device 530 from the second image. The prompt can be expressed as, for example, “please identify the ‘control device’ from the following image”, and the acquired image and the prompt are submitted to the model.

[0082] The model can process the image, and in the case that the second image includes the control device, the model can output the location where the control device is located (e.g., the region coordinates of the object in the image, and / or directly output the image of the region where the control device is located, etc.). If the image does not include the control device, the model can output an answer such as “not found”. With some implementations of the present disclosure, whether the image includes the control device can be detected based on various manners, thereby improving the performance of the robotic device in acquiring the control device.

[0083] According to some implementations of the present disclosure, in response to determining that the second image does not include the control device in the physical space 160 (i.e., the first physical space), a second physical space associated with the first physical space is determined. It should be understood that the second physical space here is a potential physical space that can include the first object. For example, in a home environment, since the control device can be placed in the drawer of a tea table or the drawer of a TV cabinet, the second physical space can be determined as the tea table and / or the TV cabinet. Specifically, in the process of determining the second physical space associated with the first physical space, a prompt for locating the control device can be acquired based on the second image and the control device; and a response of the machine learning model to the prompt is received so as to determine the second physical space.

[0084] According to some implementations of the present disclosure, the language model 420 is a trained and fine-tuned model, and has rich knowledge of performing tasks in multiple domains. The language model 420 can determine that the image 510 includes the TV cabinet 520, and identify the TV cabinet 520 as the second physical space. FIG. 5 is merely illustrative, and the language model 420 can process one or more images from different capturing devices, and find one or more second physical spaces that can include the control device 530. For example, assume that another image includes a tea table, and the tea table can be determined as the second physical space.

[0085] According to some implementations of the present disclosure, a fine-tuning operation can be performed on the model, for example, the robot device can be instructed to pre-capture images of various parts in the first physical space, for example, the robot device 110 can open the TV cabinet (or the tea table, a drawer, etc.), and capture various object-related images in the TV cabinet. Then, the model can be fine-tuned with the captured images. In this way, the model can master the specific storage locations of various objects, and thus improve the accuracy of determining the second physical space.

[0086] With some implementations of the present disclosure, the powerful processing capability and rich knowledge of the model can be utilized to determine the second physical space that can include the control device. Then, the robot device can be instructed to go to the second physical space to continue searching for the control device. Compared with the prior art solution that can only search for the control device in the visible physical space through image recognition, the technical solution of the present disclosure can search for the target object in more potential physical spaces, and thus find the hidden object that is not directly exposed to the coverage of the capturing device. In this way, the efficiency and accuracy of locating the control device can be improved, thereby improving the overall efficiency of performing the user task.

[0087] According to some implementations of the present disclosure, the second physical space can be determined based on image recognition. Specifically, in the process of determining the second physical space associated with the first physical space, a plurality of second objects can be recognized from the second image, a second object can be selected from the plurality of second objects based on a knowledge base, and a space where the second object is located can be taken as the second physical space. Here, the knowledge base can be predefined, and include the association relationship between the physical space and the object. For example, the knowledge base can include: (control device, TV cabinet), (control device, tea table), etc. The robot device can be instructed to pre-capture images of various parts in the first physical space, for example, the robot device can open the TV cabinet (or the tea table, a drawer, etc.), and capture various object-related images in the TV cabinet, and thus determine the content of the knowledge base.

[0088] With some implementations of the present disclosure, image recognition techniques can be utilized to determine various objects in a physical space, and in turn, determine a potential physical space that can include a target object. In this way, the efficiency and accuracy of locating a target object can be improved, thereby improving the overall efficiency of performing a user task.

[0089] In some cases, the robotic device can directly enter the second physical space and retrieve the target object. Assuming the second physical space is a "bedroom", the robotic device can directly enter the bedroom and take the control device 530 from the bedroom. In some cases, the robotic device can not be able to directly enter the second physical space. Assuming the second physical space is a "TV cabinet", the robotic device needs to determine the specific way to open the drawer of the TV cabinet and access the interior space of the TV cabinet. According to some implementations of the present disclosure, the robotic device can determine the access way for accessing the second physical space from the second image. In turn, the robotic device can be instructed to access the second physical space according to the access way.

[0090] Specifically, the handle of the TV cabinet drawer can be recognized from the second image, at which point it can be determined that the handle needs to be pulled to open the drawer. Alternatively and / or additionally, a corresponding prompt can be constructed and the model can be asked how to open the drawer of the TV cabinet. The prompt can be expressed as, for example, "determine the way to open the drawer of the TV cabinet from the following image", the prompt and the corresponding image can be sent to the model. At this point, the model can return that the handle needs to be pulled. In turn, the robotic device can be instructed to pull the handle to open the drawer of the TV cabinet and find the control device. With some implementations of the present disclosure, the powerful processing capability of the model can be called upon to solve unknown problems in complex environments, and in turn, determine the actions that need to be performed by the robotic device. In this way, the ability of the robotic device to handle complex tasks can be improved, thereby performing user tasks in a more accurate manner.

[0091] According to some implementations of the present disclosure, an action model can be utilized to determine the specific actions performed by the robotic device. More details are described with reference to FIG. 6, which shows a block diagram 600 of a process of calling an action model according to some implementations of the present disclosure. As shown in FIG. 6, an action model 630 can be provided, which can determine the specific actions to be performed by the robotic device based on the current state of the robotic device and the instructions, and which can be a pre-trained and fine-tuned model.

[0092] It should be appreciated that the current state can include data of multiple aspects, such as an image of the robotic device, an image of the environment of the robotic device, pose data of the robotic arm (e.g., positions of various joints of the robotic arm (POS1, …)), and a state of a tool (e.g., a gripper, a cutter, etc.) fixed at the end of the robotic arm. For example, 0 can be used to represent a closed state of the gripper, and 1 can be used to represent an open state of the gripper. The instruction and the current state can be input to the action model 630, and then the action model is used to determine an action to be performed by the robotic device based on the instruction and the current state. Here, the action can represent a difference between a current pose of the robotic device and a next pose, and a difference between a current state of the tool and a next state, etc.

[0093] The instruction 610 (e.g., “open the drawer of the TV cabinet”) can be input to the action model 630, where the instruction 610 can be expressed in natural language and can be determined from a response of the language model. Further, a current state of the robotic device can be obtained, and the action model 630 can determine a corresponding action 640 based on the input data. For example, orientations, positions, velocities, accelerations, etc. of various joints in the arm, and / or wheels and / or other movable devices of the robotic device at a next point in time can be determined. Further, the determined action 640 can be used to control a state of the robotic device at the next point in time.

[0094] With some implementations of the present disclosure, a correlation can be established between the language model and the action model, and a user task initially input by a user in natural language can be converted into a specific action executable by the robotic device. In this way, actions of the robotic device can be precisely controlled, and the user task can be performed with higher efficiency.

[0095] According to some implementations of the present disclosure, if the control device is blocked by other items, the items can be removed first, and then the control device is retrieved. Specifically, in response to determining that the second image represents that the target object (e.g., the control device) is blocked by other objects in the second physical space, the other objects can be moved to obtain the target object. More details are described with reference to FIG. 7, which shows a block diagram 700 of a process of moving objects according to some implementations of the present disclosure. As shown in FIG. 7, in the image 710, the object 720 is a control device to be retrieved, and the object 730 is located in front of the object 720 and blocks the object 720. At this time, the robotic device 110 can be instructed to move the object 730 from the current position to a position that does not hinder the retrieval of the object 720.

[0096] According to some implementations of the present disclosure, the target position can be determined, and the robotic device can be instructed to move the object 730 to the target position. At this time, the motion model will generate actions to control the robotic device to move the object 730 from the current position to the target position. In this way, the robotic device can be supported to handle complex problems in a complex environment, so as to perform the user task in a more accurate manner.

[0097] According to some implementations of the present disclosure, the exit manner for exiting the second physical space can be determined, and the robotic device can be instructed to exit the second physical space according to the exit manner. Specifically, after the robotic device has retrieved the object 720, a new image can be captured and a new prompt word can be constructed to query the language model for the next instruction. The prompt word may, for example, be expressed as: “please determine the next instruction based on the following image”, “what to do next”, and the like. The language model can return “close the TV cabinet drawer”, at which time corresponding actions can be generated based on the instruction “close the TV cabinet drawer” and the current state of the robotic device to instruct the robotic device to close the refrigerator.

[0098] According to some implementations of the present disclosure, after the robotic device closes the TV cabinet drawer, the robotic device 110 can be instructed to go to the location of the user 120. Specifically, an image can be captured in real time and the location of the user 120 can be located in the image. Further, a corresponding instruction can be determined based on the current location (e.g., location A) of the robotic device 110 and the location (e.g., location B) of the user, and the corresponding instruction can be determined. At this time, the instruction can be expressed as: move from location A to location B. At this time, the action model 630 will generate corresponding actions that can control the robotic device 110 to move from location A to location B according to the determined trajectory. In this way, the robotic device 110 can complete the task of “acquiring the control device”.

[0099] It should be understood that although the above describes one example implementation of the present disclosure in a Chinese language environment. Alternatively and / or additionally, the technical solutions of one example implementation of the present disclosure can be performed in a variety of language environments. For example, the robot can be controlled in a Chinese, English, Japanese, French, and the like environment. Specifically, the robot can be controlled in different language application environments based on the multi-language capabilities provided by the machine learning technology. Further, although the above describes the process of performing the user task using the robotic device by taking the operation of the TV as an example, alternatively and / or additionally, the robotic device can be controlled to perform other user tasks, such as finding other objects in the room, placing a certain object to a specified location, and the like.

[0100] According to some implementations of the present disclosure, a user can interact with the robotic device via language, motion, gesture, etc. For example, the user can speak a user task that he / she desires to be performed, predefine a certain motion to specify the user task, etc. Specifically, the user can make a motion pointing to a TV set, which can be used as a user task to trigger the robotic device to turn on the TV set. When the motion is recognized from the captured image sequence, the robotic device can automatically ask the user whether he / she needs to turn on the TV set, and in case of a positive reply, the robotic device can turn on the TV set and provide the user with relevant data of media items.

[0101] Alternatively and / or additionally, the user can interact with the robotic device via the interaction unit 112, e.g., the user inputs a task in text and / or image representation, and controls the robotic device to perform the task. Alternatively and / or additionally, the user can specify a condition for performing the task, e.g., to perform the task immediately, to perform the task after a predetermined time, or to perform the task upon determining that a predetermined condition is met (e.g., after the user sits in the living room), etc.

[0102] According to some implementations of the present disclosure, a variety of positioning algorithms can be utilized to determine the position of the robotic device, and of various objects in the physical environment. For example, a global positioning system (GPS) can be deployed at the robotic device, and satellite signals can be used to determine the precise position of the robotic device. Alternatively and / or additionally, a communication unit can be deployed at the robotic device, by means of which signals between the communication unit and a base station, and utilizing a communication network, the position of the robotic device can be determined. Alternatively and / or additionally, Wi-Fi access points can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspots in order to determine the position via Wi-Fi signal strength and known positions of the Wi-Fi access points. Alternatively and / or additionally, the communication unit at the robotic device can support Bluetooth functionality, in which case Bluetooth signals and known positions of Bluetooth devices can be used to determine the position of nearby devices. An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, and in turn determine the position of the robotic device.

[0103] Alternatively and / or additionally, a visual positioning system can be used to determine the position of the robotic device and / or the individual objects. A map of the physical space can be pre-acquired, and the positions of the individual objects can be marked in the map. The robotic device can detect the distance from the surrounding objects with the echo detection unit, and determine the specific positions of the individual objects in combination with the acquired images and the map of the physical space. Specifically, computer aided design (CAD) and geographic information system (GIS) can be used, and a positioning algorithm can be used to determine the positions. Alternatively and / or additionally, a tracking unit can be deployed at important objects in the physical space, for example, a tracking unit can be added at the remote controller of a household appliance (e.g., a television remote controller, an air conditioner remote controller), so that the robotic device can timely acquire the accurate positions of the important objects, and the like.

[0104] According to some implementations of the present disclosure, the original position of the robotic device itself and the destination position to which the robotic device is expected to go can be determined based on the methods described above. The robotic device can determine a path from the original position to the destination position. For example, the surrounding environment images can be constantly acquired, and the path can be constantly updated while ensuring to avoid obstacles, and the robotic device can be caused to move along the path to the destination position.

[0105] According to some implementations of the present disclosure, after reaching the destination position, the robotic device can perform a specified task. For example, a specified object can be acquired and moved to a corresponding position. The constraint conditions, i.e., the constraint conditions that should be followed during the execution of the task, can be determined with the language model and / or the knowledge base. For example, the image and the corresponding prompt word can be acquired, and the image and the prompt word can be input to the language model, and then the constraint conditions can be received from the language model. For example, the prompt word can be determined as: “please determine the constraint conditions that should be followed during the movement of the XXX object based on the following image”, or “please determine the precautions during the movement of the XXX object”, and the like.

[0106] At this time, it can be determined that the original posture of the object (e.g., a bottled water, a plate, a bowl, and the like) should be maintained (e.g., the vertical direction should be maintained, and the object should not be tilted) during the movement of the object. Further, the constraint conditions can be input to the action model, and a series of actions output by the action model will perform the corresponding task while ensuring the constraint conditions. With some implementations of the present disclosure, the safety during the operation of the robotic device can be ensured, so that the accidental damage to an object can be avoided, and the like.

[0107] Using the exemplary implementations of this disclosure, robotic devices can perform user tasks in complex physical spaces. In this way, the robotic device can activate electronic devices within the physical space based on a user's activation request, acquire media items provided by the electronic devices after activation, and then provide the user with relevant data for the media items upon receiving a query request. This approach improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby completing the intended user task.

[0108] Example process

[0109] Figure 8 illustrates a flowchart of a method 800 for performing a user task according to some implementations of this disclosure. At block 810, a media item provided by an electronic device is acquired. At block 820, in response to receiving a query request from a user, key information is identified from the media item, the key information including at least one of the following: characters, subtitles, scenes, sounds in the media item, and the provider of the media item. At block 830, relevant data of the media item is determined based on the key information; and at block 840, a robotic device provides the relevant data of the media item to the user.

[0110] According to some implementations of this disclosure, obtaining the media item provided by the electronic device includes: in response to receiving a start request from a user, the robot device starts the electronic device in the physical space, wherein the user and the robot device are located in the physical space; and obtaining the media item provided by the electronic device.

[0111] According to some implementations of this disclosure, obtaining a media item includes: obtaining an image of the media item, and determining relevant data based on key information includes: obtaining prompt words based on the image, key information, and query request, wherein the prompt words are used to determine relevant data; and determining relevant data based on a first response from a machine learning model to the prompt words.

[0112] According to some implementations of this disclosure, the method 800 further includes: in response to receiving a user's request to acquire relevant data or media items, providing a response to the acquisition request to the user.

[0113] According to some implementations of this disclosure, the method 800 further includes: determining the user's level of attention to a media item based on the user's state data; and providing additional data about the media item in response to determining that the level of attention is higher than a predetermined threshold.

[0114] According to some implementations of this disclosure, the method 800 further includes: in response to determining that the level of attention is below a predetermined threshold, providing a message to the user, the message asking the user whether they wish to switch media items; in response to receiving a switching request from the user for switching media items, determining a data channel matching the switching request; and the robotic device switching the data source of the electronic device to the data channel.

[0115] According to some implementations of this disclosure, determining the data source includes: determining multiple types of multiple media items provided by multiple data sources of an electronic device; and selecting a data source that matches the switching request from the multiple data sources.

[0116] According to some implementations of this disclosure, the robot device switches the data source of an electronic device to a data channel by: the robot device acquiring a control device for switching the data source of the electronic device; and providing the control device to a user so that the user can input a switching command for switching the data source.

[0117] According to some implementations of this disclosure, the robot device acquiring the control device includes: acquiring an image of the physical space; locating the position of the control device based on the image; and the robot device acquiring the control device from the position.

[0118] According to some implementations of this disclosure, the relevant data includes at least one of the following: audio data, video data, image data, and text data.

[0119] Example devices and equipment

[0120] Figure 9 shows a block diagram of an apparatus 900 for performing user tasks according to some implementations of the present disclosure. The apparatus 900 includes: an acquisition module 910 configured to acquire media items provided by an electronic device; an identification module 920 configured to identify key information from the media items in response to receiving a query request from a user, the key information including at least one of the following: characters, subtitles, scenes, sounds in the media items, and the provider of the media items; a determination module 930 configured to determine relevant data of the media items based on the key information; and a providing module 940 configured to cause a robotic device to provide the relevant data of the media items to the user.

[0121] According to some implementations of this disclosure, the acquisition module is further configured to: in response to receiving a start request from a user, the robot device starts an electronic device in a physical space, wherein the user and the robot device are located in the physical space; and acquire the media item provided by the electronic device.

[0122] According to some implementations of this disclosure, the acquisition module 910 is further configured to: acquire an image of a media item, and the determination module 930 is further configured to: acquire prompt words based on the image, key information, and query request, the prompt words being used to determine relevant data; and determine relevant data based on a first response to the prompt words from a machine learning model.

[0123] According to some implementations of this disclosure, the determining module 930 is further configured to: in response to receiving a user's request to acquire relevant data or media items, provide a response to the acquisition request to the user.

[0124] According to some implementations of this disclosure, the determining module 930 is further configured to: determine the user's level of attention to the media item based on the user's state data; and, in response to determining that the level of attention is higher than a predetermined threshold, provide additional data about the media item.

[0125] According to some implementations of this disclosure, the determining module 930 is further configured to: in response to determining that the level of attention is below a predetermined threshold, provide a message to the user, the message asking the user whether they wish to switch media items; in response to receiving a switching request from the user for switching media items, determine a data channel matching the switching request; and cause the robot device to switch the data source of the electronic device to the data channel.

[0126] According to some implementations of this disclosure, the determining module 930 is further configured to: determine multiple types of multiple media items provided by multiple data sources of the electronic device; and select a data source that matches the switching request from the multiple data sources.

[0127] According to some implementations of this disclosure, the determining module 930 is further configured to: enable the robot device to acquire a control device for switching the electronic device's data source; and provide the control device to the user so that the user can input a switching command for switching the data source.

[0128] According to some implementations of this disclosure, the determining module 930 is further configured to: acquire an image of the physical space; locate the position of the control device based on the image; and cause the robot device to acquire the control device from the position.

[0129] According to some implementations of this disclosure, the relevant data includes at least one of the following: audio data, video data, image data, and text data.

[0130] Figure 10 shows a block diagram of a device 1000 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1000 shown in Figure 10 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 1000 shown in Figure 10 can be used to implement the methods described above.

[0131] As shown in Figure 10, the computing device 1000 is in the form of a general-purpose computing device. Components of the computing device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage devices 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. The processing unit 1010 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1020. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 1000.

[0132] Computing device 1000 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1030 can be removable or non-removable media and may include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within computing device 1000.

[0133] The computing device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 10, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 1020 may include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.

[0134] The communication unit 1040 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 1000 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0135] Input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1060 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 1000 can also communicate with one or more external devices (not shown) via communication unit 1040 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 1000, or with any device (e.g., network card, modem, etc.) that enables computing device 1000 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interface (not shown).

[0136] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.

[0137] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0138] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0139] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0141] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for performing a user task, comprising: obtaining a media item provided by an electronic device; in response to receiving a query request from a user, identifying key information from the media item, the key information comprising at least any of a character, a caption, a scene, a sound, and a provider of the media item in the media item; determining relevant data of the media item based on the key information; and providing the relevant data of the media item to the user by a robotic device. 2.The method of claim 1, wherein obtaining the media item provided by the electronic device comprises: in response to receiving a start request from a user, the user being located in a physical space with the robotic device, starting the electronic device in the physical space by the robotic device; and obtaining the media item provided by the electronic device. obtaining an image of the media item, and determining the relevant data based on the key information comprises: obtaining a cue word based on the image, the key information, and the query request, the cue word being used to determine the relevant data; and 3. The method of claim 1, wherein obtaining the media item comprises: determining the relevant data based on a first response to the cue word from a machine learning model. in response to receiving a request for the relevant data or the media item from the user, providing a response to the request to the user. 5.The method of claim 4, further comprising:

4. The method of claim 1, further comprising: determining a degree of attention of the user to the media item based on state data of the user; and in response to determining that the degree of attention is higher than a predetermined threshold, providing additional data about the media item. 6.The method of claim 5, further comprising: in response to determining that the degree of attention is lower than the predetermined threshold, providing a message to the user, the message asking whether the user wants to switch the media item; in response to receiving a switch request for switching the media item from the user, determining a data channel matching the switch request; and switching, by the robotic device, a data source of the electronic device to the data channel. 7.The method of claim 6, wherein determining the data source comprises: determining a plurality of types of a plurality of media items provided by a plurality of data sources of the electronic device, respectively; and selecting a data source matching the switch request from the plurality of data sources. 8.The method of claim 6, wherein switching, by the robotic device, the data source of the electronic device to the data channel comprises: obtaining, by the robotic device, a control device for switching the data source of the electronic device; and providing the control device to the user for inputting a switch instruction for switching the data source by the user. 9.The method of claim 6, wherein obtaining, by the robotic device, the control device comprises: obtaining an image of the physical space; locating a position of the control device based on the image; and obtaining, by the robotic device, the control device from the position. ​ ​ ​ ​ ​ ​ ​ 10.The method of claim 1, wherein the related data comprises at least any one of audio data, video data, image data, and text data. 11.An apparatus for performing a user task, comprising: an obtaining module configured to obtain a media item provided by an electronic device; an identifying module configured to identify, in response to receiving a query request from a user, key information from the media item, the key information comprising at least any one of a person, a caption, a scene, a sound in the media item, and a provider of the media item; a determining module configured to determine, based on the key information, related data of the media item; and a providing module configured to cause a robotic device to provide the related data of the media item to the user. 12.An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1-10. 13.A computer-readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to implement the method according to any one of claims 1-10. 14.A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-10. ​ ​

Citation Information

Patent Citations

  • Program recommendation method and device

    CN105959806A

  • Interaction method and device based on artificial intelligence in video playing process

    CN106937172A

  • Intelligent television set-top box based on voice recognition

    CN111741369A

  • Program recommendation method, television and storage medium

    CN113286199A

  • Spring Structure For Mouse Cradle

    KR102489764B1