Task marking method, voice interaction method and electronic device

By displaying and marking the task characteristics of voice or text input on electronic devices, the problem of users being unable to correct parsing deviations in a timely manner is solved, thus improving the user experience of human-computer interaction.

WO2025261440A1PCT designated stage Publication Date: 2025-12-26HUAWEI TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102066
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-20
Filing Date
2025-06-19
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

During human-computer interaction, users cannot correct the interpretation deviations of electronic devices for voice or text input in a timely manner, resulting in a poor user experience, especially when inputting long voice or text messages.

Method used

Electronic devices display a text stream and highlight task characteristics in a prominent way, such as replacing person names with avatars, application names with application icons, and marking task content in real time. This allows users to quickly identify whether the electronic device's parsing is correct and make adjustments as necessary.

Benefits of technology

By marking task characteristics in real time, users can correct deviations promptly, improving the user experience of human-computer interaction and reducing trial-and-error costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102066_26122025_PF_FP_ABST
    Figure CN2025102066_26122025_PF_FP_ABST
Patent Text Reader

Abstract

A task marking method, a voice interaction method and an electronic device, which are used to mark in real time tasks in text streams that need to be executed, so as to prompt a user to pay attention to whether the electronic device has correctly parsed the tasks. For example, the electronic device acquires a first text stream. The electronic device displays the first text stream, wherein the first text stream comprises a first task; the first task comprises N task characteristics, N being a positive integer; the N task characteristics comprise at least one of task type, task content, task object, required application, and required device; the N task characteristics in the first text stream are displayed in different display manners; and the display manner for the N task characteristics is different from the display manner for content in the first text stream apart from the N task characteristics. The electronic device executes the first task.
Need to check novelty before this filing date? Find Prior Art

Description

A task labeling method, a voice interaction method, and an electronic device

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202410807659.1, filed on June 20, 2024, entitled "A Task Marking Method and Electronic Device", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of terminal technology, and in particular to a task marking method, a voice interaction method, and an electronic device. Background Technology

[0004] In human-computer interaction, electronic devices need to parse user input, such as voice and text, to determine the user's intent before responding. Currently, when users input voice or text, they are unsure whether the electronic device will correctly parse it; they can only wait for the device to respond and then judge the result. If the response does not meet the user's expectations, they need to input voice or text again for the device to respond once more. Therefore, this human-computer interaction method is not conducive to users correcting deviations in a timely manner, resulting in a poor user experience, especially with long voice or text inputs, where the impact on the user experience is more pronounced. Summary of the Invention

[0005] This application provides a task marking method, a voice interaction method, and an electronic device. The electronic device can display a text stream, and the tasks that the electronic device needs to perform in the text stream can be displayed in a prominent manner so that the user can quickly focus on the task and thus quickly know whether the electronic device's parsing of the task is correct. This helps the user to correct deviations in a timely manner and improves the user experience.

[0006] Firstly, a task tagging method is provided, applied to an electronic device, such as a mobile phone. The method includes: acquiring a first text stream; displaying the first text stream, the first text stream including a first task, the first task including N task characteristics, where N is a positive integer, the N task characteristics including at least one of task type, task content, task object, required application, and required device; displaying the N task characteristics in the first text stream using different display methods, and the display methods of the N task characteristics being different from the display methods of other content in the first text stream besides the N task characteristics; and executing the first task.

[0007] In this embodiment, during the user's input of voice or text, the electronic device can display a text stream and mark tasks within the text stream. For example, different task characteristics can correspond to different marking methods (i.e., display methods). In this way, the user can quickly focus on the various task characteristics and thus quickly know whether the electronic device's parsing of the task is correct, helping the user to correct deviations in a timely manner and improve the user experience.

[0008] In one possible design, when the first text stream contains a person's name, the person's name is replaced with the corresponding avatar; when the first text stream contains an application name, the application name is replaced with the corresponding application icon; when the first text stream contains a device type, the device type is replaced with the corresponding device icon.

[0009] In this embodiment, when an electronic device marks tasks in a text stream, it can replace names with avatars, application names with application icons, and device types with device icons. Compared to text, icons are more eye-catching, allowing users to quickly determine whether the electronic device is correctly parsing the task, helping them correct deviations in a timely manner and improving the user experience.

[0010] In one possible design, the first text stream further includes a second task, which includes M task features, where M is a positive integer. The display methods corresponding to the M task features are different from those corresponding to the N task features. The method further includes executing the second task.

[0011] In this embodiment, the electronic device can mark different tasks in a text stream, and the marking methods for different tasks can be different. For example, the task characteristics of the first task and the task characteristics of the second task can correspond to different display methods. In this way, users can quickly focus on different tasks, thereby judging whether the electronic device's parsing of each task is correct and improving the user experience.

[0012] In one possible design, the first text stream further includes a third task for canceling the first task, and the method further includes: canceling the execution of the first task if it has not been completed; and adding a first marker to the first text stream to indicate that the first task has been canceled.

[0013] In this embodiment, the electronic device can cancel one task from one task in a text stream and add a marker to the text stream to indicate that the other task has been canceled. This interaction method provides a better user experience.

[0014] In one possible design, the first text stream further includes a third task, which is used to cancel the first task. The method further includes: if the first task has been completed, outputting a first prompt message or executing a fourth task, wherein the first prompt message is used to indicate that the first task has been completed, and the fourth task is used to adjust the executed first task.

[0015] In this embodiment, the electronic device can cancel the first task based on the third task in the text stream. If the first task has been completed, it can prompt the user or execute a fourth task to adjust the executed first task. This interaction method eliminates the need for the user to manually trigger the fourth task, improving convenience and providing a better user experience.

[0016] In one possible design, the method further includes: when it is determined that the first task cannot be adjusted, outputting a second prompt message, the second prompt message being used to indicate that the first task cannot be adjusted.

[0017] In this embodiment, if the first task has been completed, the fourth task is executed to adjust the first task. If adjustment is not possible, the user is prompted that adjustment is not possible so that the user can provide other solutions.

[0018] In one possible design, the first text stream further includes a fifth task, which includes K task features, where K is a positive integer. Before displaying the first text stream, the method further includes: determining that there are identical task features among the N task features and the K task features, wherein the identical task features include identical task content; merging the first task and the third task so that the identical task features appear only once in the first text stream.

[0019] In the embodiments of this application, the electronic device can merge multiple tasks in the text stream to make the text stream simple, clear and easy to understand, making it convenient for users to quickly determine whether the electronic device's parsing of the tasks is correct, resulting in a better user experience.

[0020] In one possible design, executing the first task includes: executing the first task upon receiving a user's instruction to confirm the execution of the first task; or, starting a countdown and executing the first task when the countdown reaches 0, wherein the countdown duration is a preset duration.

[0021] In this embodiment of the application, the electronic device can perform the first task after user confirmation or after a preset time period, so as to ensure that the first task is accurate as much as possible, so that the electronic device performs the accurate task.

[0022] In one possible design, after the countdown is started, the method further includes: pausing the countdown when the user's gaze is detected on the display screen of the electronic device; and resuming the countdown when the user's gaze is detected leaving the display screen of the electronic device.

[0023] In this embodiment, after the electronic device starts the countdown, it pauses the countdown if it detects that the user is looking at the screen, because the user may modify the task and needs to be given time to do so; when it detects that the user is no longer paying attention to the screen, the countdown resumes. This approach allows the user to modify the task, is relatively intelligent, and provides a better user experience.

[0024] In one possible design, the first text stream includes an application that executes the first task, and the application is a first application. Executing the first task includes: when it is determined that the electronic device does not contain the first application, using a second application in the electronic device to execute the first task.

[0025] In this embodiment of the application, if the application executing the first task is the first application, but the electronic device does not contain the first application, a second application can be used to execute it, so as to ensure that the first task can be completed smoothly.

[0026] In one possible design, before using the second application in the electronic device to perform the first task, the method further includes: outputting a third prompt message, the third prompt message being used to prompt whether the first application is not installed and whether to use the second application to perform the first task; and receiving an instruction to confirm using the second application to perform the first task.

[0027] In this embodiment of the application, if the application executing the first task is the first application, but the electronic device does not contain the first application, a second application can be used to execute it with the user's consent, so as to ensure that the first task can be completed smoothly.

[0028] In one possible design, the first text stream does not include the application executing the first task, and executing the first task includes: identifying a first application in the electronic device, the first application having historically executed other tasks with the same task characteristics as the first task; outputting a fourth prompt message, the fourth prompt message being used to prompt whether to use the first application to execute the first task; and receiving an instruction to confirm using the first application to execute the first task.

[0029] In this embodiment, with the user's consent, the electronic device can use a first application that has historically performed other tasks with the same task characteristics as the first task to execute the first task. This method allows the electronic device to intelligently determine the application to execute even if the user does not input the application, which is convenient and provides a better user experience.

[0030] In one possible design, the first text stream includes an execution device for the first task, and the device is a first device. Executing the first task includes: sending the first task to the first device and executing the first task through the first device.

[0031] In this embodiment, a user inputs voice or text on a device, which displays a text stream and marks tasks within it. If a task requires execution by another device, it is sent to that device. This allows the user to quickly monitor whether a task assigned to another device has been correctly parsed, improving the user experience.

[0032] In one possible design, the first text stream does not include a device for executing the first task, and executing the first task includes: when it is detected that there are other devices around the electronic device capable of executing the first task, the electronic device sends the first task to the other devices; when it is detected that there are no other devices around the electronic device capable of executing the first task, the electronic device executes the first task.

[0033] In this embodiment, a user inputs voice or text on a device, which displays a text stream and marks tasks within the stream. If the task can be executed by another device, the system determines if a nearby device can execute the task. If so, the task is sent; otherwise, the task is executed locally. This allows for flexible execution of tasks across devices or locally, resulting in a better user experience.

[0034] In one possible design, performing the first task includes performing the first task while displaying the first text stream.

[0035] In this embodiment of the application, tasks in the text stream can be executed while the electronic device is displaying the text stream, resulting in timely response and a better user experience.

[0036] In one possible design, the method further includes: displaying a sixth task while displaying the first text stream, the sixth task being a task that is related to the first task based on the contextual semantics of the first text stream; and executing the sixth task.

[0037] In this embodiment, the electronic device can determine and execute associated tasks based on the contextual semantics of the text stream. For example, after receiving a call from Cai Cai, the electronic device captures the user's voice message, "Reply to her directly, I'm driving now, I'll call her back later." Based on the contextual semantics of the text stream, the electronic device determines that the associated task is to hang up the phone, displays the text "Hang up the phone," and executes the task. In this way, although the user does not explicitly give the task of "hanging up the phone," the electronic device can infer the task, display it, and execute it, making it more intelligent and providing a better user experience.

[0038] In one possible design, before executing the first task, the method further includes: acquiring a second text stream; when it is determined that the second text stream is used to modify the first task, modifying the first task according to the second text stream; and executing the first task, including: executing the modified first task.

[0039] In this embodiment, the electronic device can mark tasks in the text stream. If the user determines that the electronic device is parsing a task incorrectly, the task can be modified. The electronic device then executes the modified task. This approach helps the user correct deviations in a timely manner, ensuring that the electronic device executes tasks accurately, resulting in a better user experience.

[0040] In one possible design, the method further includes displaying a modification flag for the first task.

[0041] In this embodiment of the application, if the electronic device modifies the task, it can display a modification mark to show the modified content to the user in a conspicuous manner, so as to ensure that the electronic device performs the correct task.

[0042] In one possible design, the display method corresponding to each of the N task features is different; wherein, the different display methods corresponding to each of the N task features include: at least one of the font color, font style, and font background color of each of the N task features is different.

[0043] In this embodiment of the application, the electronic device can mark tasks in the text stream, and different task characteristics correspond to different display methods, so as to facilitate users to distinguish different task characteristics and to facilitate users to judge whether the electronic device parses the task correctly, thereby improving the user experience.

[0044] In one possible design, the method further includes: outputting a fifth prompt message, the fifth prompt message being used to prompt the user to confirm the first task, wherein the task characteristics of the first task are marked in the fifth prompt message; and receiving a confirmation instruction, the confirmation instruction being used to confirm the first task.

[0045] In this embodiment of the application, the electronic device can prompt the user to confirm the first task to ensure that the first task is accurate, so that the electronic device can perform the accurate task.

[0046] In one possible design, before executing the first task, the method further includes: when the first task lacks a first task characteristic, determining the first task characteristic based on the contextual semantics of the first text stream; adding the first task characteristic to the first text stream, wherein the first task characteristic includes at least one of the task type, task content, task object, required application, and required device of the first task.

[0047] In this embodiment of the application, the electronic device can supplement incomplete tasks (tasks lacking task characteristics) according to the contextual semantics of the text stream, so as to ensure that the task executed by the electronic device is complete and accurate.

[0048] In one possible design, displaying the first text stream includes: synchronously transmitting the first text stream to a display unit and a task identification unit; the display unit is used to display the first text stream; the task identification unit is used to identify a first task and the task characteristics of the first task in the first text stream; and, after identifying the task characteristics of the first task, for the already displayed portion of the first text stream, adjusting the task characteristics of the first task contained in the already displayed portion from the current first display mode to a second display mode corresponding to the task characteristics; for the undisplayed portion of the first text stream, displaying the task characteristics of the first task contained in the undisplayed portion according to a third display mode corresponding to the task characteristics; and continuing to display other content in the undisplayed portion according to the first display mode.

[0049] In this embodiment, text stream display and task recognition can be performed simultaneously. After a task is recognized, the display method of the task characteristics contained in the already displayed portion of the text stream is adjusted, and the task characteristics contained in the undisplayed portion are displayed according to the corresponding display method. In this way, the electronic device marks tasks while displaying the text stream, providing a timely response and allowing users to quickly see whether the electronic device's task parsing is correct, resulting in a better user experience.

[0050] In one possible design, the method further includes: displaying the execution progress of the first task.

[0051] In this embodiment, the electronic device displays the task execution progress, which makes it convenient for users to understand the progress and provides a better user experience.

[0052] In one possible design, the first task is to send a message to a target object using a first instant messaging application. Before performing the first task, the method further includes: displaying a chat window with the target object in the first instant messaging application; and entering the message in the chat window.

[0053] In this embodiment of the application, if the first task is to send information through an instant messaging application, a chat window can be displayed while the electronic device is displaying the text stream, and text can be entered in real time in the chat window, giving the user the feeling that they are entering information in the chat window, resulting in a better user experience.

[0054] In one possible design, the electronic device is in a network-free state, which includes: no SIM card inserted or no connection to a wireless communication network. The first text stream also includes a seventh task that requires network access. The method further includes: outputting a sixth prompt message to indicate that the electronic device is in a network-free state.

[0055] In this embodiment of the application, if the electronic device has no network, but the execution of the first task requires a network, the electronic device can prompt the user that there is no network, so as to avoid the task being unable to be executed without informing the user and affecting the user experience.

[0056] Secondly, a voice interaction method is also provided, applied to an electronic device, such as a mobile phone or tablet computer. The method includes: acquiring a first voice stream; displaying a first text stream, the first text stream being obtained from the first voice stream, the first text stream including a first task; starting a timer; when a user's gaze is detected on the display screen of the electronic device, controlling the timer to pause; when a user's gaze is detected leaving the display screen of the electronic device, controlling the timer to resume; and when the timer reaches a preset time, executing the first task.

[0057] In this embodiment, the electronic device can display a text stream converted from a speech stream. After recognizing the first task in the text stream, a timer is started. If user attention is detected, the timer pauses, allowing time for the user to review or modify the task. If user attention is no longer detected, the countdown resumes. This approach incorporates an attention mechanism during the countdown to ensure sufficient time for the user to review or modify the task, making it more intelligent and providing a better user experience.

[0058] In one possible design, the method further includes: displaying a first button; and immediately executing the first task upon receiving an operation on the first button.

[0059] In this embodiment of the application, the user can also trigger the first button to make the electronic device execute the task immediately without waiting for the countdown to reach 0, which provides a better user experience.

[0060] In one possible design, the method further includes: displaying the timing progress of the timer, the timing progress being located within the display area where the first button is located.

[0061] In this embodiment of the application, the timing progress is displayed in the display area where the first button is located, which can avoid excessive obstruction of the display screen and improve the interactive experience.

[0062] Thirdly, an electronic device is also provided, comprising:

[0063] Processor, memory, and one or more programs;

[0064] The one or more programs are stored in the memory, and the one or more programs include instructions that, when executed by the processor, cause the electronic device to perform the method provided in the first or second aspect above.

[0065] Fourthly, a computer-readable storage medium is also provided for storing a computer program that, when run on a computer, causes the computer to perform the methods provided in the first or second aspect above.

[0066] Fifthly, a computer program product is also provided, comprising a computer program that, when run on a computer, causes the computer to perform the methods provided in the first or second aspect above.

[0067] In a sixth aspect, a chip is also provided, which is coupled to a memory in an electronic device for calling a computer program stored in the memory and executing the technical solutions provided in the first or second aspect of the embodiments of this application. In the embodiments of this application, "coupling" means that two components are directly or indirectly combined with each other.

[0068] In a seventh aspect, a chip system is also provided, the chip system including a processing circuit and a storage medium, the storage medium storing instructions; when the instructions are executed by the processing circuit, they implement the method described in the first aspect above.

[0069] For the technical effects that can be achieved in the second to seventh aspects mentioned above, please refer to the description of the technical effects that can be achieved by the corresponding design scheme in the first aspect mentioned above. This application will not repeat them here. Attached Figure Description

[0070] Figure 1 is a schematic diagram of a voice interaction process provided in an embodiment of this application;

[0071] Figure 2 is another schematic diagram of the voice interaction process provided in an embodiment of this application;

[0072] Figure 3 is another schematic diagram of a voice interaction process provided in an embodiment of this application;

[0073] Figures 4A to 4C are schematic diagrams of a task marking method provided in an embodiment of this application;

[0074] Figures 5A and 5B are schematic diagrams of another task marking method provided in an embodiment of this application;

[0075] Figures 6A to 6G are schematic diagrams of another task marking method provided in an embodiment of this application;

[0076] Figures 7A to 7D are schematic diagrams of multitasking provided in an embodiment of this application;

[0077] Figures 8A and 8B are schematic diagrams illustrating a task cancellation method provided in an embodiment of this application;

[0078] Figure 9 is a schematic diagram of multi-task merging provided in an embodiment of this application;

[0079] Figures 10A to 10C are schematic diagrams of cross-device task execution provided in an embodiment of this application;

[0080] Figure 11 is a flowchart illustrating a task marking method provided in an embodiment of this application;

[0081] Figure 12 is a flowchart illustrating a task marking method provided in an embodiment of this application;

[0082] Figure 13 is a schematic diagram of an electronic device provided in an embodiment of this application;

[0083] Figure 14 is another schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0084] The following explanations of some terms used in the embodiments of this application are provided to facilitate understanding by those skilled in the art.

[0085] The embodiments of this application involve at least one, including one or more; where "multiple" means two or more. Furthermore, it should be understood that in the description of this specification, terms such as "first," "second," and "third" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order. For example, "first task" and "second task" do not represent the degree of importance of the two or their order, but are merely for descriptive distinction. In the embodiments of this application, "and / or" merely describes an association relationship, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0086] The directional terms mentioned in the embodiments of this application, such as "up", "down", "left", "right", "inner", and "outer", are only for reference to the directions in the accompanying drawings. Therefore, the directional terms used are for better and clearer explanation and understanding of the embodiments of this application, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this application.

[0087] References to "one embodiment," "in some examples," or "some embodiments" as described in the embodiments of this application mean that one or more embodiments of this specification include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in some examples," "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0088] The voice interaction method provided in this application can be applied to electronic devices. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, personal computer (PC), ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), or other portable devices; or it can be a wearable device such as a watch or bracelet; or it can be a smart home device such as a television or refrigerator; or it can be a smart office device such as a printer; or it can be a means of transportation, such as various types of vehicles, trains, or aircraft; or it can be an in-vehicle device, such as a smart cockpit or various systems within a smart cockpit, such as an in-vehicle infotainment system (IVI) or other systems; or it can be a virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, etc. In short, this application does not limit the specific type of electronic device. Furthermore, the operating system of the electronic device can be any operating system, such as Android. System, HarmonyOS system, system, system, system, system, system, The embodiments of this application are not limited to systems, etc.

[0089] In this embodiment, the electronic device has a voice interaction function. The voice interaction function is the function of the electronic device responding to voice signals (short for: voice) input by the user. For example, if the user says "Call Cai Cai," the electronic device will collect the voice and respond by dialing Cai Cai's number. Optionally, the voice interaction function can include two states: a wake-up state and a non-wake-up state. To avoid wasting power and accidental triggering, the voice interaction function is in a non-wake-up state by default, and enters the wake-up state upon receiving a wake-up command. In the wake-up state, the electronic device can collect voice and then respond to the voice. The wake-up command can be a wake-up operation or a wake-up word. Taking a wake-up operation as an example, the wake-up operation can be an operation on a specific button. The specific button can be the home button, volume button, a specific button, or other buttons. The operation on the specific button can be a long press operation or other operations. Taking a wake-up word as an example, the wake-up word can be a customized noun such as "Xiao Yi Xiao Yi" or "YOYO." Optionally, the voice interaction function can be integrated into a first application in the electronic device. The first application can be a system application or a third-party application; this embodiment does not limit this. Taking a system application as an example, the first application can be "Huawei Smart Assistant."

[0090] Optionally, once the voice interaction function of an electronic device enters the wake-up state, there are multiple ways to implement it.

[0091] In the first method, the electronic device collects voice input and responds after the voice input ends. For example, as shown in Figure 1(a), the electronic device displays interface 100. Interface 100 can be any interface, such as a lock screen, a black screen, a desktop, a negative one screen, an application interface, or other interfaces. In Figure 1(a), the electronic device enters a wake-up state when it receives a wake-up command (e.g., a long press operation on a specific button). Optionally, after entering the wake-up state, the electronic device can provide certain prompts, which can be voice prompts, text prompts, graphic prompts, etc. For example, the electronic device displays interface 101 as shown in Figure 1(b). Interface 101 includes a graphic 102 to prompt the user to speak. For example, the user speaks the voice "How is the weather today?" The electronic device can collect voice input and respond after determining that the voice input has ended (e.g., a long press operation on a specific button has ended). For example, the electronic device displays interface 103 as shown in Figure 1(c). Interface 103 includes the weather information retrieved by the electronic device. In the first approach, the electronic device remains unresponsive while the user is speaking. For example, in Figure 1(b), the electronic device does not respond while the user is speaking. In this case, the user cannot be certain whether their voice will be correctly interpreted and must wait for the electronic device to respond (e.g., outputting weather information) before judging the result. If the response does not meet the user's expectations, the user needs to speak again to get the electronic device to respond again. Therefore, this voice interaction method has high trial-and-error costs and a poor user experience, especially with long voice messages, where the impact on the user experience is more pronounced.

[0092] The second method involves the electronic device capturing voice input, converting it into text in real time, and displaying the real-time text. It then responds after the voice input ends. For example, as shown in Figure 2(a), the electronic device displays interface 200. Interface 200 operates on the same principle as interface 100 described earlier and will not be repeated. When the electronic device receives a wake-up command (e.g., a long press on a specific button), it displays interface 201 as shown in Figure 2(b). Interface 201 includes a prompt message 202 indicating that the electronic device is capturing voice input. For example, if a user says "What's the weather like today?", the electronic device captures the voice input, converts it into text in real time, and outputs the text in real time, as shown in Figures 2(c) and 2(d). After determining that the voice input has ended (e.g., a long press on a specific button has ended), the electronic device responds, for example, by displaying interface 203 as shown in Figure 2(e). Interface 203 includes weather information. In the first method, the electronic device needs to convert speech to text in real time, which can be achieved through Automatic Speech Recognition (ASR) or other technologies. In the second method, the electronic device displays the text converted from speech in real time, so users can see if the device has correctly interpreted their speech, improving the user experience. However, in this method, all text is displayed indiscriminately, making it difficult for users to quickly focus on important content (e.g., tasks or instructions that the electronic device needs to perform), especially when there is a lot of text.

[0093] In view of this, embodiments of this application provide a voice interaction method in which an electronic device can display text converted from speech in real time, and important content in the text (e.g., tasks or instructions that the electronic device needs to perform) is marked so that the user can quickly focus on the content, thereby determining whether the content has been correctly parsed, reducing trial and error costs, and improving user experience. Optionally, the important content can be tasks or instructions in the text that the electronic device needs to perform, hereinafter referred to as "tasks". For example, the task may include: communicating with a contact, opening content, playing multimedia, etc. Communicating with a contact may include: sending a message to a contact through an instant messaging application, making a phone call to a contact, etc. Opening content may include opening a webpage, calendar, weather, memo, etc. Playing multimedia may include playing music, playing video, etc.

[0094] For example, please refer to Figure 3, which is a schematic diagram of a voice interaction process provided in an embodiment of this application. As shown in Figure 3(a), the electronic device displays interface 300. Interface 300 is in the same principle as interface 100 described above, and will not be repeated. When the electronic device receives a wake-up command (e.g., a long press operation for a specific button), it displays interface 301 as shown in Figure 3(b). Interface 301 includes prompt information 302, which is used to prompt the electronic device that it is collecting voice. Taking the user's voice "Please send a message to Cai Cai on WeChat, saying that I can't attend the banquet today" as an example, the electronic device can collect the voice and display the text converted from the voice in real time. Moreover, during the display of the text, the task in the text is marked in real time. For example, as shown in Figure 3(c), at time T1 (e.g., 00:06), the electronic device outputs the text: Please send a message to Cai Cai on WeChat. At this time, the text "Please send a message to Cai Cai on WeChat" is displayed without distinction (e.g., the font, background, and fill colors are the same). As shown in Figure 3(d), at time T2 (e.g., 00:07), the electronic device continues to output the following text: "Just say today." Assuming that before time T2, the electronic device has already identified the task contained in the text as "Send a message to Cai Cai via WeChat, I can't attend the banquet today," the electronic device can mark the task in the text. Since a portion of the text, "Send a message to Cai Cai via WeChat, just say today," has already been displayed at time T2, the electronic device can mark the task-related content within this portion. For example, it can adjust the font background color of "Send a message to Cai Cai via WeChat" and "today" to gray or another color. Content unrelated to the task (e.g., words like "please" and "just say") continues to maintain its original display method, i.e., the font background fill color is not adjusted. As shown in Figure 3(e), at time T3 (e.g., 00:09), the electronic device continues to output the following text: "Can't attend the banquet." Since "Can't attend the banquet" has been identified as task-related content, when the electronic device displays the text "Can't attend the banquet," the background fill color is adjusted to gray in real time. Once the electronic device determines that the voice input has ended (e.g., the long press operation on a specific key has ended), it can execute the marked task. Therefore, in the process shown in Figures 3(c) to 3(e), the electronic device displays a text stream and marks the tasks in the text stream in real time so that the user can quickly focus on the task and thus determine whether the task has been correctly parsed, reducing trial and error costs.

[0095] As mentioned earlier, when an electronic device displays a text stream, the tasks or instructions within the text stream are marked in real time. Optionally, the marking methods can include multiple methods, such as circling with a specific graphic (rectangle, ellipse, etc.), or using at least one of a specific font color, a specific font background color, or a specific font style. Optionally, the specific font style can include at least one of a specific font (e.g., SimSun, KaiTi, etc.), italics, underline, bold, or enlarged font.

[0096] Let's take the example of using a specific graphic to circle the text. As shown in Figure 4A(a), at time T1, the electronic device outputs the text: "Please send a message to Cai Cai via WeChat." At this time, the text "Please send a message to Cai Cai via WeChat" is displayed without distinction. As shown in Figure 4A(b), at time T2, the electronic device continues to output the following text: "Just say today." Assuming that before time T2, the electronic device has already recognized that the task contained in the text is "Send a message to Cai Cai via WeChat, I can't attend the banquet today," the electronic device can mark the task in the text. Since a part of the text, "Send a message to Cai Cai via WeChat, just say today," has already been displayed at time T2, the electronic device can mark the content related to the task in this part, such as circling "Send a message to Cai Cai via WeChat" and "today" with rectangles or ellipses, while content unrelated to the task (such as words like "please" and "just say") does not need to be circled. As shown in Figure 4A(c), at time T3, the electronic device continues to output the following text: "Can't attend the banquet." Since "cannot attend the banquet" has been identified as task-related content, it is highlighted in real time when the electronic device displays the text "cannot attend the banquet".

[0097] Taking a specific font style (e.g., underline) as an example. As shown in Figure 4B(a), at time T1, the electronic device outputs the text: Please send a message to Cai Cai via WeChat. At this time, the text "Please send a message to Cai Cai via WeChat" has no underline. As shown in Figure 4B(b), at time T2, the electronic device continues to output the following text: Just say today. Assume that before time T2, the electronic device has already recognized that the task contained in the text is "Send a message to Cai Cai via WeChat, I can't attend the banquet today", so the electronic device can mark the task in the text. Since part of the text, namely "Send a message to Cai Cai via WeChat, just say today", has already been displayed at time T2, the electronic device can mark the content related to the task in this part, such as adding underlines to "Send a message to Cai Cai via WeChat" and "today", while content unrelated to the task (e.g., words like "please" and "just say") does not need to be underlined. As shown in Figure 4B(c), at time T3, the electronic device continues to output the following text: I can't attend the banquet. Since "cannot attend the banquet" has been identified as task-related content, electronic devices will underline the text "cannot attend the banquet" by default.

[0098] Taking a specific font style (e.g., italics) as an example. As shown in Figure 4C(a), at time T1, the electronic device outputs the text: "Please send a message to Cai Cai via WeChat." At this time, the font of "Please send a message to Cai Cai via WeChat" is not italic. As shown in Figure 4C(b), at time T2, the electronic device continues to output the following text: "Just say today." At this time, the fonts of "Send a message to Cai Cai via WeChat" and "Today" are adjusted to italics, for the same reason as described above, which will not be repeated here. As shown in Figure 4C(c), at time T3, the electronic device continues to output the following text: "Cannot attend the banquet." The font of "Cannot attend the banquet" is italic.

[0099] In the above embodiments, during the display of the text stream by the electronic device, tasks within the text stream are marked in real time. In this embodiment, a task may include at least one of the following: task type, task content, task object, and required application. The task type may include sending a message, making a phone call, opening content, playing multimedia, etc. The task content may include: what message to send, what content to open, what multimedia to play, etc. The task object may include: who to send the message to, who to call. The required application is the application used to perform the task, such as WeChat, telephone, etc. Continuing with the example of the task "Send a message to Cai Cai on WeChat saying I can't attend the banquet today," this task may include four parts: task type, task object, task content, and required application. For ease of understanding, please refer to Table 1 below:

[0100] Table 1

[0101] Optionally, the four parts—task type, task content, task object, and required application—can use the same or different tagging methods. Continuing with Table 1 above as an example, the task type can use tagging method 1, the task object can use tagging method 2, the task content can use tagging method 3, and the required application can use tagging method 4. Tagging methods 1 through 4 can be the same or different.

[0102] For example, as shown in Figure 5A(a), at time T1, the electronic device outputs the text: "Please send a message to Cai Cai using WeChat." At this time, "Please send a message to Cai Cai using WeChat" is displayed indiscriminately. As shown in Figure 5A(b), at time T2, the electronic device continues to output the following text: "Just say today." Assuming that before time T2, the electronic device has recognized that the task contained in the text is "Send a message to Cai Cai using WeChat, I can't attend the banquet today," and has determined that the task type is "send a message," the task object is "Cai Cai," the required application is "WeChat," and the task content is "I can't attend the banquet today," then the electronic device can determine (for example, according to Table 1 above) that the marking method corresponding to the task type is marking method 1, the marking method corresponding to the task object is marking method 2, the marking method corresponding to the task content is marking method 3, and the marking method corresponding to the required application is marking method 4. Therefore, the electronic device can use marking method 1 to mark the task type, marking method 2 to mark the task object, marking method 3 to mark the task content, and marking method 4 to mark the required application. Since a portion of the text, "Send a message to Cai Cai via WeChat today," has already been displayed at time T2, the electronic device can mark the task-related content within that portion using corresponding marking methods. For example, the required application "WeChat" is circled (marking method 4), the task type "send a message" is underlined (marking method 1), the background color of the task object "Cai Cai" is adjusted to gray (marking method 2), and the background color of the task content "today" is adjusted to gray (marking method 3). As shown in Figure 5A(c), at time T3, the electronic device continues to output the subsequent text: "Cannot attend the banquet." Since "Cannot attend the banquet" belongs to the task content, it is marked using marking method 3, i.e., the background color of the font is adjusted to gray. Therefore, through the method shown in Figure 5A, users can quickly distinguish between the task type, task content, task object, and required application, allowing users to judge the accuracy of the parsing results for each task type, task content, task object, and required application, thus improving the user experience.

[0103] Understandably, the task may include names such as application names, contact names, etc. Optionally, the electronic device can replace these names with corresponding icons. For example, an application name can be replaced with its corresponding application icon. Similarly, a contact name can be replaced with the contact's profile picture. Continuing with the task "Send a message to Cai Cai via WeChat saying you can't attend the banquet today," the task includes the application name "WeChat," which can be replaced with the WeChat application icon. The task also includes the contact name "Cai Cai," which can be replaced with the corresponding profile picture. For example, as shown in Figure 5B(a), at time T1, the electronic device outputs the text: "Please send a message to Cai Cai via WeChat." At this time, "Please send a message to Cai Cai via WeChat" is displayed indiscriminately. As shown in Figure 5B(b), at time T2, the electronic device continues to output the following text: "Just say today." Assuming that before time T2, the electronic device has recognized that the task includes the application name "WeChat," it can replace this application name with its corresponding icon, as shown in Figure 5B(b). Furthermore, an underline is added to the task type "Send Message," and the background color of the font for the task object "CaiCai" and the task content "Today" is adjusted to gray. The principle behind this has been described previously and will not be repeated here. As shown in Figure 5B(c), at time T3, the electronic device continues to output the following text: "Cannot attend the banquet." The background color of the font for "Cannot attend the banquet" is gray.

[0104] In some embodiments, the electronic device can also output the execution progress of the task. The execution progress can include "in progress" or "completed". Continuing with the example of the task "Send a message to Cai Cai via WeChat saying you can't attend the banquet today". As shown in Figure 6A(a), at time T1, the electronic device outputs the text: "Please send a message to Cai Cai via WeChat". At this time, "Please send a message to Cai Cai via WeChat" is displayed indiscriminately. As shown in Figure 6A(b), at time T2, the electronic device continues to output the following text: "Just say today". Assuming that before time T2, the electronic device has recognized the task contained in the text as "Send a message to Cai Cai via WeChat saying you can't attend the banquet today", the electronic device can pop up the chat window of the WeChat application. The chat window includes the task object, "Cai Cai", and also includes a text input box for real-time input of the task content, "today". As shown in Figure 6A(c), at time T3, the electronic device continues to output the following text: "Can't attend the banquet". Moreover, "Can't attend the banquet" is also entered in real-time within the chat window. Therefore, while the electronic device is outputting text, the user is also inputting text in real-time within the chat window, giving the user the experience of real-time input of dialogue content.

[0105] In some embodiments, the electronic device can execute a task when it recognizes a complete task. Optionally, the electronic device can determine whether a task is complete in several ways. Method A: The electronic device determines whether a task is complete through contextual semantic recognition. For example, in Figure 6A(b), the electronic device determines that "Send a message to Cai Cai on WeChat saying 'Today'" is an incomplete sentence, thus determining the task to be incomplete. For example, in Figure 6A(c), the electronic device determines that "Send a message to Cai Cai on WeChat saying 'I can't attend the banquet today'" is a complete task. Method B: The electronic device determines that the task is complete when it detects the end of voice input. Assuming that in Figure 6A(b), the electronic device detects the end of voice input (e.g., the end of a long press on a specific button), then it determines that "Send a message to Cai Cai on WeChat saying 'Today'" is a complete task. Assuming that in Figure 6A(c), the electronic device detects the end of voice input (e.g., the end of a long press on a specific button), then it determines that "Send a message to Cai Cai on WeChat saying 'I can't attend the banquet today'" is a complete task. When an electronic device performs a task, it can display a prompt message, such as the "Sending" message shown in Figure 6A(d). When the electronic device determines that the task has been completed, it can also display a prompt message, such as the "Sent" message shown in Figure 6A(e).

[0106] As mentioned above, when an electronic device recognizes a complete task, it can execute that task. Optionally, when an electronic device recognizes a complete task, it can execute the task automatically or based on manual user input.

[0107] Taking automatic execution as an example, for instance, in Figure 6A(c), when the electronic device recognizes the complete task, it can immediately execute the task automatically, or start a countdown with a preset duration (e.g., 3s or 5s). When the countdown reaches 0, the task is executed automatically.

[0108] Taking manual execution as an example, for instance, in Figure 6A(c), when the electronic device recognizes the complete task, it can display the "Confirm Send" button as shown in Figure 6B. When the electronic device receives an operation on this button, it executes the task, i.e., sends the information. Alternatively, when the electronic device receives a voice command to confirm sending, it executes the task, i.e., sends the information.

[0109] The two methods described above can be used individually or in combination. Taking combined use as an example, when the electronic device recognizes a complete task, it can display a "Confirm Send" button as shown in Figure 6C(a) and start a countdown for a preset duration. Optionally, the electronic device can also display the countdown progress. To avoid obscuring the interface, the countdown progress can be displayed on the "Confirm Send" button. For example, as shown in Figure 6C(b), the "Confirm Send" button includes a countdown progress bar. When the countdown reaches 0, as shown in Figure 6C(c), the countdown progress bar in the "Confirm Send" button is full, and the electronic device can automatically send the information. It is understood that before the countdown reaches 0, for example, in Figure 6C(b), if the electronic device receives an operation on the "Confirm Send" button, it can immediately execute the task, i.e., send the information.

[0110] Considering that users may modify the information to be sent, the electronic device can pause the countdown when it determines that the user intends to modify it. For example, in Figure 6C(b), when the electronic device detects that the user's gaze is on the display screen, it can pause the countdown, and the progress bar in the "Confirm Send" button will pause. During the paused countdown, if the electronic device receives a modification operation on the text, it will modify the text. When the electronic device detects that the user's gaze has left the display screen, the countdown resumes, and the progress bar in the "Confirm Send" button continues. When the countdown reaches 0, the electronic device automatically sends the message. During this process, the user can see whether the electronic device has correctly interpreted their voice, and can also modify the task, improving the user experience.

[0111] It should be noted that Figure 6C uses manual text modification by the user as an example. Optionally, text can also be modified via voice. For example, see Figure 6D(a), where the user speaks "Play Fei Fei's songs." After the electronic device captures this speech, it displays the corresponding text stream "Play Fei Fei's songs," and the task in the text stream is marked. It should be noted that the electronic device may parse words with the same pronunciation incorrectly. For example, the user wants to express "fei," but the electronic device parses it as "fei." In this case, after seeing the parsing result of the electronic device, i.e., the text stream in Figure 6D(a), the user can modify the text via voice. For example, the user speaks "It's the 'fei' of 'flying.'" In this case, as in Figure 6D(b), the electronic device continues to output the text: "It's the 'fei' of 'flying.'" The electronic device can modify the preceding text, for example, displaying the modification method shown in Figure 6D(c).

[0112] Optionally, before modifying the text, the electronic device can also output a prompt message to indicate whether the text should be modified. Taking Figure 6D(b) as an example, the electronic device can also display the prompt message "Do you want to modify Feifei?" as shown in Figure 6D(d). When the electronic device receives the instruction to confirm the modification, it displays the modification method shown in Figure 6D(c). Optionally, when the electronic device outputs the prompt message, it can also use the task marking method provided in the embodiments of this application to mark the task characteristics in the prompt message (please refer to the previous description for task characteristics). For example, in Figure 6D(d), "Feifei" is underlined, and "Feifei" is in italics and bold.

[0113] It should be noted that during the process shown in Figure 6D, the user utters two voice messages. One is "Play Fei Fei's songs," and the other is "It's the Fei of flying." Optionally, these two voice messages can be different parts of the same voice message. For example, in an exemplary scenario, the user presses and holds a specific button (e.g., the home button) and utters the voice message "Play Fei Fei's songs." The electronic device displays a text stream in real time. The user does not release the button at this point. When the user sees the text stream displayed on the electronic device as "Play Fei Fei's songs," they continue to utter the voice message "It's the Fei of flying." At this point, the user can release the button. Alternatively, the user can release the button after seeing the electronic device correct any typos. Optionally, the two voice messages can also be two independent voice messages. For example, in an exemplary scenario, the user presses and holds a specific button (e.g., the home button) and utters the voice message "Play Fei Fei's songs," then immediately releases the button. After the electronic device displays the text stream in real time, the user sees the text stream displayed as "Play Fei Fei's song". The user then presses and holds a specific button again to produce the voice message "It's Fei as in flying", and immediately releases the button. The electronic device then captures the next voice message ("It's Fei as in flying") and modifies the text stream of the previous voice message.

[0114] It is understandable that in Figure 6D, there is a possible scenario where the electronic device has already executed the task before modifying it. In this case, the electronic device can cancel the execution of the task, then modify the task and execute the modified task. For example, see Figure 6E(a), where the user utters the voice "Play Fei Fei's songs," which is captured by the electronic device, displaying the corresponding text stream "Play Fei Fei's songs," and the task in the text stream is marked. As shown in Figure 6E(b), at time T2, the electronic device has already started executing the task "Play Fei Fei's songs," and displays the prompt message "Searching for Fei Fei's songs." Assume that before time T2, the electronic device captured the voice "It's Fei Fei (as in Fei Xiang)," and outputs the text stream "It's Fei Fei (as in Fei Xiang)" at time T2. As shown in Figure 6E(c), at time T3, the electronic device modifies the task. The electronic device can cancel the task "Play Fei Fei's songs" and display a prompt message indicating task cancellation. The electronic device can also execute the task "Play Fei Fei's songs" and display the prompt message "Searching for Fei Fei's songs."

[0115] In practical applications, there exists a situation where an electronic device determines that a task lacks a specific feature. In this case, it can set up slots (or empty spaces) in the text stream to prompt the user to add the missing feature. For example, a user might say, "Send a message on WeChat saying I'm busy today." The electronic device captures the speech and displays the text converted from the speech, as shown in Figure 6F(a), and the task is marked. The electronic device then determines the task recipient (i.e., who to send the message to) within the task "Send a message on WeChat saying I'm busy today." In this scenario, the electronic device can display a slot. After seeing the slot, the user can say, "Send it to Cai Cai." The electronic device displays the text "Send it to Cai Cai," as shown in Figure 6F(b). The electronic device can adjust the word order of the text, for example, by filling the slot with "Cai Cai," as shown in Figure 6F(c). It's understandable that the user may not necessarily notice the slot, so the electronic device can also remind the user to add the required feature, such as the task recipient, to the slot. For example, in Figure 6F(a), due to the presence of a slot, i.e., the lack of a task feature (i.e., a task object), the electronic device can output a prompt message to remind the user to complete it.

[0116] In some embodiments, the task requires an electronic device to be connected to a network to execute; if the electronic device is without a network, the task cannot be executed. Optionally, being without a network can include: no SIM card inserted, no SIM card signal, or no access to a wireless communication network (e.g., Wi-Fi, 4G, 5G, 6G, etc.). In this case, the electronic device can output a prompt message to indicate that there is currently no network. For example, see Figure 6G(a), where the electronic device outputs a text stream: "Send a message to Cai Cai to say goodbye today." Assuming the electronic device recognizes that the task requires an SMS application, it can determine whether the SIM card has a signal. Assuming it recognizes that no SIM card is inserted, it outputs the prompt message shown in Figure 6G(b) to indicate that there is no SIM card. In this case, the electronic device can determine whether there are other applications that can send messages to Cai Cai. If so, it can recommend other applications to the user. For example, if the electronic device determines that it is currently connected to Wi-Fi and that WeChat is installed on the device, and that there is a contact named "Cai Cai" in WeChat, the electronic device can recommend that the user use WeChat to send a message to Cai Cai. If the electronic device receives a user consent instruction, it uses WeChat to send a message to Cai Cai.

[0117] For example, an electronic device outputs a text stream: Sending a message to Cai Cai via WeChat saying that I can't attend the banquet today. The electronic device recognizes that this task requires the WeChat application. It can determine if it's currently connected to the internet; if not, it can automatically connect (e.g., by turning on mobile data). Alternatively, it can output a prompt asking the user whether to connect. After receiving the user's consent, it connects to the network. Once connected, the electronic device executes the task.

[0118] In practical applications, there is a possibility that a user's voice message may contain multiple different tasks. As mentioned earlier, a task can include at least one of the following: task type, task object, task content, and required application. Different tasks can include at least one of the following: different task types, different task objects, different task content, and different required applications. For example, if a user's voice message is "First, send a message to Cai Cai on WeChat saying I can't attend the banquet today, then send a message to Zhao Si saying I can attend tonight's meeting," this voice message contains two tasks: Task 1 is "Send a message to Cai Cai on WeChat saying I can't attend the banquet today"; Task 2 is "Send a message to Zhao Si saying I can attend tonight's meeting." Task 1 and Task 2 have different task content (i.e., different message content) and different task objects (i.e., different recipients), therefore they are different tasks.

[0119] In cases where a voice message includes multiple tasks, one of the tasks may be incomplete. In such situations, the electronic device can use other tasks to complete the incomplete task. Completing a task can include at least one of the following: task type, task object, task content, required application, and required device. For example, let's continue with the example of a user sending the voice message, "First, send a message to Cai Cai on WeChat saying I can't attend the banquet today, then send a message to Zhao Si saying I can attend the meeting tonight." Task 1 is "Send a message to Cai Cai on WeChat saying I can't attend the banquet today"; Task 2 is "Send a message to Zhao Si saying I can attend the meeting tonight." Task 2 doesn't specify which application to use, so it can be considered incomplete. In this case, the electronic device can complete Task 2 based on Task 1. For example, if Task 1 uses WeChat to send a message, then the electronic device determines that Task 2 also uses WeChat to send a message. Therefore, the completed Task 2 is "Send a message to Zhao Si on WeChat saying I can attend the meeting tonight."

[0120] Optionally, when a voice recording includes multiple tasks, the electronic device can use the same marking method for multiple tasks during real-time task marking, such as using the same font color, font style, or font background color. Alternatively, to facilitate user differentiation, multiple tasks can use different marking methods. Optionally, using different marking methods for multiple tasks can include at least one of the following: using different shapes to circle, using different font colors, different font background colors, or different font styles. The shapes include squares, rectangles, circles, ellipses, etc.

[0121] Continuing with the example of a user's voice message: "First, send a message to Cai Cai on WeChat saying I can't attend the banquet today, then send a message to Zhao Si saying I can attend tonight's meeting," as shown in Figure 7A(a), at time T1, the electronic device outputs the text: "Please send a message to Cai Cai on WeChat." At this time, "Please send a message to Cai Cai on WeChat" is displayed without any difference. As shown in Figure 7A(b), at time T2, the electronic device continues to output the following text: "Just say today." At this time, "WeChat" is replaced with an icon, "Send a message" is underlined, and the background color of "Cai Cai" and "Weather" is adjusted to gray. As shown in Figure 7A(c), at time T3, the electronic device continues to output the following text: "Can't attend the banquet." The background color of "Can't attend the banquet" is gray. As shown in Figure 7A(d), at time T4, the electronic device continues to output the text "Then send it to Li Si," which is displayed without any difference. As shown in Figure 7A(e), at time T5, the electronic device continues to output the text "I can attend." Assuming that before time T5, the electronic device has recognized that the speech contains two tasks, and the previous task has been marked, then the electronic device can use different marking methods to mark the next task. For example, at time T5, "Li Si, I can participate" is circled. As shown in Figure 7A(f), at time T6, the electronic device continues to output the text "Tonight's meeting." Here, "Tonight's meeting" is circled. As shown in Figure 7A, the electronic device can display a text stream converted from speech and mark the tasks in the text stream in real time, using different marking methods for different tasks, allowing users to quickly distinguish between different tasks and providing a better user experience.

[0122] In some embodiments, when a voice message includes multiple tasks, the electronic device can output the execution progress of each task separately. For example, as shown in Figure 7B, at time T1, the electronic device outputs the text: "Please send a message to Cai Cai using WeChat." At this time, "Please send a message to Cai Cai using WeChat" is displayed indiscriminately. Assuming that the electronic device has recognized that the task type is sending a message, the task recipient is Cai Cai, and the application used is WeChat, the electronic device will pop up a WeChat chat window. The chat window includes the recipient, "Cai Cai," and also includes a text input box for real-time input of dialogue content. It should be understood that the text input box is empty at this time. At time T2, the electronic device continues to output the text: "Just say today." Moreover, the electronic device enters the text "Today" in the WeChat chat window. At time T3, the electronic device continues to output the text: "Can't attend the banquet," and also enters "Can't attend the banquet" in real-time in the chat window. At time T4, the electronic device continues to output the text: "Then send it to Li Si." Assuming that before time T4, the electronic device has determined that the previous task is a complete task, then at time T4, the electronic device can automatically execute the previous task, that is, send a message to Cai Cai. Therefore, at time T4, the electronic device can display a "Sending" message. At time T5, the electronic device continues to output the text: "I can attend." At this time, another chat window pops up, including the recipient "Li Si" and a text input box for real-time input of the dialogue content "I can attend." Assuming that the electronic device has determined that the previous task has been completed before time T5, then at time T5, the electronic device can display a "Sent" message. At time T6, the electronic device continues to output the text: "Tonight's meeting." Furthermore, the dialogue content "Tonight's meeting" is added to the chat window in real-time. It should be understood that if the electronic device determines that the next task is complete, it can automatically execute that task, i.e., send the message to Li Si (this process is not shown in Figure 7B).

[0123] It should be noted that in Figure 7B above, the example is that the electronic device automatically executes each complete task it detects immediately. Optionally, the electronic device can also start a countdown when it detects a complete task, and automatically execute the task when the countdown reaches 0. For example, if Figure 7B includes two tasks, two countdowns can be started, one for each task. Alternatively, when the electronic device detects the first complete task, it can not start the countdown initially, but start the countdown when it detects the second complete task, and execute both tasks together when the countdown reaches 0. In this way, one countdown is sufficient for two tasks. In other embodiments, when the electronic device detects the first complete task, it can not execute the task immediately, but instead display a "Confirm Send" button. The electronic device executes the first task when it receives an operation on the "Confirm Send" button or receives a voice command from the user confirming the send. The same principle applies to the next task. When the electronic device detects a second complete task, it can initially delay execution, displaying a "Confirm Send" button. The second task is executed only upon receiving an action on the "Confirm Send" button or a voice command from the user confirming the send. This method requires displaying two "Confirm Send" buttons because the voice input includes two tasks. Alternatively, when the electronic device detects the first complete task, it can delay execution, checking if the voice input has ended (e.g., whether a long press on a specific button has ended). If not, it waits. Similarly, when the electronic device detects the second complete task, it delays execution, checking if the voice input has ended. If it has ended (e.g., whether a long press on a specific button has ended), it displays a single "Confirm Send" button. Upon receiving an action on the "Confirm Send" button or a voice command from the user confirming the send, both tasks can be executed simultaneously. This method requires only one "Confirm Send" button for each task.

[0124] For another example, an electronic device receives a call from Cai Cai. The user can hang up the phone via voice and reply to Cai Cai. As shown in Figure 7C, at time T1, the electronic device outputs the text: "Hang up the phone, then." At this time, the text is displayed without any difference. At time T2, the electronic device continues to output the text: "Reply to him." At this time, "Hang up the phone" is marked. At time T3, the electronic device continues to output the text: "I am driving now." Here, "Reply to him, I am driving now" is marked, and the marking method can be different from the marking method of the previous task (hanging up the phone). The electronic device also displays a chat window and enters in the chat window: "I am driving now." Assuming that at time T3, the electronic device begins to execute the previous task, i.e., "hang up the phone," it will display a prompt message indicating that it is in progress, such as "Hanging up." At time T4, the electronic device continues to output the text: "I'll call you back later." This text is marked. Moreover, the electronic device enters in real time in the chat window: "I'll call you back later." Assuming that at time T4, the electronic device has successfully hung up the phone, it will display a message indicating that the phone was successfully hung up. Assuming at time T5, the electronic device begins executing the task of sending information for the next day, it will display a "Sending" message. At time T6, the electronic device confirms that the information has been sent and displays a "Sent successfully" message.

[0125] It should be noted that in Figure 7C, the user's voice contains the task of "hanging up the phone." It's understandable that there's a possibility that the user didn't explicitly say "hang up the phone," but instead said, "Reply to him directly, I'm driving now, I'll call him back later." That is, the voice doesn't explicitly state the task of "hanging up the phone," but the context of the electronic device's voice indicates an implicit task: hanging up the phone. In this case, the electronic device can output the implicit task. For example, as shown in Figure 7D, at time T1, the electronic device outputs the text: "Reply to him directly." At this time, the text is displayed without distinction. At time T2, the electronic device continues to output the text: "I'm driving now." Assuming that the electronic device recognized the implicit task "hang up the phone" before time T2, then at time T2, it can output the implicit task "hang up the phone," and "hang up the phone" will be marked. Assuming that at time T2, the electronic device has already started executing the "hang up the phone" task, then it displays a prompt message indicating that it is being executed, such as "hanging up." In addition, the electronic device also displays a chat window and allows the user to type: "I'm driving now." At time T3, the electronic device continues to output the text: "I'll call you back later." The electronic device also inputs "I'll call you back later" in real-time within the chat window. Assuming that at time T4, the electronic device has successfully hung up the phone, it will display a message indicating successful hang-up.

[0126] In practical applications, there's a possibility that a user's voice message might contain multiple tasks, with later tasks canceling earlier ones. For example, a voice message might include Task 1 and Task 2, with Task 1 being sent earlier and Task 2 cancelling Task 1. In this case, when the electronic device detects that Task 2 is a complete task, it can execute Task 2, thus canceling Task 1. Optionally, if Task 1 hasn't been executed yet, it can be canceled; if Task 1 has been completed, there are two possible handling methods. Method A: The electronic device outputs a prompt indicating that Task 1 has been completed. In this case, the user might send another voice message, and the electronic device can respond accordingly. Method B: The electronic device automatically executes a corresponding remedial strategy. Taking Task 1 as an example of hanging up the phone, if Task 1 has been completed, the electronic device can execute a remedial strategy, i.e., dial the other party's number.

[0127] For example, as shown in Figure 8A, at time T1, the electronic device outputs the text: "Hang up the phone, then." At this time, the text is displayed without distinction. At time T2, the electronic device continues to output the text: "Reply to him." Here, "Hang up the phone" is marked. At time T3, the electronic device continues to output the text: "I am driving now." Here, "Reply to him, I am driving now" is marked, and the marking method is different from the marking method of the previous task. The electronic device also displays a chat window and enters in the chat window: "I am driving now." Assuming that at time T3, the electronic device has started executing the "Hang up the phone" task, it displays the "Hanging up" prompt message. At time T4, the electronic output continues to output the text: "Don't reply, answer the phone." Here, "Don't reply" and "Answer the phone" are marked respectively, and the marking method can be different from the marking method of the previous task. Therefore, the electronic device executes the task of canceling the reply message, and optionally, it can also display the cancel reply prompt message. Moreover, the electronic device also executes the tasks of canceling hanging up the phone and answering the phone, and optionally, it can also display the cancel hang-up prompt message and the "Answering" prompt message. Optionally, in Figure 8A, the electronic device can mark the canceled task with strikethrough or other markers.

[0128] For example, as shown in Figure 8B, at time T1, the electronic device outputs the text: "Hang up the phone, then." At this point, the text is displayed indiscriminately. At time T2, the electronic device continues to output the text: "Reply to him." Here, "Hang up the phone" is marked. At time T3, the electronic device continues to output the text: "I am driving now." Here, "Reply to him, I am driving now" is also marked, but in a different way than the previous task. The electronic device also displays a chat window and enters: "I am driving now." Assuming that at time T3, the electronic device has started executing the "Hang up the phone" task, it displays the message "Hanging up." At time T4, the electronic device continues to output the text: "Don't reply, answer the phone." Here, "Don't reply" and "Answer the phone" are marked in different ways. Therefore, the electronic device executes the task of canceling the reply message, and optionally, it can also display a message indicating that the reply has been canceled. Assuming that at time T4, the electronic device has successfully hung up the phone and displays the message "Hang up successfully." The electronic device can employ a remedial strategy; for example, at time T5, the electronic device initiates a call, and optionally, it can also display a message indicating that a call has been initiated. Optionally, the electronic device can output a voice prompt before initiating a call, such as "Hanged up, do you want to start a call?" If the electronic device receives a voice command confirming the start of the call, then the call is initiated. Optionally, in Figure 8B, the electronic device can mark canceled tasks with strikethrough or other markers.

[0129] In practical applications, when a single voice recording includes multiple tasks, the text output by the electronic device is often lengthy, making it difficult for the user to focus on the tasks. To simplify the text, the electronic device can merge multiple tasks. As mentioned earlier, a task includes at least one of the following: task type, task object, task content, and required application. Optionally, the electronic device can merge multiple tasks that share at least one of the following characteristics: task type, task object, task content, and required application.

[0130] Let's take the example of a user sending the voice message, "Turn the bedroom air conditioner to 25 degrees Celsius, adjust the bedroom air conditioner, send a WeChat message to Cai Cai telling her I can't attend the banquet today, and then send it to Li Si, 'I can't attend the banquet today.'" For example, as shown in Figure 9(a), the electronic device outputs the text: "Turn the bedroom air conditioner to 25 degrees Celsius, adjust the bedroom air conditioner, send a WeChat message to Cai Cai telling her I can't attend the banquet today, and then send it to Li Si, 'I can't attend the banquet today.'" This text includes four tasks: Task 1 "Turn the bedroom air conditioner to 25 degrees Celsius"; Task 2 "Adjust the bedroom air conditioner"; Task 3 "Send a WeChat message to Cai Cai telling her I can't attend the banquet today"; Task 4 "Send it to Li Si, 'I can't attend the banquet today.'" If these four tasks are not combined, as shown in Figure 9(a), the text is too long and obscures the display interface significantly. Therefore, the electronic device can combine the four tasks. For example, Task 1 and Task 2 have the same task type, "adjust the air conditioner temperature," and the same task content, "adjust to 25 degrees Celsius." Therefore, the electronic device can merge Task 1 and Task 2, as shown in Figure 9(b). Furthermore, the electronic device determines that Task 3 and Task 4 have the same task type, "send a message," and the same task content, "I cannot attend the banquet today." Therefore, the electronic device can merge Task 3 and Task 4, as shown in Figure 9(b). Comparing Figure 9(a) and Figure 9(b), the merged interface is cleaner and more organized.

[0131] In the above embodiment, the user's voice message is "Send a message to Cai Cai via WeChat saying that I can't attend the banquet today." It should be understood that in actual use, the user's voice message might be "Send a message to Cai Cai saying that I can't attend the banquet today." That is, the user's voice message doesn't explicitly specify the application. In this case, the electronic device can determine which application to send the message to Cai Cai through. One possible approach is for the electronic device to iterate through the contact lists of various instant messaging applications to determine which instant messaging application the contact "Cai Cai" is in. If she is in WeChat, the message is sent via WeChat; if she is in the address book, the message is sent via SMS. Another possible scenario is that the contact "Cai Cai" is in multiple instant messaging applications. In this case, the electronic device can select one of the multiple instant messaging applications and send the message to "Cai Cai" through that application. Optionally, selecting one of the multiple instant messaging applications can include: randomly selecting an application, selecting an application frequently used by the user, or selecting an application with chat history with "Cai Cai."

[0132] In other embodiments, consider the possibility that the user's voice message is, "Send a message to Cai Cai on WeChat saying I can't attend the banquet today," but WeChat is not installed on the electronic device. In this case, the electronic device can output a prompt message to indicate that WeChat is not installed. For example, the electronic device might play the voice message: "WeChat is not installed." Optionally, the electronic device can also output recommendation information to suggest other applications. For example, the electronic device might play the voice message: "We recommend you use SMS." When the electronic device receives the user's consent, it uses the recommended application (i.e., SMS) to send a message to "Cai Cai." Of course, if the electronic device determines that WeChat is not installed, it can also directly use other applications (such as SMS) to send a message to "Cai Cai" without prompting the user.

[0133] In the above embodiments, taking the voice interaction method provided in this application embodiment as an example of its application to electronic devices, optionally, the voice interaction method provided in this application embodiment can also be applied to a communication system. The communication system includes N electronic devices, where N is an integer greater than or equal to 2. Among the N electronic devices, there is a central device with voice interaction functionality. The central device is connected to all other devices among the N electronic devices except the central device, and can control the other devices through voice interaction. Taking a smart home scenario as an example, the central device can be, for example, a mobile phone, wearable device, or smart speaker. Other devices can be various home appliances, such as televisions, smart refrigerators, smart lights, smart air conditioners, smart curtains, smart projectors, etc. In short, this application embodiment does not limit the specific types of the N electronic devices. For ease of understanding, the following description uses a mobile phone as the central device.

[0134] A central device (e.g., a mobile phone) can detect wake words and enter a wake-up state upon detection. Once in the wake-up state, the central device can capture speech and display the text converted from speech in real time, while also marking tasks within the text in real time. For example, as shown in Figure 10A, at time T1, the central device outputs the text: "Use the TV." At this time, the text is displayed without distinction. At time T2, the central device continues to output the text: "Play program X." At this time, "TV playing program X" is marked. At time T3, the central device executes the task and displays the execution progress. For example, if the TV is not yet turned on, the central device can display the message "TV is turning on." If the TV is already turned on, the central device can display the message "Opening program X."

[0135] Optionally, a task may include at least one of the following: task type, task content, task object, required application, and required device. Continuing with the task "Play program X on a television," it includes the following components: task type: play multimedia; task content: program X; required device: television set. In some embodiments, the five parts—task type, task content, task object, required application, and required device—may correspond to the same or different labeling methods.

[0136] In some embodiments, when the task includes the name of the required device (e.g., television), the electronic device can replace the name with the corresponding icon. For example, the device name can be replaced with the corresponding device icon. For instance, as shown in Figure 10B, at time T1, the central device outputs the text: "Use a television." At this time, the text is displayed without distinction. At time T2, the central device continues to output the text: "Play program X." Here, "television" is replaced with an icon, "play" is underlined, and "program X" is circled. This is because, before time T2, the electronic device recognized the task type ("play"), the task content ("program X"), and the required device ("television"), and marked the task type and task content differently, replacing the name of the required device with the corresponding icon. At time T3, the central device executes the task and displays the execution progress.

[0137] In the above embodiment, the user's voice indicated the desired device, namely, the television. It should be understood that in practical applications, the user's voice might be "play program X," meaning the desired device isn't explicitly specified. One possible scenario is that a user is at home, and their home includes a television. The user wants to use a central device (e.g., a mobile phone) to turn on the television and play program X. In this case, the user utters the voice command "play program X." After the central device captures this voice, it can display the text converted from the voice in real time and mark the task within the text. Considering that the central device also has the function of playing program X, one possible approach is that the central device can determine whether the task should be performed by itself or by another device. If it should be performed by another device, the central device will send the task to that device for execution.

[0138] For example, please refer to Figure 10C, which is a flowchart illustrating a voice interaction method provided in an embodiment of this application. As shown in Figure 10C, the process includes:

[0139] S11, the central equipment collected the voice message "Play program X".

[0140] S12, the central device determines whether there is a device (e.g., a television) specifically designed to perform this task in the vicinity. If it exists, proceed to S13; otherwise, proceed to S15.

[0141] S13, the central device determines whether the television meets the conditions. If it does, proceed to S14; otherwise, proceed to S15.

[0142] Optionally, the conditions may include at least one of the following: being powered on, having been used more than a preset number of times, or having a distance from the central device that is less than a preset distance.

[0143] S14, the central device sends a task to the television set to make the television set play program X.

[0144] Optionally, before S14, the central device can output a prompt message to indicate whether to use the television to play program X. For example, the prompt message can be the text "Use the television to play interface X?". Optionally, when the central device outputs the prompt message, it can mark the task characteristics in the prompt message (see the previous description for task characteristics), such as replacing "television" with an icon, using a certain font color for "play", and using italics, bold, or other fonts for "program X". When the central device receives the instruction to confirm using the television to play the program, it sends the task to the television.

[0145] S15, the central device performs the task locally, namely, playing program X.

[0146] Figure 11 is a schematic diagram of a voice interaction method provided in an embodiment of this application. As shown in Figure 11, the electronic device includes a voice acquisition unit, a voice recognition unit (e.g., an ASR unit), a display unit, and a rich text streaming processing unit (also referred to as a task recognition unit or other names). The voice acquisition unit can be a microphone, which can be a microphone array, used to acquire voice signals. The ASR unit can be integrated into a processor (e.g., CPU, NPU, etc.) to convert voice into a text stream. Optionally, the ASR unit can also be located on the cloud side; in other words, the electronic device can send the voice signal acquired by the voice acquisition unit to the cloud side, where the ASR unit converts the voice into a text stream and then sends the text stream to the electronic device. The display unit can be a touch screen for displaying the text stream. The rich text streaming processing unit is used to mark tasks in the text stream. Optionally, the rich text streaming processing unit can be a function of an application in the electronic device, which can be a third-party application or a system application. Optionally, the rich text streaming processing unit can be integrated into a processor (e.g., CPU, NPU, etc.).

[0147] It should be noted that for the traditional voice interaction process (for example, the interaction process in Figure 2 mentioned above), the electronic device does not include a rich text streaming processing unit, and its display principle is as follows: (1), The user issues voice. For example, the voice is: Send a message to Caicai on WeChat and say that I can't attend the banquet today. (2), The voice collection unit is used to collect the real-time voice stream. The real-time voice stream is transmitted to the ASR unit. (3), The ASR unit is used to perform speech recognition on the real-time voice stream to obtain a text stream. (4), The display unit is used to display the text stream, but the text stream is displayed without distinction, and it is difficult for the user to quickly notice whether the electronic device's parsing of the task is correct.

[0148] In the task marking method provided by the embodiments of the present application, since a rich text streaming processing unit is added to the electronic device, the voice interaction process of the electronic device can include the following:

[0149] (1), The user issues voice. For example, the voice is: Send a message to Caicai on WeChat and say that I can't attend the banquet today.

[0150] (2), The voice collection unit is used to collect the real-time voice stream. The real-time voice stream is transmitted to the ASR unit.

[0151] (3), The ASR unit is used to perform speech recognition on the real-time voice stream to obtain a text stream. The text stream is synchronously input to the display unit and the rich text streaming processing unit.

[0152] It should be noted that (1)-(3) can be carried out synchronously. For example, during the user's speech, the ASR unit has already started to convert the voice into a text stream.

[0153] (4), The display unit is used to display the text stream. The display unit can display one character at a time, and moreover, at the beginning (for example, the first 3 characters or the first 5 characters, etc.), the text stream can be displayed without distinction. Displaying without distinction means that the font color, font style, background color, etc. of the text are all the same. For example, in Figure 11, the first four characters "Send a message" are displayed without distinction.

[0154] (5), The rich text streaming processing unit is used to identify the task and task characteristics in the text stream, for example, identifying the task as "Send a message to Caicai on WeChat and I can't attend the banquet today". After identifying the task, the rich text streaming processing unit can mark the task in the text stream.

[0155] It should be understood that the above steps (4) and (5) can be executed synchronously. In other words, during the process of the display unit displaying the text stream, the rich text streaming processing unit performs task identification. Therefore, when the task is identified, a part of the text stream may already be displayed.

[0156] For the portion of the text stream already displayed, the rich text streaming unit adjusts the display mode of the task characteristics in the displayed portion from the current display mode to the display mode corresponding to the task characteristics. For the portion of the text stream not yet displayed, the rich text streaming unit can control the display unit to display the task characteristics in the undisplayed portion according to the corresponding display mode, while other content in the undisplayed portion continues to use the original display mode (i.e., the display mode when displaying without distinction).

[0157] For example, in Figure 11, at time T1, the display unit shows "Send via WeChat," which is displayed without distinction. Assuming that before time T2, the rich text streaming unit has identified the task characteristic as "WeChat," then at time T2, the display method of "WeChat" in the already displayed portion ("Send via WeChat") is adjusted, for example, by replacing it with the WeChat application icon. At time T3, the undisplayed portion includes "Cannot attend the banquet." Since the rich text streaming unit has determined that "Cannot attend the banquet" is a task characteristic, it can control the display unit to display "Cannot attend the banquet" according to the corresponding display method (i.e., bold display). In other words, the text stream of "Cannot attend the banquet" is already bold when it appears, and does not need to be adjusted from non-bold to bold.

[0158] It should be noted that in the embodiments described above, the example given is that a user's voice is captured by an electronic device, the electronic device displays a text stream, and the tasks in the text stream are marked in real time. Optionally, the user can also input the text stream through an input device (e.g., a physical keyboard or a virtual keyboard). In this case, the electronic device can also display the text stream and mark the tasks in the text stream in real time. The marking principle is the same as the marking principle described above, and will not be repeated here.

[0159] Figure 12 is a schematic flowchart of a task marking method provided in an embodiment of this application. This task marking method can be applied to the scenarios shown in Figures 3 to 11. As shown in Figure 12, the process includes:

[0160] S101, Obtain the first text stream.

[0161] Optionally, in one implementation of S101, the user emits voice, the electronic device acquires a first voice stream, and then obtains a first text stream based on the first voice stream. In another implementation of S101, the user inputs text through an input device (e.g., a physical keyboard or a virtual keyboard), and the electronic device obtains the first text stream through the input device.

[0162] S102, display a first text stream, the first text stream includes a first task, the first task includes N task characteristics, N is a positive integer, the N task characteristics include at least one of task type, task content, task object, required application and required device; the N task characteristics in the first text stream are displayed using different display methods, and the display methods of the N task characteristics are different from the display methods of other content in the first text stream except for the N task characteristics.

[0163] For details regarding task types, task content, task objects, required applications, and required devices, please refer to the previous descriptions; they will not be repeated here.

[0164] For example, as shown in Figure 5A above, different task characteristics (e.g., "WeChat", "Send Message", "Cai Cai, I can't attend the banquet today") use different display methods. Optionally, the names in the first text stream can be replaced with the corresponding avatars. Optionally, the application names in the first text stream can be replaced with the corresponding application icons, as shown in Figure 5B. Optionally, the device types in the first text stream can be replaced with the corresponding device icons, as shown in Figure 10B.

[0165] In some embodiments, when the first text stream further includes a second task, the first task and the second task can be displayed in different ways. For example, the second task includes M task characteristics, where M is a positive integer, and the display method corresponding to the M task characteristics is different from the display method corresponding to the N task characteristics of the first task. For example, as shown in Figure 7A, the first task is "Send a message to Cai Cai on WeChat saying that I can't attend the banquet today," and the second task is "Send a message to Li Si saying that I can attend the meeting tonight." These two tasks use different display methods.

[0166] In one embodiment, the first text stream further includes a third task, which is used to cancel the first task. One possible scenario is that if the first task is not completed, its execution is canceled; a first marker is added to the first text stream to indicate that the first task has been canceled. For example, as shown in Figure 8A, after the electronic device cancels the task, a strikethrough is used to indicate that the task has been canceled. Another possible scenario is that the first task has been completed. In this case, the electronic device can output a first prompt message or execute a fourth task. The first prompt message indicates that the first task has been completed, and the fourth task is used to adjust the executed first task. For example, as shown in Figure 8B, the first task is "hang up the phone," and the second task is "answer the phone." Since the first task has been completed, the electronic device can execute the fourth task, namely, "initiate a call." As another example, the first task could also be "send a message to Cai Cai saying that I can't attend the banquet today." If the first task has been completed, the fourth task could be to retract the message. There is a possibility that the executed first task cannot be adjusted; for example, the message cannot be retracted. In this case, the electronic device can output a second prompt message to indicate that the first task cannot be adjusted.

[0167] In some embodiments, the first text stream further includes a fifth task, which comprises K task features, where K is a positive integer. Before displaying the first text stream, the electronic device can merge the first task and the fifth task. For example, it is determined that there are identical task features among the N and K task features, where the identical task features include the same task content; the first task and the third task are merged so that the identical task feature appears only once in the first text stream. For example, in Figure 9, the first task is "set the living room air conditioner to 25 degrees," and the fifth task is "set the bedroom air conditioner as well." These two tasks have the same task characteristics, namely, the task object "air conditioner" and the task content "25 degrees." Therefore, these two tasks can be merged, for example, into "set the living room and bedroom air conditioners to 25 degrees."

[0168] In some embodiments, the first text stream includes an application executing the first task, and the application is the first application. In this case, if the electronic device determines that the first application is not included, it can use a second application in the electronic device to execute the first task. Optionally, before using the second application to execute the first task, the electronic device can output a third prompt message to indicate that the first application is not installed and whether to use the second application to execute the first task. When the electronic device receives an instruction confirming the use of the second application to execute the first task, it uses the second application to execute the first task. In other embodiments, the first text stream does not include an application executing the first task. In this case, the electronic device can identify the first application, and the first application may have historically executed other tasks with the same task characteristics as the first task; output a fourth prompt message to indicate whether to use the first application to execute the first task; and when it receives an instruction confirming the use of the first application to execute the first task, it uses the first application to execute the first task.

[0169] In some embodiments, the first text stream includes a device for executing the first task, and the executing device is the first device. In this case, the electronic device sends the first task to the first device, and the first device executes the first task. In other embodiments, the first text stream does not include a device for executing the first task. In this case, as shown in FIG10C, when the electronic device detects the presence of other devices (e.g., a television) capable of executing the first task in the vicinity, it sends the first task to the other devices; when it detects that no other devices capable of executing the first task are present in the vicinity, the electronic device executes the first task.

[0170] In some embodiments, the first text stream may include implicit tasks. Therefore, the electronic device can determine a sixth task related to the first task based on the contextual semantics of the first text stream; and display the text of the sixth task while displaying the first text stream, and then execute the sixth task. For example, as shown in Figure 7D above, after receiving a call from Cai Cai, the electronic device captures the user's voice message, "Reply to her directly, I'm driving now, I'll call her later." Based on the contextual semantics of the text stream, the electronic device determines that the associated task is to hang up the phone and executes the task of hanging up the phone. In this way, although the user does not explicitly give a task, the electronic device can infer the task and execute it, which is relatively intelligent and provides a better user experience.

[0171] In some embodiments, the electronic device can also modify the first task. For example, the electronic device acquires a second text stream; when it determines that the second text stream is used to modify the first task, it modifies the first task according to the second text stream, and then executes the modified first task. For example, as shown in FIG6E, the electronic device modifies the first task "play Feifei's songs" to "play Feifei's songs". Optionally, the electronic device can also display a modification mark for the first task, such as a strikethrough in FIG6E(c).

[0172] In some embodiments, to ensure task accuracy, the electronic device may output a fifth prompt message before executing the first task, prompting the user to confirm the first task. The task characteristics of the first task are highlighted in the fifth prompt message. For example, in Figure 6D(c), the prompt message is: "Change Feifei to Feifei?" Here, "Feifei" and "Feifei" are highlighted. As another example, if the first task is "Dial Caicai's phone number," the electronic device may output a prompt message asking "Dial Caicai's phone number?", in which the task is highlighted. After the electronic device outputs the fifth prompt message, if a confirmation command is received, the first task is executed.

[0173] In some embodiments, some tasks may be incomplete. For example, a first task may lack a first task characteristic. The first task characteristic includes at least one of the following: task type, task content, task object, required application, and required device. In this case, the electronic device can determine the first task characteristic based on the contextual semantics of the first text stream and add the first task characteristic to the first text stream. For example, a user sends the voice message "Send to Cai Cai, saying I'm not going to the banquet today, and also send to Zhao Si." Here, the task of "Send to Zhao Si" is incomplete, so the first task characteristic (i.e., task content) is determined. The electronic device can infer from the context that the information content sent to Zhao Si is "I'm not going to the banquet today," and add "I can't go to the banquet" to the first text stream. Therefore, the final first text stream is "Send to Cai Cai, saying I'm not going to the banquet today, and also send to Zhao Si, I can't go to the banquet today."

[0174] S103, execute the first task.

[0175] In some embodiments, the electronic device executes the first task when it receives a user's confirmation to perform the first task. Alternatively, the electronic device starts a countdown and executes the first task when the countdown reaches 0, wherein the countdown duration is a preset duration. Optionally, an attention mechanism can be introduced during the countdown. For example, after starting the countdown, if the electronic device detects that the user's gaze is focused on the electronic device's display screen, the countdown is paused; if the electronic device detects that the user's gaze leaves the electronic device's display screen, the countdown continues.

[0176] Optionally, S102 and S103 can be executed synchronously, that is, the first task is executed while the first text stream is being displayed.

[0177] Please refer to Figure 13, which is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device can be the electronic device mentioned above, such as a mobile phone. As shown in Figure 13, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0178] Processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. The controller may serve as the nerve center and command center of the electronic device. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. Processor 110 may also include memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that processor 110 has just used or is repeatedly used. If processor 110 needs to reuse the instruction or data, it can directly retrieve it from the memory. This avoids repeated access, reduces the waiting time of processor 110, and thus improves system efficiency.

[0179] In some embodiments, the processor 110 may execute the task marking method provided in the embodiments of this application.

[0180] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0181] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.

[0182] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0183] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0184] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.

[0185] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the electronic device 100 to capture images. The processor 110 and the display screen 194 communicate via the DSI interface to enable the electronic device 100 to display images.

[0186] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0187] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0188] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0189] The wireless communication function of the electronic device can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in conjunction with a tuning switch.

[0190] The mobile communication module 150 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G in electronic devices. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0191] The wireless communication module 160 can provide solutions for wireless communication applications in electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0192] In some embodiments, antenna 1 of the electronic device is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling the electronic device to communicate with networks and other devices via wireless communication technology.

[0193] The display screen 194 is used to display the application's interface, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device may include one or N display screens 194, where N is a positive integer greater than 1.

[0194] The electronic device 100 can perform shooting functions through an ISP, a camera 193, a video codec, a GPU, a display 194, and an application processor. The ISP is used to process the data fed back by the camera 193.

[0195] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system and software code for at least one application program. The data storage area may store data generated during the use of the electronic device (e.g., images, videos, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, general-purpose flash memory, etc.

[0196] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, images, videos, and other files can be saved on the external memory card.

[0197] Electronic devices can implement audio functions such as music playback and recording through audio modules 170, speakers 170A, receivers 170B, microphones 170C, headphone jacks 170D, and application processors.

[0198] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0199] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls and other external playback scenarios through one or more speakers 170A.

[0200] The receiver 170B, also known as a "handpiece," can be one or more, and is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 170B can be brought close to the ear to listen to the voice.

[0201] The microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals.

[0202] The 170D headphone jack is used to connect wired headphones.

[0203] The pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 180A may be disposed on the display screen 194.

[0204] The gyroscope sensor 180B can be used to determine the motion attitude of an electronic device. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device about three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization.

[0205] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0206] The magnetic sensor 180D includes a Hall effect sensor. Electronic devices can use the magnetic sensor 180D to detect the opening and closing of a flip cover.

[0207] The 180E accelerometer can detect the magnitude of acceleration in various directions (typically three axes) of electronic devices. When the electronic device is stationary, it can detect the magnitude and direction of gravity.

[0208] The 180F distance sensor is used to measure distance. Electronic devices can measure distance using infrared or laser.

[0209] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device emits infrared light outward through the LED. The electronic device uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that an object is near the electronic device. When insufficient reflected light is detected, the electronic device can determine that no object is near the electronic device.

[0210] An ambient light sensor 180L is used to detect ambient light levels. Electronic devices can adaptively adjust the brightness of the display screen 194 based on the detected ambient light levels.

[0211] The fingerprint sensor 180H is used to collect fingerprints.

[0212] The 180J temperature sensor is used to detect temperature.

[0213] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K is used to detect touch operations applied to or near it. The touch sensor can then transmit the detected touch operation to the application processor to determine the type of touch event.

[0214] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals from the vibrating bone segments of the human vocal cords.

[0215] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch buttons. The electronic device can receive button inputs and generate key signal inputs related to user settings and function control. Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the electronic device.

[0216] It is understood that the components shown in Figure 13 do not constitute a specific limitation on the electronic device. The electronic device in embodiments of the present invention may include more or fewer components than those shown in Figure 13. Furthermore, the combination / connection relationships between the components in Figure 13 can also be adjusted and modified.

[0217] Figure 14 is a schematic diagram of the structure of an electronic device 1400 provided in an embodiment of this application. The electronic device 1400 can be one of the aforementioned electronic devices (e.g., a mobile phone). As shown in Figure 14, the electronic device 1400 may include: one or more processors 1401; one or more memories 1402; a communication interface 1403; and one or more computer programs 1404. These devices can be connected via one or more communication buses 1405. The one or more computer programs 1404 are stored in the aforementioned memories 1402 and configured to be executed by the one or more processors 1401. The one or more computer programs 1404 include instructions. For example, when the electronic device 1400 is one of the aforementioned electronic devices (e.g., a mobile phone), the instructions can be used to perform relevant steps of the electronic device as described in the corresponding embodiments above, such as performing the relevant steps of the electronic device in Figures 3 to 12. The communication interface 1403 is used to enable communication between the electronic device 1400 and other devices; for example, the communication interface can be a transceiver.

[0218] In the embodiments provided above, the methods provided by the embodiments of this application are described from the perspective of an electronic device (e.g., a mobile phone) as the executing entity. To implement the functions of the methods provided in the embodiments of this application, the electronic device may include hardware structures and / or software modules, implementing the above functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular function is implemented in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.

[0219] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)). Where there is no conflict, the solutions in the above embodiments can be combined.

[0220] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0221] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0222] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0223] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0224] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope and intent of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations.

Claims

1. A task marking method, characterized in that, Applied to electronic devices, the method includes: Obtain the first text stream; The first text stream is displayed, which includes a first task. The first task includes N task characteristics, where N is a positive integer. The N task characteristics include at least one of task type, task content, task object, required application, and required device. The N task characteristics in the first text stream are displayed using different display methods, and the display methods of the N task characteristics are different from the display methods of other content in the first text stream besides the N task characteristics. Perform the first task.

2. The method according to claim 1, characterized in that, When the first text stream contains a person's name, the person's name is replaced with the corresponding avatar; When the first text stream includes an application name, the application name is replaced with the corresponding application icon; When the first text stream includes a device type, the device type is replaced with the corresponding device icon.

3. The method according to claim 1 or 2, characterized in that, The first text stream also includes a second task, which includes M task characteristics, where M is a positive integer. The display method corresponding to the M task characteristics is different from the display method corresponding to the N task characteristics. The method further includes: executing the second task.

4. The method according to any one of claims 1-3, characterized in that, The first text stream further includes a third task, which is used to cancel the first task. The method further includes: If the first task is not completed, the execution of the first task will be cancelled; Add a first marker to the first text stream, the first marker being used to indicate that the first task has been cancelled.

5. The method according to any one of claims 1-3, characterized in that, The first text stream further includes a third task, which is used to cancel the first task. The method further includes: If the first task has been completed, output a first prompt message or execute a fourth task. The first prompt message is used to indicate that the first task has been completed, and the fourth task is used to adjust the first task that has been executed.

6. The method according to claim 5, characterized in that, The method further includes: When it is determined that the first task cannot be adjusted, a second prompt message is output, which indicates that the first task cannot be adjusted.

7. The method according to any one of claims 1-6, characterized in that, The first text stream further includes a fifth task, which includes K task features, where K is a positive integer. Before displaying the first text stream, the method further includes: It is determined that there are common task features among the N task features and the K task features, and the common task features include the same task content; The first task and the third task are merged so that the same task feature appears only once in the first text stream.

8. The method according to any one of claims 1-7, characterized in that, The execution of the first task includes: Upon receiving a user's confirmation to execute the first task, execute the first task; or, A countdown is started, and when the countdown reaches 0, the first task is executed, wherein the countdown duration is a preset duration.

9. The method according to claim 8, characterized in that, After starting the countdown, the method further includes: The countdown is paused when the user's gaze is detected on the display screen of the electronic device; The countdown continues when the user's gaze leaves the display screen of the electronic device.

10. The method according to any one of claims 1-9, characterized in that, The first text stream includes the application that executes the first task, and the application is the first application. Executing the first task includes: If it is determined that the electronic device does not contain the first application, the first task is performed using the second application in the electronic device.

11. The method according to claim 10, characterized in that, Before performing the first task using the second application in the electronic device, the method further includes: Output a third prompt message, which is used to prompt whether the first application is not installed and whether to use the second application to perform the first task. Received an instruction confirming the use of the second application to execute the first task.

12. The method according to any one of claims 1-9, characterized in that, The first text stream does not include the application executing the first task, and executing the first task includes: Identify a first application in the electronic device, the first application having historically performed other tasks with the same task characteristics as the first task; Output a fourth prompt message, which is used to prompt whether to use the first application to perform the first task; Receive an instruction confirming the use of the first application to perform the first task.

13. The method according to any one of claims 1-12, characterized in that, The first text stream includes an execution device for the first task, and the device is a first device. Executing the first task includes: Send the first task to the first device, and execute the first task through the first device.

14. The method according to any one of claims 1-12, characterized in that, The first text stream does not include the execution device for the first task, and the execution of the first task includes: When the presence of other devices capable of performing the first task is detected in the vicinity of the electronic device, the electronic device sends the first task to the other devices; When it is detected that there are no other devices around the electronic device capable of performing the first task, the electronic device performs the first task.

15. The method according to any one of claims 1-14, characterized in that, The execution of the first task includes: While displaying the first text stream, the first task is performed.

16. The method according to any one of claims 1-15, characterized in that, The method further includes: While displaying the first text stream, a sixth task is also displayed, which is a task that is related to the first task and is determined based on the contextual semantics of the first text stream. Perform the sixth task.

17. The method according to any one of claims 1-16, characterized in that, Before performing the first task, the method further includes: Obtain the second text stream; When it is determined that the second text stream is used to modify the first task, the first task is modified according to the second text stream; Performing the first task includes: performing the modified first task.

18. The method according to claim 17, characterized in that, The method further includes: Display the modification markers for the first task.

19. The method according to any one of claims 1-18, characterized in that, The display method for each of the N task features is different; wherein, the different display method for each of the N task features includes: at least one of the following: font color, font style, and font background color for each of the N task features is different.

20. The method according to any one of claims 1-19, characterized in that, The method further includes: A fifth prompt message is output, which prompts the user to confirm the first task. The task characteristics of the first task are marked in the fifth prompt message. A confirmation instruction is received, which is used to confirm the first task.

21. The method according to any one of claims 1-19, characterized in that, Before performing the first task, the method further includes: When the first task lacks the first task characteristic, the first task characteristic is determined based on the contextual semantics of the first text stream; The first task characteristic is added to the first text stream. The first task characteristic includes at least one of the following: task type, task content, task object, required application, and required device.

22. The method according to any one of claims 1-21, characterized in that, The display of the first text stream includes: The first text stream is synchronously transmitted to the display unit and the task identification unit. The display unit displays the first text stream; the task identification unit identifies the first task in the first text stream and the task characteristics of the first task. After identifying the task characteristics of the first task, For the portion of the first text stream that has already been displayed, the task characteristics of the first task contained in the displayed portion are adjusted from the current first display mode to the second display mode corresponding to the task characteristics; For the undisplayed portion of the first text stream, the task characteristics of the first task contained in the undisplayed portion are displayed according to the third display method corresponding to the task characteristics, and the other content in the undisplayed portion continues to be displayed according to the first display method.

23. A voice interaction method, characterized in that, Applied to electronic devices, the method includes: Obtain the first audio stream; Display a first text stream, which is obtained from the first speech stream, and the first text stream includes a first task; Start the timer; When the system detects that the user's gaze is focused on the display screen of the electronic device, it controls the timer to pause. When the system detects that the user's gaze has left the display screen of the electronic device, it controls the timer to continue counting. When the timer reaches the preset time, the first task is executed.

24. The method according to claim 23, characterized in that, The method further includes: Display the first button; Upon receiving an operation on the first button, the first task is executed immediately.

25. The method according to claim 24, characterized in that, The method further includes: The timer's progress is displayed within the display area where the first button is located.

26. An electronic device, characterized in that, include: Processor, memory, and one or more programs; The one or more programs are stored in the memory, and the one or more programs include instructions that, when executed by the processor, cause the electronic device to perform the steps of the method as described in any one of claims 1-25.

27. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 25.

28. A computer program product, characterized in that, Includes a computer program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 25.

Citation Information

Patent Citations

  • Multi-command single utterance input method

    CN106471570A

  • Display method, device and terminal for speech input control instruction

    CN107122160A

  • Voice control text display method and device

    CN107155121A

  • A voice information interaction method and an intelligent electric appliance

    CN109215645A

  • Task establishing method and mobile terminal

    CN110223695A