Interface interaction method and apparatus, electronic device, and storage medium
By combining voice commands and touch operations, and utilizing large language models and intent recognition technology, the problem of inaccurate content selection by users on electronic device display interfaces has been solved, achieving higher operational accuracy.
Patent Information
- Application Number
- PCT/CN2025/117625
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-30
- Filing Date
- 2025-08-28
- Publication Date
- 2026-03-05
AI Technical Summary
In existing technologies, when users select content on the display interface of electronic devices through touch operations, it is difficult for them to accurately indicate points of interest and ranges, causing the electronic devices to fail to accurately execute the user's operation requirements.
By combining voice commands and touch operations, and utilizing large language models and intent recognition technology, the target interface content is determined from the display interface of the electronic device, and the corresponding operation is executed.
It improves the accuracy of electronic devices in meeting user operation needs, ensuring accurate selection of interface content and execution of operations.
Smart Images

Figure CN2025117625_05032026_PF_FP_ABST
Abstract
Description
Interface interaction methods, devices, electronic devices and storage media
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese application No. 202411216939.1, filed on August 30, 2024, the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0003] This application relates to the field of electronic device technology, and more specifically, to an interface interaction method, apparatus, electronic device, and storage medium. Background Technology
[0004] With the rapid advancement of technology and living standards, electronic devices (such as smartphones and tablets) have become one of the most commonly used electronic products in people's lives. Currently, when using electronic devices, people typically perform operations such as searching, saving, translating, copying, and sharing content displayed on the screen. Summary of the Invention
[0005] This application proposes a user interface interaction method, apparatus, electronic device, and storage medium.
[0006] In a first aspect, embodiments of this application provide an interface interaction method, the method comprising: displaying a first interface; responding to a detected voice command and a touch operation in the first interface, performing an operation corresponding to the voice command based on target interface content in the first interface, wherein the target interface content is determined from the first interface based on the voice command and the touch operation.
[0007] Secondly, embodiments of this application provide an interface interaction device, the device comprising: an interface display module and an operation execution module, wherein the interface display module is used to display a first interface; the recognition output module is used to respond to a detected voice command and a touch operation in the first interface, and execute the operation corresponding to the voice command based on the target interface content in the first interface, wherein the target interface content is determined from the first interface based on the voice command and the touch operation.
[0008] Thirdly, embodiments of this application provide an electronic device, including: one or more processors; a memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to perform the interface interaction method provided in the first aspect above.
[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing program code, which can be invoked by a processor to execute the interface interaction method provided in the first aspect above. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 shows a flowchart of an interface interaction method according to an embodiment of this application.
[0012] Figure 2 shows a schematic diagram of an interface provided in an embodiment of this application.
[0013] Figure 3 shows a schematic diagram of an interface provided in an embodiment of this application.
[0014] Figure 4 shows a flowchart of an interface interaction method according to another embodiment of this application.
[0015] Figure 5 shows a schematic diagram of an interface provided in an embodiment of this application.
[0016] Figure 6 shows a schematic diagram of an interface provided in an embodiment of this application.
[0017] Figure 7 shows a flowchart of an interface interaction method according to yet another embodiment of this application.
[0018] Figure 8 shows a schematic diagram of an interface provided in an embodiment of this application.
[0019] Figure 9 shows a schematic diagram of an interface provided in an embodiment of this application.
[0020] Figure 10 shows a schematic diagram of an interface provided in an embodiment of this application.
[0021] Figure 11 shows a schematic diagram of an interface provided in an embodiment of this application.
[0022] Figure 12 shows a schematic diagram of an interface provided in an embodiment of this application.
[0023] Figure 13 shows a schematic diagram of an interface provided in an embodiment of this application.
[0024] Figure 14 shows a schematic diagram of an interface provided in an embodiment of this application.
[0025] Figure 15 shows a schematic diagram of an interface provided in an embodiment of this application.
[0026] Figure 16 shows a schematic diagram of an interface provided in an embodiment of this application.
[0027] Figure 17 shows a schematic diagram of an interface provided in an embodiment of this application.
[0028] Figure 18 shows a schematic diagram of an interface provided in an embodiment of this application.
[0029] Figure 19 shows a schematic diagram of an interface provided in an embodiment of this application.
[0030] Figure 20 shows a schematic diagram of an interface provided in an embodiment of this application.
[0031] Figure 21 shows a flowchart of an interface interaction method according to another embodiment of this application.
[0032] Figure 22 shows a flowchart of an interface interaction method according to another embodiment of this application.
[0033] Figure 23 shows a flowchart of an interface interaction method according to yet another embodiment of this application.
[0034] Figure 24 shows a block diagram of an interface interaction device according to an embodiment of the present application.
[0035] Figure 25 is a block diagram of an electronic device for performing an interface interaction method according to an embodiment of the present application.
[0036] Figure 26 is a storage unit according to an embodiment of the present application for storing or carrying program code that implements the interface interaction method according to an embodiment of the present application. Detailed Implementation
[0037] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0038] Currently, when using electronic devices, people often need to perform operations such as searching, saving, translating, and copying content on the display interface. In related technologies, users can indicate the location and range of points of interest (POIs) by performing touch operations on the screen (e.g., drawing circles, lines, scribbling, clicking, etc.). The electronic device then extracts metadata such as images and text from the determined range as the POI, i.e., the interface content to be operated on, and then executes the corresponding operation indicated by the user's voice command. However, this method requires high precision in the user's touch operation. When selecting interface content through touch operations, the user may not accurately indicate the location and range of the POI, resulting in the electronic device determining that the POI is not the user's desired POI, and consequently, failing to accurately execute the user's requested operation.
[0039] To address the aforementioned problems, the inventors have proposed an interface interaction method, apparatus, electronic device, and storage medium as provided in the embodiments of this application. These methods can determine the interface content requiring operation based on user-input voice commands and touch operations on the display interface, thereby ensuring the accuracy of the operated interface content and improving the accuracy of the electronic device in executing user-required operations. The specific interface interaction method will be described in detail in subsequent embodiments.
[0040] The embodiments of this application provide an interface interaction method, the method comprising: displaying a first interface; responding to a detected voice command and a touch operation in the first interface, performing an operation corresponding to the voice command based on target interface content in the first interface, wherein the target interface content is determined from the first interface based on the voice command and the touch operation.
[0041] According to an embodiment of this application, the step of responding to a detected voice command and a touch operation in the first interface, and executing the operation corresponding to the voice command based on the target interface content in the first interface, includes: responding to a detected voice command and a touch operation in the first interface, determining the content selected by the touch operation from the first interface as the target interface content based on the intent information corresponding to the voice command and the position information of the touch operation; and executing the operation corresponding to the voice command based on the target interface content.
[0042] According to an embodiment of this application, determining the content selected by the touch operation from the first interface as the target interface content based on the intent information corresponding to the voice command and the location information of the touch operation includes: determining interface content that matches the intent information and the location information from the first interface as the target interface content.
[0043] According to an embodiment of this application, determining the interface content that matches the intent information and the location information from the first interface as the target interface content includes: performing content recognition on the interface image of the first interface to obtain a recognition result; and determining the interface content that matches the location information and the intent information from the recognition result as the target interface content.
[0044] According to an embodiment of this application, the location information includes the operation trajectory information of the touch operation. Before determining the content selected by the touch operation from the first interface based on the intent information corresponding to the voice command and the location information of the touch operation as the target interface content, the method further includes: displaying the operation trajectory information corresponding to the touch operation in the interface; after determining the content selected by the touch operation from the first interface based on the intent information corresponding to the voice command and the location information of the touch operation as the target interface content, the method further includes: updating the operation trajectory information displayed in the interface based on the target interface content, wherein the updated operation trajectory information matches the area where the target interface content is located.
[0045] According to an embodiment of this application, updating the operation trajectory information displayed on the interface based on the target interface content includes: drawing the edge of the target interface content based on the range of the target interface content in the first interface, and displaying the drawn edge on the first interface so that the displayed operation trajectory information matches the area where the target interface content is located.
[0046] According to an embodiment of this application, displaying the operation trajectory information corresponding to the touch operation in the interface includes: displaying the operation trajectory information in the interface with a target color, wherein the target color is different from the color of other content in the interface.
[0047] According to an embodiment of this application, before executing the operation corresponding to the voice command based on the target interface content, the method further includes: displaying a first prompt message in the interface, the first prompt message being used to indicate that the target interface content is being determined from the interface.
[0048] According to an embodiment of this application, determining the content selected by the touch operation from the first interface as the target interface content based on the intent information corresponding to the voice command and the location information of the touch operation includes: if the interface content matching the intent information and the location information is determined from the first interface to include multiple contents, displaying candidate information corresponding to each of the multiple contents in the interface; if a selection operation for the target candidate information is detected, determining the interface content corresponding to the target candidate information as the target interface content.
[0049] According to an embodiment of this application, the step of responding to a detected voice command and a touch operation in the first interface, and executing the operation corresponding to the voice command based on the target interface content in the first interface, includes: responding to a detected voice command and a touch operation in the first interface, using a large language model, and based on the text corresponding to the voice command, touch position information, and the interface image of the first interface, determining the target interface content; and executing the operation corresponding to the voice command based on the target interface content.
[0050] According to an embodiment of this application, the step of using a large language model and determining the target interface content based on the text corresponding to the voice command, touch position information, and the interface image of the first interface includes: inputting the text corresponding to the voice command, touch position information, the interface image of the first interface, application information corresponding to the first interface, and the result of content segmentation of the interface image into the large language model to obtain the target interface content output by the large language model.
[0051] According to an embodiment of this application, the step of responding to a detected voice command and a touch operation in the first interface, and executing the operation corresponding to the voice command based on the target interface content in the first interface, includes: responding to the voice command and the touch operation, extracting text information from the target interface content, and executing the operation corresponding to the voice command on the text information; or
[0052] In response to the voice command and the touch operation, a region image of the area where the target interface content is located is obtained, and the operation corresponding to the voice command is executed on the region image.
[0053] According to an embodiment of this application, the step of responding to a detected voice command and a touch operation in the first interface, and executing the operation corresponding to the voice command based on the target interface content in the first interface, includes: responding to the voice command and the touch operation, executing the operation corresponding to the voice command on a target object corresponding to the target interface content, wherein the target object is a file or an application.
[0054] According to an embodiment of this application, in response to a detected voice command and a touch operation in the first interface, an operation corresponding to the voice command is executed based on the target interface content in the first interface, including: in response to the simultaneous detection of the voice command and the touch operation, after detecting that the input of the voice command has stopped and the input of the touch operation has stopped, the operation corresponding to the voice command is executed based on the target interface content in the first interface.
[0055] According to an embodiment of this application, the method further includes: displaying a second prompt message on the interface, the second prompt message being used to indicate that an operation corresponding to the voice command is being executed.
[0056] According to an embodiment of this application, the application corresponding to the first interface is a first application. In response to a detected voice command and a touch operation in the first interface, based on the target interface content in the first interface, the operation corresponding to the voice command is executed, including: in response to the voice command and the touch operation, calling a second application, and through the second application, executing the operation corresponding to the voice command based on the target interface content in the first interface; the method further includes: displaying a second interface corresponding to the second application, wherein the second interface is an interface displayed by the second application for the operation result corresponding to the operation.
[0057] According to embodiments of this application, the touch operation includes one or a combination of several of the following: selection operation, line drawing operation, smearing operation, and click operation.
[0058] The interface interaction method provided in the embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0059] Please refer to Figure 1, which shows a flowchart of an interface interaction method provided in one embodiment of this application. In a specific embodiment, the interface interaction method is applied to the interface interaction device 600 shown in Figure 24 and the electronic device 100 (Figure 25) configured with the interface interaction device 600. The specific flow of this embodiment will be described below using an electronic device as an example. Of course, it is understood that the electronic device used in this embodiment can be a smartphone, tablet computer, e-reader, laptop computer, smartwatch, etc., and is not limited here. The flow shown in Figure 1 will be described in detail below. The interface interaction method specifically includes the following steps:
[0060] Step S110: Display the first interface.
[0061] The first interface can be any interface displayed on the electronic device, such as a document interface, browser interface, video playback interface, social application interface, etc., and the specific interface is not limited. It is understood that users may have needs to search, save, translate, copy, share, and perform other operations on the content displayed on different interfaces of the electronic device. Therefore, the interface interaction method provided in this application embodiment can be executed when the electronic device displays any interface.
[0062] In this embodiment of the application, when the electronic device displays a first interface, it can detect the user's voice commands and touch operations on the interface. In order to determine the corresponding content from the first interface based on the voice commands and touch operations, and then perform corresponding operations on the determined content.
[0063] Step S120: In response to the detected voice command and the touch operation in the first interface, the operation corresponding to the voice command is executed based on the target interface content in the first interface, wherein the target interface content is determined from the first interface based on the voice command and the touch operation.
[0064] In this embodiment, when the electronic device displays a first interface, after detecting a voice command and a touch operation on the first interface, it can respond to the voice command and touch operation, determine the target interface content that the user wants the electronic device to process according to their intention based on the voice command and touch operation, and perform the operation corresponding to the voice command on the target interface content. The touch operation is used to indicate the content selected by the user on the displayed interface; it is understood that when the user indicates the content selected by touch operation on the displayed interface, the precision of the touch operation may be insufficient and cannot reflect the actual interface content that the user wants to select. Therefore, the interface content to be operated on can be determined jointly based on the input voice command and the above touch operation.
[0065] Understandably, when an electronic device displays a first interface, if the first interface contains content for which the user wants to perform a target operation, the user can input a voice command and then perform a touch operation on the desired content within the first interface. Correspondingly, the electronic device can detect the input voice command and touch operation, and based on these, retrieve the target content from the first interface that needs to be operated on.
[0066] In some implementations, the electronic device can collect audio from the environment using an audio acquisition device and identify voice commands within the collected audio. Optionally, the electronic device can detect whether the collected audio includes speech; if the collected audio includes speech, the speech in the collected audio is converted into text to obtain the recognized voice command; if the collected audio does not include speech, it indicates that the user has not input a voice command, and therefore no further processing is required. For example, if the electronic device converts the detected speech into text, and the resulting text is "Search for the price of these shoes," then the user's voice command would be "Search for the price of these shoes."
[0067] In one possible implementation, the electronic device may be equipped with a smart assistant. This smart assistant can be an application based on artificial intelligence technology. It typically provides various services and assistance to users and usually possesses capabilities such as speech recognition, natural language processing, and machine learning. It can converse with users, understand their needs, and then provide corresponding information, suggestions, or perform tasks. When the electronic device displays the first interface, upon detecting a call to the smart assistant, it can respond to the call by waking the smart assistant and displaying its interactive area on the first interface, allowing the user to input voice commands via the smart assistant.
[0068] For example, please refer to Figures 2 and 3. When the electronic device displays the first interface A1 as shown in Figure 2, after detecting the call operation to the smart assistant, it can display the interface B1 as shown in Figure 3. The interface B1 displays the interactive area A2 of the smart assistant, and after detecting the voice input by the user to the smart assistant, the text content corresponding to the input voice can be displayed in the interactive area A2.
[0069] Optionally, the electronic device can detect input voice commands and touch operations on the first interface when the smart assistant is awake and the target mode (e.g., selection mode) in the smart assistant is enabled.
[0070] In one possible implementation, the electronic device can detect user-inputted voice commands and touch operations on the first interface in real time, or it can detect user-inputted voice commands and touch operations on the first interface at preset time intervals, or it can detect user-inputted voice commands and touch operations on the first interface at preset time points, or it can detect user-inputted voice commands and touch operations on the interface according to other preset conditions, etc., without limitation.
[0071] In some implementations, touch operations include one or a combination of one or more of the following: selection, drawing, smudge, and clicking.
[0072] In one possible implementation, the touch operation may consist only of a circle drawing operation, in which case the user can indicate the content to be selected on the first interface by drawing a circle.
[0073] In one possible implementation, the touch operation may consist only of a line drawing operation, in which case the user can indicate the content to be selected in the first interface by drawing a line.
[0074] In one possible implementation, the touch operation may consist only of a swiping operation, in which case the user can indicate the content to be selected on the first interface by swiping.
[0075] In one possible implementation, the touch operation may consist only of a click operation, in which case the user can indicate the content to be selected on the first interface by clicking.
[0076] In one possible implementation, the touch operation may also include both selection and clicking operations. In this case, the user can select the content that needs to be selected on the first interface by selecting and clicking.
[0077] Of course, in the embodiments of this application, the above touch operation may also include more combinations of operations, which will not be elaborated here.
[0078] In some implementations, when the electronic device determines the target interface content to be operated from the first interface based on the above voice commands and touch operations, it can analyze and understand the content in the first interface, such as recognizing and parsing various elements such as text, images, and icons, and performing intent understanding on the above voice commands. Then, based on the results of the analysis and understanding of the first interface, the results of intent understanding on the voice commands, and the operation position and operation area of the touch operation, it determines the target interface content that matches the user's intent (i.e., the interface content that the user actually wants to select).
[0079] In some implementations, the electronic device performs operations corresponding to the voice commands on the target interface content, such as searching, saving, translating, copying, and sharing the target interface content.
[0080] In one possible implementation, considering that when the electronic device executes the operation corresponding to the above voice command, since the application corresponding to the first interface may not be able to perform the operation required by the user, the electronic device can call other applications to perform the operation corresponding to the above voice command.
[0081] The interface interaction method provided in this application embodiment can determine the interface content that needs to be operated on based on the user's voice input and the touch operation on the display interface. This achieves the effect of optimizing and supplementing the content selected by the touch operation based on the voice command, thereby ensuring the accuracy of the interface content of the operation being performed and improving the accuracy of the electronic device in performing the user's required operation.
[0082] Please refer to Figure 4, which shows a flowchart of an interface interaction method provided in another embodiment of this application. This interface interaction method is applied to the aforementioned electronic device. The flowchart shown in Figure 4 will be described in detail below. Specifically, the interface interaction method may include the following steps:
[0083] Step S210: Display the first interface.
[0084] In this embodiment, step S210 can be referred to the content of other embodiments, and will not be repeated here.
[0085] Step S220: In response to the detected voice command and the touch operation in the first interface, based on the intent information corresponding to the voice command and the position information of the touch operation, determine the content selected by the touch operation from the first interface as the target interface content.
[0086] In this embodiment, when the electronic device displays a first interface, in response to detected voice commands and touch operations, when determining the target interface content to be operated on in the first interface, it can determine the content selected by the touch operation from the first interface as the target interface content based on the intent information corresponding to the voice command and the position information of the touch operation. It is understood that when the user performs the above touch operation, it is equivalent to indicating the selected interface content to be operated on through the operation position on the screen. However, the operation position on the screen may not accurately indicate the interface content that the user actually wants to select. Therefore, the selected content can be determined from the first interface by combining the intent information corresponding to the voice command (i.e., the user's current operation intent).
[0087] In the above embodiments, the intent information can be obtained by performing intent recognition on the voice command. Intent recognition is an important task in Natural Language Processing (NLP) and speech processing. Its goal is to identify the true intent behind user input (whether text or speech). These intents can be executing a specific command (e.g., "play music"), querying certain information (e.g., "What's the weather like tomorrow?"), expressing a certain emotion or need (e.g., "I need help"), or conducting a certain transaction (e.g., "buy this product"), etc. Optionally, the intent information can be obtained by performing intent recognition on the above voice commands using a pre-trained intent recognition model.
[0088] In some implementations, when an electronic device determines the target interface content based on the intent information corresponding to a voice command and the location information of a touch operation, it can select interface content from a first interface that matches the intent information and location information as the target interface content. In other words, the determined target interface content satisfies the intent expressed by the user's voice command, and the location of the target interface content matches the location of the user's input, thus accurately determining the interface content the user intends to perform.
[0089] In one possible implementation, when the electronic device determines the interface content matching the intent information and location information from the first interface, it can perform content recognition on the interface image of the first interface to obtain a recognition result; then, it can determine the interface content matching the location information and intent information from the recognition result as the target interface content. The interface image of the first interface can be an image obtained by taking a screenshot of the first interface, and the content recognition on the interface image can be the recognition of text, images, or other elements in the interface image.
[0090] For example, if the text corresponding to the above voice command is "What is the price of this computer?", then the intent of the voice command is to "search for the price of this computer". In order to search for the price of the computer, it is necessary to know the name of the computer or a picture of the computer. Therefore, content recognition is performed on the interface image to obtain the recognition result. Then, based on the recognition result, the name of the computer or the picture of the computer in the recognition result is determined. Then, from the determined name of the computer or the picture of the computer, the name of the computer or the picture of the computer whose position matches the position information of the touch operation is determined, thereby obtaining the target interface content that needs to be determined.
[0091] Of course, in the embodiments of this application, the specific method by which the electronic device determines the target interface content from the first interface based on voice commands and touch operations is not limited. Optionally, the electronic device may also use a large language model (LLM) and determine the target interface content based on the text corresponding to the above voice commands, the position information of the above touch operations, and the above interface images.
[0092] For example, when determining the target interface content using a large language model, in addition to inputting the text corresponding to the voice command, the location information of the touch operation, and the interface image into the large language model, the application information or scene information corresponding to the first interface, as well as the results of content segmentation or recognition of the interface image of the first interface, can also be input into the large language model. This can provide more references for the large language model when determining the target interface content, and further improve the accuracy of the determined interface content that needs to be operated.
[0093] In some implementations, during the process of determining the target interface content in response to the above voice commands and touch operations, the electronic device may also display a first prompt message on the interface. This first prompt message indicates that the target interface content is being determined from the interface, thereby allowing the user to understand the processing stage of the voice commands and touch operations and preventing the user from interrupting the current processing flow by performing other operations. For example, referring to Figure 5, the electronic device may display interface B2 as shown in Figure 5. Interface B2 may include a "Understanding screen" prompt message to indicate that the interface content to be operated is being determined from the displayed interface.
[0094] In one possible implementation, when the electronic device displays the first prompt information on the first interface, it can also highlight the first prompt information so that the user can more easily notice it.
[0095] Optionally, the electronic device displays other areas of the first interface besides the first prompt information at a first display brightness, and displays the first prompt information at a second display brightness, which is greater than the first display brightness, so that the first prompt information is more prominent than other areas.
[0096] Optionally, the first prompt message is a text prompt message, and the electronic device displays the first prompt message in a first color, which may be a colored font color that is different from the other font colors on the screen.
[0097] Of course, there is no limitation on the specific way the electronic device highlights the first prompt information. For example, the electronic device can also adjust the transparency of other areas on the first interface other than the first prompt information to make the other areas more transparent than the first prompt information, thereby making the first prompt information more prominent. Or, for example, the electronic device can also adjust the blurriness of other areas on the first interface other than the first prompt information to make the other areas blurry, thereby making the first prompt information more prominent.
[0098] Step S230: Based on the target interface content, execute the operation corresponding to the voice command.
[0099] In this embodiment of the application, after the electronic device determines the above target interface content, it can perform the operation corresponding to the above voice command based on the target interface content.
[0100] In some implementations, after determining the target interface content, the electronic device, while executing the operation corresponding to the voice command based on the target interface content, can also display a second prompt message on the interface. This second prompt message indicates that the operation corresponding to the voice command is being executed, thus allowing the user to understand the processing stage of the electronic device for the voice command and touch operation, preventing the user from interrupting the current processing flow for the voice command and touch operation by performing other operations. For example, referring to Figure 6, the electronic device can display interface B3 as shown in Figure 6. Interface B3 may include a prompt message "Operating application B" to indicate that the required operation is being performed through application A.
[0101] In one possible implementation, when the electronic device displays the second prompt information on the first interface, it can also highlight the second prompt information to make the first prompt information more easily noticeable to the user. The specific method by which the electronic device highlights the second prompt information can be found in the aforementioned method for highlighting the first prompt information, and will not be repeated here.
[0102] The interface interaction method provided in this application embodiment can determine the interface content that needs to be operated on based on the user's voice input and the touch operation on the display interface. This achieves the effect of optimizing and supplementing the content selected by the touch operation based on the voice command, thereby ensuring the accuracy of the interface content of the operation being performed and improving the accuracy of the electronic device in performing the user's required operation.
[0103] Please refer to Figure 7, which shows a flowchart of an interface interaction method provided in another embodiment of this application. This interface interaction method is applied to the aforementioned electronic device. The flowchart shown in Figure 7 will be described in detail below. Specifically, the interface interaction method may include the following steps:
[0104] Step S310: Display the first interface.
[0105] In this embodiment, step S310 can be referred to the content of the foregoing embodiments, and will not be repeated here.
[0106] Step S320: In response to the detected voice command and the touch operation in the first interface, display the operation trajectory information corresponding to the touch operation in the first interface.
[0107] In this embodiment, the location information of the touch operation in the foregoing embodiments may include the operation trajectory information corresponding to the touch operation. The operation trajectory information may be the touch path or touch trajectory recorded by the operating system based on the touch position when a touch operation is performed on the touch screen. After detecting the above voice command and touch operation, the electronic device may display the operation trajectory information corresponding to the touch operation on the first interface so that the user knows the selected position and range indicated by the touch operation.
[0108] In some implementations, when displaying the above operation trajectory information, the electronic device can render the trajectory of the above touch operation based on the operation trajectory information and display the rendered trajectory on the first interface.
[0109] In one possible implementation, the electronic device can display the above operation trajectory information in a target color on a first interface, the target color being different from the colors of other content on the first interface. Optionally, the grayscale corresponding to the target color can be greater than the target grayscale, thereby making the above operation trajectory information darker and enabling the user to more clearly perceive the operation trajectory information.
[0110] Step S330: Based on the intent information corresponding to the voice command and the position information of the touch operation, determine the content selected by the touch operation from the first interface as the target interface content.
[0111] In some implementations, after the electronic device determines the interface content matching the intent information and location information from the first interface based on the intent information corresponding to the voice command and the location information of the touch operation, it considers that the determined interface content matching the intent information and location information may include multiple contents. For example, the text corresponding to the voice command input by the user is "search for the price of this piece of clothing", and the interface currently displayed by the electronic device includes pictures of multiple pieces of clothing. The location information of the user's touch operation also covers the area where the pictures of multiple pieces of clothing are located. Therefore, the interface content determined by the electronic device may include multiple contents, while the user actually needs to operate on one of the contents. Therefore, when the determined interface content matching the intent information and location information includes multiple contents, the electronic device can display candidate information corresponding to each of the multiple contents in the first interface. After displaying the candidate information, the electronic device can detect the operation in the first interface. After detecting the selection operation for the target candidate information, it can respond to the selection operation, determine the interface content corresponding to the target candidate information selected by the selection operation, and take the interface content corresponding to the target candidate information as the target interface content to be operated on this time.
[0112] In one possible implementation, when the electronic device displays candidate information corresponding to multiple items, it can display a selection area for each item in the first interface. That is, each item referred to by the entity is within a selection area, each selection area can be selected, and each selection area can serve as a candidate information. The electronic device can detect operations in the first interface to detect selection operations on the corresponding selection areas.
[0113] In one possible implementation, after displaying a selection area for each of the above-mentioned items, the electronic device may also highlight each selection area to make it easier for the user to see these selection areas and select the corresponding selection area.
[0114] Optionally, the electronic device can also control the area in the first interface other than the selected area to be unselectable, so as to prevent the user from accidentally triggering operations on other content in the first interface, thereby avoiding accidental interruption of the selection of the selected area.
[0115] In one possible implementation, when the electronic device displays a selection area for each of the above-mentioned items, it can also display a prompt message on the first interface to prompt the user to select the selection area, thereby determining the target interface content to be operated on based on the selection area selected by the user.
[0116] In one possible implementation, when the electronic device displays candidate information corresponding to multiple items, it can display a list of items on a first interface. This list includes options corresponding to each of the multiple items; that is, each option is considered a candidate. The electronic device can detect operations on the first interface to detect the selection of the corresponding option.
[0117] In one possible implementation, the above selection operation can be a click operation on the target candidate information. That is, when the electronic device detects a click operation on the target candidate information, it can determine the content corresponding to the target candidate information as the target interface content to be operated on this time.
[0118] Of course, the specific operation type for selecting the above target candidate information is not limited. For example, the above selection operation can also be a long press operation on the target candidate information. The long press operation can be a press operation with a press duration longer than the target press duration.
[0119] In one possible implementation, if the determined interface content that matches the intent information and location information includes multiple contents, the multiple contents can also be determined as the target interface content at the same time, and the operation corresponding to the voice command can be executed based on the multiple contents.
[0120] Step S340: Update the operation trajectory information displayed in the first interface based on the target interface content, and match the updated operation trajectory information with the area where the target interface content is located.
[0121] In this embodiment of the application, after the electronic device determines the target interface content from the first interface, it can know the true range of the interface content that the user wants to select. Therefore, the electronic device can also update the operation trajectory information displayed in the first interface based on the target interface content, so that the operation trajectory information displayed in the first interface matches the area where the target interface content is located. That is, the updated operation trajectory information is the true range of the interface content that the user wants to select, so that the user knows the area of the target interface content finally determined by the electronic device.
[0122] In some implementations, when updating the operation trajectory information displayed on the first interface, the electronic device can draw the edge of the target interface content based on the range of the target interface content in the first interface, and display the drawn edge on the first interface, so that the displayed operation trajectory information matches the area where the target interface content is located.
[0123] For example, please refer to Figures 3 and 6 simultaneously. The interface B1 shown in Figure 3 of the electronic device includes operation trajectory information A3. If the target interface content is determined to be the image of the "chair" in the first interface, then the smallest area containing the "chair" can be determined, and the edge of the smallest area can be drawn. Then, the interface B3 shown in Figure 6 is displayed, and the operation trajectory information A3 is updated to the edge of the smallest area.
[0124] In some implementations, after the electronic device updates the operation trajectory information in the first interface, since the user can see the range of the target interface content determined by the electronic device based on the displayed operation trajectory information, if the user believes that the range of the determined target interface content is inaccurate, the electronic device can also respond to the detected adjustment operation on the operation trajectory information, update the displayed operation trajectory information again, and redetermine the target interface content based on the updated operation trajectory information.
[0125] Step S350: Based on the target interface content, execute the operation corresponding to the voice command.
[0126] In this embodiment of the application, after the electronic device determines the above target interface content, since the content that the user wants to operate may be the target interface content itself, or it may be the object referred to by the target interface content (such as a file, application, etc.), the electronic device can also obtain the content that the user actually wants to operate based on the target interface content, and perform the operation corresponding to the above voice command on the obtained content.
[0127] In some implementations, if the intent information in the voice command is determined to be an operation on the target interface content, the electronic device can extract the text information in the target interface content and execute the operation corresponding to the voice command on the text information; the electronic device can also acquire a region image of the area where the target interface content is located and execute the operation corresponding to the voice command on the region image.
[0128] In one possible implementation, when the electronic device operates on the target interface content itself, it can determine whether to operate on the text or the image of the target interface content based on the intent information parsed from the voice command. If it is determined that to operate on the text of the target interface content, that is, if the user's point of interest in the first interface is the text of the target interface content, then the text information of the target interface content can be extracted, and the operation corresponding to the voice command can be executed on the text information. If it is determined that to operate on the image of the target interface content, that is, if the user's point of interest in the first interface is the image of the target interface content, then a region image of the area where the target interface content is located can be obtained, and the operation corresponding to the voice command can be executed on the region image.
[0129] In some implementations, if the intent information in the voice command is determined to be an operation on an object corresponding to the target interface content, the electronic device can execute the operation corresponding to the voice command on the target object corresponding to the target interface content. Here, the target object is a file or an application.
[0130] The interface interaction method provided in the embodiments of this application will be further illustrated by examples below.
[0131] For example, referring to Figures 2, 3, and 6 simultaneously, when the electronic device displays the first interface A1 as shown in Figure 2, after detecting the call operation to the smart assistant, it can display the interface B1 as shown in Figure 3. Interface B1 displays the interactive area A2 of the smart assistant, and after detecting the content input by the user to the smart assistant, the input content can be displayed in the interactive area A2. After detecting the user's voice command "Is this chair for sale in application B?" and the selection operation of the area where "chair" is located, the operation trajectory information A3 is displayed in interface B1. After determining that the target interface content is the "chair picture" in the first interface, the smallest area containing the "chair picture" can be determined, and the edge of the smallest area can be drawn. Then, interface B3 as shown in Figure 6 is displayed, and the operation trajectory information A3 is updated to the edge of the smallest area. Then, the image of the area surrounded by the smallest area is searched through application B, and the prompt message "Operating application B" can be displayed in interface B3 to indicate that the "chair picture" is being searched through application B.
[0132] For example, referring to Figures 8, 9, 10, and 11 simultaneously, when the electronic device displays the first interface A1 as shown in Figure 8, after detecting a call to the smart assistant, it can display interface B4 as shown in Figure 9. Interface B4 displays the interactive area A2 of the smart assistant, and after detecting the user's input to the smart assistant, the input content can be displayed in the interactive area A2. After detecting the user's voice command "Save this picture to the album" and the selection operation of the area where the picture is located, the operation trajectory information A3 can be displayed in interface B4. Then, based on the voice command and the selection operation, the interface can determine the area to be operated. The target interface content is determined, and the electronic device can display interface B5 as shown in Figure 10, and display the prompt message "Understanding the screen" in interface B5 to indicate that the target interface content is being determined from the interface; after determining that the target interface content is the image in the first interface, the smallest area containing the image can be determined, and the edge of the smallest area can be drawn. Then, interface B6 as shown in Figure 11 is displayed, and the operation trajectory information A3 is updated to the edge of the smallest area. Then, the image of the area surrounded by the smallest area is saved, and the prompt message "Saving image" can be displayed in the interface to indicate that the image in the interface is being saved.
[0133] For example, please refer to Figures 12, 13, 14, 15, and 16 simultaneously. When the electronic device displays the first interface A1 as shown in Figure 12, after detecting the call operation to the smart assistant, it can display interface B7 as shown in Figure 13. Interface B7 displays the interactive area A2 of the smart assistant, and after detecting the content input by the user to the smart assistant, the input content can be displayed in the interactive area A2. After detecting the user's voice command "Translate this article" and the selection operation of the area where the article is located, the operation trajectory information A3 can be displayed in interface B7. Then, based on the voice command and the selection operation, the target interface content to be operated on can be determined, and the electronic device can display interface B8 as shown in Figure 14, and display the prompt message "Ongoing" in interface B8. The system is "Understanding the screen" to indicate that it is identifying the target interface content from the interface. After determining that the target interface content is the area where the selected article is located in the interface, the system can determine the smallest area containing the image and draw the edge of the smallest area. Then, the system displays interface B9 as shown in Figure 15, and updates the operation trajectory information A3 to the edge of the smallest area. Then, for the image of the area enclosed by the smallest area, the system retrieves the corresponding article file. The system can also display the prompt message "Collecting resources" to indicate that the system is retrieving the file of the selected article. After the file is retrieved, the system can translate the file. The electronic device can display interface B10 as shown in Figure 16 and display the prompt message "Translating" to indicate that the system is translating the file of the selected article.
[0134] Regarding the previous example, the electronic device can also handle situations where the user selects multiple files. Please refer to the figures, specifically Figures 12, 17, 18, 19, and 20. When the electronic device displays the first interface A1 as shown in Figure 12, after detecting an operation to invoke the smart assistant, it can display interface B11 as shown in Figure 17. Interface B11 displays the interactive area A2 of the smart assistant, and after detecting the user's input to the smart assistant, the input content can be displayed in the interactive area A2. After detecting the user's voice command "summarize these 3 articles" and the selection operation of the area containing the 3 articles, the operation trajectory information A3 is displayed in interface B11. Then, based on the voice command and the selection operation, the target interface content to be operated on can be determined, and the electronic device can display interface B12 as shown in Figure 18, and display the interface B12. In step 12, the prompt message "Understanding the screen" is displayed to indicate that the target interface content is being determined from the interface. After determining that the target interface content is the area where the selected article is located in the interface, the smallest area containing the image can be determined, and the edge of the smallest area can be drawn. Then, interface B13 as shown in Figure 19 is displayed, and the operation trajectory information A3 is updated to the edge of the smallest area. Then, for the image of the area enclosed by the smallest area, and for the image of the area, the files of the three corresponding articles are obtained. The prompt message "Collecting resources" can also be displayed in the interface to indicate that the files of the selected articles are being obtained. After obtaining the files of the three articles, the files of the three articles can be summarized. The electronic device can display interface B14 as shown in Figure 20, and display the prompt message "Summarizing" in interface B14 to indicate that the content of the selected article files is being summarized.
[0135] The interface interaction method provided in this application can determine the interface content to be operated on based on the user's voice input and touch operations on the display interface. This achieves the effect of optimizing and supplementing the content selected for touch operations based on voice commands, thereby ensuring the accuracy of the interface content to be operated and improving the accuracy of the electronic device in performing user-required operations. In addition, during the response to the user's voice input and touch operations on the display interface, the operation trajectory information of the touch operation is also displayed. After determining the interface content to be operated on, the displayed operation trajectory information is updated, so that the user can easily know the interface content selected from the display interface to be operated.
[0136] Please refer to Figure 21, which shows a flowchart of an interface interaction method provided in another embodiment of this application. This interface interaction method is applied to the aforementioned electronic device. The flowchart shown in Figure 21 will be described in detail below. Specifically, the interface interaction method may include the following steps:
[0137] Step S410: Display the first interface.
[0138] In this embodiment, step S410 can be referred to the content of the foregoing embodiments, and will not be repeated here.
[0139] Step S420: In response to the simultaneously detected voice command and touch operation, after detecting that the voice command input has stopped and the touch operation input has stopped, the operation corresponding to the voice command is executed based on the target interface content in the first interface, wherein the target interface content is determined from the first interface based on the voice command and the touch operation.
[0140] In this embodiment, when the electronic device displays a first interface, it can simultaneously detect the input voice and the operation on the interface. When both the input voice command and the touch operation are detected simultaneously, it can detect in real time whether to stop inputting the current voice command and whether to stop inputting the current touch operation. After detecting that the input of the above voice command and the input of the above touch operation have stopped, it can process the detected voice command and touch operation. That is, after determining the target interface content in the first interface based on the voice command and the touch operation, it executes the operation corresponding to the voice command based on the determined target interface content.
[0141] In some implementations, the electronic device may be equipped with a smart assistant, which can simultaneously detect voice input and touch operations on the interface when the smart assistant is awake and the selection mode in the smart assistant is enabled.
[0142] The interface interaction method provided in this application can determine the interface content to be operated on based on the user's input voice commands and touch operations on the display interface. This achieves the effect of optimizing and supplementing the content selected for touch operations based on voice commands, thereby ensuring the accuracy of the interface content to be operated and improving the accuracy of the electronic device in performing the user's required operations. In addition, the electronic device can simultaneously detect the input voice commands and touch operations on the interface, allowing the user to input voice commands while simultaneously performing touch operations on the display interface, thereby improving the user's operating efficiency and enhancing the user experience of the electronic device.
[0143] Please refer to Figure 22, which shows a flowchart of another embodiment of the interface interaction method provided in this application. This interface interaction method is applied to the aforementioned electronic device. The flowchart shown in Figure 22 will be described in detail below. Specifically, the interface interaction method may include the following steps:
[0144] Step S510: Display the first interface, where the application corresponding to the first interface is the first application.
[0145] In this embodiment of the application, the first interface displayed by the electronic device can be the application interface of the first application. That is, in this embodiment of the application, the scenario of operating on the interface content in the displayed interface can be the scenario where the user needs to operate on the corresponding content in the application interface of the first application when the electronic device displays the application interface of the first application.
[0146] Step S520: In response to the detected voice command and the touch operation in the first interface, a second application is invoked, and through the second application, the operation corresponding to the voice command is executed based on the target interface content in the first interface, wherein the target interface content is determined from the first interface based on the voice command and the touch operation.
[0147] In this embodiment of the application, when the electronic device displays the first interface, it responds to the voice commands and touch operations. After obtaining the content of the target interface, it can call the second application and execute the operation corresponding to the voice command based on the target interface content in the first interface through the second application.
[0148] In some implementations, before performing the operation corresponding to the above voice command through the second application, the electronic device may generate a script command for the second application based on the skill template of the second application and the above voice command, and then execute the script command to automatically perform the operation.
[0149] In some implementations, after acquiring the target interface content, the electronic device can determine the application used to perform the operation corresponding to the voice command, i.e., determine the second application to be invoked. Since the application to be invoked needs to be able to complete the operation required by the voice command, and different applications can perform different tasks—for example, shopping applications can be used to search for products, social applications can be used to share content, and translation applications can be used to translate documents—the electronic device can determine the second application to be invoked based on the voice command.
[0150] In one possible implementation, the electronic device can recognize the required operation task in response to a voice command; the electronic device then determines an application that matches the operation task from a knowledge base. This knowledge base may include application knowledge for each application installed on the electronic device, and the application knowledge may include at least the tasks that the application can perform.
[0151] In one possible implementation, the electronic device can pre-obtain application knowledge for each of the above applications from the server and save the obtained application knowledge to the above knowledge base. Optionally, the electronic device can obtain the application knowledge of each application installed from the server and save it to the above knowledge base.
[0152] In one possible implementation, considering that there are many applications installed in the electronic device and that many applications can perform the same operation task, the electronic device may identify multiple applications based on the operation task to be performed. Therefore, in the case of identifying multiple applications, a second application can be identified from the multiple applications based on the usage data of each application.
[0153] Optionally, the usage data mentioned above may include at least one of the number of times an application is used and the frequency of application use. The electronic device can determine a first usage score for each of the multiple applications identified above, and then select the application with the highest first usage score from among these applications as the second application. The first usage score of an application is positively correlated with both the number of times and the frequency of application use. Understandably, the number of times and the frequency of application use reflect user habits to some extent; therefore, the second application used to perform the operation can be accurately determined using the above method.
[0154] Optionally, the usage data above includes the number of times the application was used to perform the above tasks. The electronic device can use this usage count to determine a second usage score for each application, and then identify the application with the highest second usage score from among the multiple applications as the second application, wherein the second usage score is positively correlated with the number of times it was used. Understandably, the number of times used reflects the user's application preferences when performing the above tasks; therefore, applications that meet the user's preferences can be identified through the above method.
[0155] In some implementations, the voice command may contain keywords corresponding to the application used for the current operation task. For example, if the user inputs a voice command like "Share this picture with user 1 in application B", the electronic device can directly determine the second application based on the keywords corresponding to the application in the voice command.
[0156] Step S530: Display the second interface corresponding to the second application. The second interface is the interface that displays the operation result of the second application for the operation.
[0157] In this embodiment, after the electronic device completes the above operations, it can switch the display interface to a second interface corresponding to the second application. This second interface displays the operation results of the second application for the above operations. In this way, the operation results can be directly displayed in the application interface of the second application where the operation was performed, allowing the user to view and further manipulate the results.
[0158] In some implementations, when the electronic device invokes the second application and performs the above operations through the second application, the process can be handled in the background. After the operation is completed through the second application, the second application is brought to the foreground and displayed, thereby showing the second interface. Optionally, the electronic device can launch the second application, control it to run in the background, and use the background-running second application to perform the operations corresponding to the above voice commands. Thus, the entire operation process is invisible to the user; the user only needs to input voice commands and perform touch operations to directly see the second interface, without witnessing the process of launching the second application and operating it to perform the operations corresponding to the above voice commands.
[0159] Foreground operation refers to the state in which an electronic device displays the application's interface on the screen. When the application is running in the foreground, the user can see the application's interface on the screen. In contrast, background operation refers to the state in which the electronic device does not display the application's interface on the screen, and the user cannot interact with the application's interface. However, the application will still consume the electronic device's system resources when it is running in the background.
[0160] The interface interaction method provided in this application can determine the interface content to be operated on based on the user's voice input and touch operations on the display interface. This achieves the effect of optimizing and supplementing the content selected by touch operations based on voice commands, thereby ensuring the accuracy of the interface content to be operated and improving the accuracy of the electronic device in performing user-required operations. In addition, the electronic device can perform operations by calling another application, and after the operation is completed, the display interface will be switched to the interface of the application that performed the operation. This allows the search results to be displayed directly in the application's interface, enabling the user to view the operation results and perform further operations on them.
[0161] The interface interaction method involved in the above embodiments will be further described below with reference to Figure 23.
[0162] As shown in Figure 23, when the intelligent assistant in the electronic device is in a wake-up state, it can detect input voice commands and touch operations on the display interface. After detecting the cessation of touch operations and voice commands, it can trigger the electronic device to record the current application / scenario, take a screenshot of the current interface, record the area selected by the touch operation, and segment the content of the display interface. Then, the voice command, the recorded application / scenario, the screenshot, the recorded area selected by the touch operation, and the result of the content segmentation of the display interface are input into the large language model to complete the intent recognition of the voice command and determine the point of interest (i.e., the target interface content to be operated) in the display interface. If the type of the point of interest is text, the selection range of the touch operation can be redrawn in the interface based on the determined point of interest, and the text content within the selection range can be extracted. If the interest point is an image, the selection area for touch operation can be redrawn on the interface based on the determined interest point, and the image data within the selection area can be extracted. If the interest point is an object (e.g., a file, application), the selection area for touch operation can be redrawn on the interface based on the determined interest point, and the target object corresponding to the selection area can be extracted. Then, after optimizing and supplementing the prompt for the voice command, the prompt, along with the extracted text data, image data, or target object, is input into the large language model. This yields the operation result output by the large language model (e.g., the result of a translation operation) or the script instruction to execute the operation. If the output of the large language model is a script instruction, it can be further executed to complete the user's requested operation. Additionally, the electronic device can display the result of the operation so that the user is aware of the outcome.
[0163] Please refer to Figure 24, which shows a structural block diagram of an interface interaction device 600 provided in an embodiment of this application. This interface interaction device 600 utilizes the aforementioned electronic device and includes: an interface display module 610 and an operation execution module 620. The interface display module 610 is used to display a first interface; the operation execution module 620 is used to, in response to a detected voice command and a touch operation in the first interface, execute the operation corresponding to the voice command based on target interface content in the first interface, wherein the target interface content is determined from the first interface based on the voice command and the touch operation.
[0164] In some implementations, the operation execution module 620 may be specifically used to: respond to a detected voice command and a touch operation in the first interface, determine the content selected by the touch operation from the first interface based on the intent information corresponding to the voice command and the position information of the touch operation, and use it as the target interface content; and execute the operation corresponding to the voice command based on the target interface content.
[0165] In one possible implementation, the operation execution module 620 can also be used to determine interface content from the first interface that matches the intent information and the location information, as the target interface content.
[0166] Optionally, the operation execution module 620 can also be used to perform content recognition on the interface image of the first interface to obtain a recognition result; and determine the interface content that matches the location information and the intent information from the recognition result as the target interface content.
[0167] In one possible implementation, the location information includes the operation trajectory information of the touch operation. The interface display module 610 can also be used to display the operation trajectory information corresponding to the touch operation in the first interface before determining the content selected by the touch operation from the first interface based on the intent information corresponding to the voice command and the location information of the touch operation as the target interface content. The interface display module 610 can also be used to update the operation trajectory information based on the target interface content after determining the content selected by the touch operation from the first interface based on the intent information corresponding to the voice command and the location information of the touch operation as the target interface content. The updated operation trajectory information matches the area where the target interface content is located.
[0168] In one possible implementation, the interface display module 610 may also be used to display a first prompt message on the first interface before the operation corresponding to the voice command is executed based on the target interface content. The first prompt message is used to indicate that the target interface content is being determined from the first interface.
[0169] In some implementations, the operation execution module 620 may be specifically used for:
[0170] In response to the voice command and the touch operation, extract text information from the target interface content and execute the operation corresponding to the voice command on the text information; or
[0171] In response to the voice command and the touch operation, a region image of the area where the target interface content is located is obtained, and the operation corresponding to the voice command is executed on the region image.
[0172] In some implementations, the operation execution module 620 may be specifically used to: in response to the voice command and the touch operation, execute the operation corresponding to the voice command on the target object corresponding to the target interface content, wherein the target object is a file or an application.
[0173] In some implementations, the operation execution module 620 may be specifically used to: in response to the simultaneously detected voice command and touch operation, after detecting that the voice command input has stopped and the touch operation input has stopped, execute the operation corresponding to the voice command based on the target interface content in the first interface.
[0174] In some implementations, the interface display module 610 may also be used to display a second prompt message on the first interface, the second prompt message being used to indicate that the operation corresponding to the voice command is being executed.
[0175] In some implementations, the operation execution module 620 may be specifically used to: in response to the voice command and the touch operation, invoke a second application, and through the second application, execute the operation corresponding to the voice command based on the target interface content in the first interface; the interface display module 610 may also be used to display a second interface corresponding to the second application, the second interface being the interface displayed by the second application for the operation result corresponding to the operation.
[0176] In some implementations, the touch operation includes one or a combination of several of the following: selection, drawing, smearing, and clicking.
[0177] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0178] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0179] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0180] In summary, the solution provided in this application, by displaying a first interface and responding to detected voice commands and touch operations on the first interface, executes the operation corresponding to the voice command based on the target interface content on the first interface. The target interface content is determined from the first interface based on the voice command and touch operations. Therefore, it is possible to determine the interface content to be operated on based on the user's input voice command and touch operations on the display interface, thereby ensuring the accuracy of the operated interface content and improving the accuracy of the electronic device in executing user-required operations.
[0181] Please refer to Figure 25, which shows a structural block diagram of an electronic device provided in an embodiment of this application. The electronic device 100 can be a smartphone, tablet computer, e-reader, laptop computer, smartwatch, or other electronic device capable of running applications. The electronic device 100 in this application may include one or more of the following components: a processor 110, a memory 120, a display screen 130, an audio acquisition device 140, and one or more applications. The one or more applications may be stored in the memory 120 and configured to be executed by one or more processors 110. The one or more applications are configured to perform the methods described in the foregoing method embodiments.
[0182] The processor 110 uses various interfaces and lines to connect various parts within the electronic device 100. It executes various functions of the electronic device 100 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120.
[0183] The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described above, etc.
[0184] The display screen 130 is a display component used for displaying images and is usually located on the front panel of the electronic device 100; the audio acquisition device 140 is used to acquire audio from the environment, and the audio acquisition device 140 may be a microphone or the like.
[0185] Please refer to Figure 26, which shows a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 800 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0186] The computer-readable storage medium 800 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 800 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 800 has storage space for program code 810 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 810 may be compressed, for example, in a suitable form.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A user interface interaction method, wherein, The method includes: Display the first screen; In response to a detected voice command and a touch operation on the first interface, an operation corresponding to the voice command is executed based on the target interface content on the first interface, wherein the target interface content is determined from the first interface based on the voice command and the touch operation.
2. The method according to claim 1, wherein, The operation corresponding to the voice command, in response to the detected voice command and the touch operation on the first interface, is executed based on the target interface content on the first interface, including: In response to the detected voice command and the touch operation in the first interface, based on the intent information corresponding to the voice command and the position information of the touch operation, the content selected by the touch operation is determined from the first interface as the target interface content; Based on the content of the target interface, execute the operation corresponding to the voice command.
3. The method according to claim 2, wherein, The step of determining the content selected by the touch operation from the first interface as the target interface content based on the intent information corresponding to the voice command and the position information of the touch operation includes: The interface content that matches the intent information and the location information from the first interface is determined as the target interface content.
4. The method according to claim 3, wherein, The step of determining the interface content that matches the intent information and the location information from the first interface as the target interface content includes: Content recognition is performed on the interface image of the first interface to obtain the recognition result; The interface content that matches the location information and the intent information is determined from the recognition results and used as the target interface content.
5. The method according to any one of claims 2-4, wherein, The location information includes the operation trajectory information of the touch operation. Before determining the content selected by the touch operation from the first interface based on the intent information corresponding to the voice command and the location information of the touch operation, and using it as the target interface content, the method further includes: The interface displays the operation trajectory information corresponding to the touch operation; After determining the content selected by the touch operation from the first interface as the target interface content based on the intent information corresponding to the voice command and the location information of the touch operation, the method further includes: The operation trajectory information displayed on the interface is updated based on the target interface content, and the updated operation trajectory information matches the area where the target interface content is located.
6. The method according to claim 5, wherein, The step of updating the operation trajectory information displayed on the interface based on the target interface content includes: Based on the range of the target interface content in the first interface, the edge of the target interface content is drawn, and the drawn edge is displayed in the first interface so that the displayed operation trajectory information matches the area where the target interface content is located.
7. The method according to claim 5 or 6, wherein, The step of displaying the operation trajectory information corresponding to the touch operation in the interface includes: The operation trajectory information is displayed in the interface in a target color, which is different from the color of other content in the interface.
8. The method according to any one of claims 2-7, wherein, Before executing the operation corresponding to the voice command based on the target interface content, the method further includes: The interface displays a first prompt message, which indicates that the target interface content is being determined from the interface.
9. The method according to any one of claims 2-8, wherein, The step of determining the content selected by the touch operation from the first interface as the target interface content based on the intent information corresponding to the voice command and the position information of the touch operation includes: If, when it is determined from the first interface that the interface content matching the intent information and the location information includes multiple contents, candidate information corresponding to each of the multiple contents is displayed on the interface. If a selection operation for target candidate information is detected, the interface content corresponding to the target candidate information is determined and used as the target interface content.
10. The method according to claim 1, wherein, The operation corresponding to the voice command, in response to the detected voice command and the touch operation on the first interface, is executed based on the target interface content on the first interface, including: In response to the detected voice command and the touch operation in the first interface, the target interface content is determined by using a large language model and based on the text corresponding to the voice command, the touch position information and the interface image of the first interface. Based on the content of the target interface, execute the operation corresponding to the voice command.
11. The method according to claim 10, wherein, The step of using a large language model and determining the target interface content based on the text corresponding to the voice command, touch position information, and the interface image of the first interface includes: The text, touch location information, interface image, application information corresponding to the first interface, and the result of content segmentation of the interface image are input into the large language model to obtain the target interface content.
12. The method according to any one of claims 1-11, wherein, The operation corresponding to the voice command, in response to the detected voice command and the touch operation on the first interface, is executed based on the target interface content on the first interface, including: In response to the voice command and the touch operation, extract text information from the target interface content and execute the operation corresponding to the voice command on the text information; or In response to the voice command and the touch operation, a region image of the area where the target interface content is located is obtained, and the operation corresponding to the voice command is executed on the region image.
13. The method according to any one of claims 1-11, wherein, The operation corresponding to the voice command, in response to the detected voice command and the touch operation on the first interface, is executed based on the target interface content on the first interface, including: In response to the voice command and the touch operation, the operation corresponding to the voice command is executed on the target object corresponding to the target interface content, where the target object is a file or an application.
14. The method according to any one of claims 1-13, wherein, In response to a detected voice command and a touch operation on the first interface, based on the target interface content on the first interface, the operation corresponding to the voice command is executed, including: In response to the simultaneously detected voice command and touch operation, after detecting that the input of the voice command has stopped and the input of the touch operation has stopped, the operation corresponding to the voice command is executed based on the target interface content in the first interface.
15. The method according to any one of claims 1-14, wherein, The method further includes: A second prompt message is displayed on the interface, which indicates that the operation corresponding to the voice command is being executed.
16. The method according to any one of claims 1-15, wherein, The application corresponding to the first interface is the first application. In response to the detected voice command and the touch operation on the first interface, based on the target interface content in the first interface, it executes the operation corresponding to the voice command, including: In response to the voice command and the touch operation, a second application is invoked, and through the second application, the operation corresponding to the voice command is executed based on the target interface content in the first interface; The method further includes: The second interface corresponding to the second application is displayed, which is the interface for displaying the operation result of the second application for the operation.
17. The method according to any one of claims 1-16, wherein, The touch operation includes one or a combination of several of the following: selection, drawing, smudge, and clicking.
18. A user interface interaction device, wherein, The device includes: an interface display module and an operation execution module, wherein... The interface display module is used to display the first interface; The operation execution module is used to respond to the detected voice command and the touch operation in the first interface, and to execute the operation corresponding to the voice command based on the target interface content in the first interface. The target interface content is determined from the first interface based on the voice command and the touch operation.
19. An electronic device, wherein, include: One or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being configured to perform the method as described in any one of claims 1-17.
20. A computer-readable storage medium, wherein, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1-17.
Citation Information
Patent Citations
Projection touch screen control method and device
CN106155513A
A search method and a search apparatus
CN108984730A
Query information processing method and device
CN110119461A
Satisfying specified intent(s) based on multimodal request(s)
US20130159001A1