Voice control method and electronic equipment

By implementing visible-talk functions and simulated click area determination technology in electronic devices, the problem of click intention execution when the control is blocked is solved, ensuring the accuracy of voice control.

CN120279903AActive Publication Date: 2025-07-08HONOR DEVICE CO LTD

Patent Information

Application Number
CN202311872502.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-08
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

When multiple windows are displayed by electronic devices, controls under voice control may be blocked, resulting in the inability to accurately execute the user's click intention.

Method used

By implementing the visible and talkable function in the electronic device, voice commands are continuously collected, and when the control is detected to be blocked, the simulated click area is determined through the comparison of the floating window with the area of the control to be clicked to ensure that the click operation is performed in the area.

Benefits of technology

When the control is blocked, the user's click intention can be accurately executed, reducing simulated click errors, and improving the accuracy of voice control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279903A_ABST
    Figure CN120279903A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice control method and electronic equipment, relates to the technical field of voice processing, and is used for accurately executing a click intention of a user in a scene that a control matched with voice is shielded. The method is applied to electronic equipment comprising a microphone, and comprises the following steps: displaying a first interface comprising a first switch; and in response to the operation of turning on the first switch, turning on the voice control function. Specifically, a recording channel corresponding to a voice control function is opened, and the recording channel is used for obtaining voice collected by a microphone. In response to a voice click instruction collected by the microphone, searching a to-be-clicked control matched with the voice click instruction in a current display interface of the electronic equipment; under the condition that the to-be-clicked control is partially shielded by the floating window, determining a first simulation click area based on an area corresponding to the floating window and an area corresponding to the to-be-clicked control; and executing a click event corresponding to the voice click instruction in the first simulation click area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of voice processing, and in particular, to a voice control method and an electronic device. Background Art

[0002] With the development of technology, voice control has gradually become a common human-computer interaction method. In particular, when the user's hands are occupied in scenarios such as driving, cooking, and reading, it is very convenient and fast to control an electronic device through voice.

[0003] In the process of an electronic device performing a click operation in response to the voice input by the user, in the related art, it is necessary to first find the control to be clicked that matches the voice. Then, the electronic device simulates a click operation on the center position of the text control to be clicked. However, in some scenarios, the electronic device displays multiple windows simultaneously, and the control to be clicked may be blocked. In this case, the electronic device may not be able to accurately click on the control to be clicked in response to the voice, and thus the user's intention cannot be realized. Therefore, there is an urgent need for a method that can accurately execute the user's click intention in a scenario where the control matching the voice is blocked. Summary of the Invention

[0004] The embodiments of the present application provide a voice control method and an electronic device for accurately executing the user's click intention in a scenario where the control matching the voice is blocked.

[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, a voice control method is provided. The method is applied to an electronic device, and the electronic device includes a microphone.

[0007] The method includes:

[0008] The electronic device displays a first interface, and the first interface includes a first switch. In response to the user's operation of turning on the first switch, the voice control function is enabled. The enabling of the voice control function is controlled by the first switch and does not require waking up the electronic device. For example, the voice control function is a visible-and-speak function. After the voice control function is enabled, when a voice click instruction collected by the microphone is received, the electronic device can search for a control to be clicked that matches the first voice instruction on the current display interface. Then, the electronic device determines whether the control to be clicked is blocked. If the control to be clicked is partially blocked by a floating window, the electronic device first determines a first simulated click area based on the area corresponding to the floating window and the area corresponding to the control to be clicked. Then, within the first simulated click area, the click event corresponding to the voice click instruction is executed. In this way, when the electronic device executes the click event corresponding to the voice click instruction, it can accurately click on the control to be clicked, avoiding the problem of incorrect simulated clicks in this scenario. Thus, the electronic device can respond to the voice containing the intention of clicking on a control input by the user and perform a simulated click operation at an accurate position, thereby accurately implementing the user's intention.

[0009] Among them, enabling the voice control function includes: enabling the recording channel corresponding to the voice control function, and the recording channel is used to obtain the voice collected by the microphone. After the recording channel corresponding to the voice control function is enabled, the voice control function can obtain the voice (audio stream) collected by the microphone through this recording channel.

[0010] In a possible implementation manner of the first aspect, the control to be clicked is a control in the focus window of the electronic device.

[0011] In a possible implementation manner of the first aspect, the above-mentioned control to be clicked is partially blocked by a floating window, which may specifically include: the current display interface of the electronic device includes a floating window, the floating window is not the focus window, and the area corresponding to the floating window and the area corresponding to the control to be clicked include an overlapping area. Among them, the overlapping area is smaller than the area corresponding to the control to be clicked, and the control to be clicked is not completely blocked.

[0012] When the floating window is the focus window, the control to be clicked is a control on the floating window. In this case, the control to be clicked will not be blocked by the floating window. Therefore, when the current display interface includes a floating window and the floating window is not the focus window, the control to be clicked may be blocked. At this time, by comparing the areas corresponding to the floating window and the control to be clicked respectively, it can be determined whether there is an overlapping area, and thus it can be determined whether the control to be clicked is blocked by the floating window. In this solution, the electronic device can quickly determine whether the control to be clicked is blocked.

[0013] In a possible implementation of the first aspect, determining the first simulated click area based on the area corresponding to the floating window and the area corresponding to the control to be clicked may specifically include: subtracting the overlapping area from the area corresponding to the control to be clicked to obtain the first simulated click area. The above-mentioned first simulated click area is the partial area in the area corresponding to the control to be clicked that is not blocked by the floating window. Executing a click event in this first simulated click area can ensure that the electronic device does not click on other windows, thereby accurately executing the voice click instruction.

[0014] In a possible implementation of the first aspect, the above method may further include: when it is detected that the current display interface of the electronic device includes a floating window, obtaining the area corresponding to the current focus window of the electronic device. Then, comparing the area corresponding to the focus window with the area corresponding to the floating window. If the area corresponding to the focus window is exactly the same as the area corresponding to the floating window, it is determined that the floating window is the focus window. In this way, it is possible to quickly determine whether the floating window is the focus window, which is convenient for determining whether the control to be clicked is blocked by the floating window.

[0015] In a possible implementation of the first aspect, the above method may further include: when the control to be clicked is not blocked, executing the click event corresponding to the voice click instruction within the second simulated click area. The second simulated click area includes the area corresponding to the control to be clicked. If the control to be clicked is not blocked, the click event can be directly executed in the area corresponding to the control to be clicked.

[0016] Among them, the control to be clicked not being blocked may specifically include: the current display interface of the electronic device does not include a floating window.

[0017] Or, the control to be clicked not being blocked may specifically include: the current display interface of the electronic device includes a floating window, and the floating window is the focus window.

[0018] Or, the control to be clicked not being blocked may specifically include: the current display interface of the electronic device includes a floating window, the floating window is not the focus window, and the area corresponding to the control to be clicked does not overlap with the area corresponding to the floating window.

[0019] In a possible implementation of the first aspect, the above method may further include: when the control to be clicked is completely blocked by the floating window, sending a prompt message. The prompt message is used to indicate that it is not clickable. In this way, the execution result of the voice can be prompted to the user.

[0020] In a possible implementation of the first aspect, finding a control to be clicked that matches the voice click instruction on the current display interface of the electronic device includes: finding a first control that matches the voice click instruction on the current display interface of the electronic device. Obtain the click attribute of the first control. If the click attribute of the first control is clickable, determine the first control as the control to be clicked. Before performing the occlusion judgment on the control to be clicked, the electronic device also judges the click attribute of the control to be clicked. In this way, it is further ensured that the control to be clicked can execute the click event. Thus, when the electronic device accurately recognizes the voice intention, the possibility that the electronic device cannot realize the true intention of the voice input by the user is reduced, and the possibility of performing an accurate simulated click operation in response to the voice input by the user is increased.

[0021] In a possible implementation of the first aspect, the voice click instruction matches at least one interface hot word set by the electronic device; in this implementation, the first control is a text control.

[0022] In a possible implementation of the first aspect, finding a control to be clicked that matches the voice click instruction on the current display interface of the electronic device may further include: if the click attribute of the first control is not clickable, find a second control that belongs to the same control group as the first control. Obtain the click attribute of the second control. If the click attribute of the second control is clickable, determine the second control as the control to be clicked. In this way, it is further ensured that the control to be clicked can execute the click event. Thus, when the electronic device accurately recognizes the voice intention, the possibility that the electronic device cannot realize the true intention of the voice input by the user is reduced, and the possibility of performing an accurate simulated click operation in response to the voice input by the user is increased.

[0023] In a possible implementation of the first aspect, finding a control to be clicked that matches the voice click instruction on the current display interface of the electronic device may further include: if the click attribute of the first control is clickable and the first control is completely occluded by the floating window, find a second control that belongs to the same control group as the first control. Obtain the click attribute of the second control. If the click attribute of the second control is clickable, determine the second control as the control to be clicked.

[0024] In a possible implementation of the first aspect, the electronic device includes a voice control application package (APK), a voice processing engine, and an Activity Manager Service (AMS). The method may further include: The voice control APK receives the current display interface of the electronic device returned by the AMS, and parses the current display interface to obtain and save the text controls and picture controls of the current display interface. The voice control APK uses the text included in the text controls of the current display interface as interface hot words and sends them to the voice processing engine. The voice control APK distributes the voice click instruction to the voice processing engine in response to receiving the voice click instruction. The voice processing engine searches for target interface hot words that match the voice click instruction. The voice processing engine returns the target interface hot words to the voice control APK. The voice control APK queries whether the electronic device includes a floating window, and when the electronic device includes a floating window, determines whether the floating window is the focus window, and determines whether the area corresponding to the floating window and the area corresponding to the control to be clicked include an overlapping area.

[0025] In another possible implementation of the first aspect, within the first simulated click area, performing the click event corresponding to the voice click instruction includes: determining a simulated click position within the first simulated click area; and performing a simulated click operation at the simulated click position.

[0026] In a second aspect, the present application further provides an electronic device. The electronic device may include: a microphone, a display screen, a processor, and a memory. The microphone, the display screen, and the memory are respectively coupled to the processor. The microphone is used to collect language, and the display screen is used to display the interface of the electronic device. The memory is used to store computer execution instructions. When the electronic device runs, the processor executes the computer execution instructions stored in the memory, so that the electronic device executes the voice control method according to any one of the above first aspects.

[0027] In a third aspect, the present application provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed by the processor of the electronic device, the electronic device can execute the voice control method according to any one of the above first aspects.

[0028] In a fourth aspect, a computer program product including instructions is provided. When it runs on an electronic device, the electronic device can execute the voice control method according to any one of the above first aspects.

[0029] In a fifth aspect, a device (for example, the device may be a chip system) is provided. The device includes a processor for supporting an electronic device to implement the functions involved in the first aspect above. In a possible design, the device further includes a memory for storing necessary program instructions and data of the electronic device. When the device is a chip system, it may be composed of chips or may include chips and other discrete devices.

[0030] Among them, for the technical effects brought about by any one of the design manners in the second aspect to the fifth aspect, reference may be made to the technical effects brought about by different design manners in the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a schematic diagram of a scenario example of a voice control method;

[0032] Figure 2A It is a schematic diagram of an example process for enabling a voice control function;

[0033] Figure 2B It is a schematic diagram of a scenario example of a voice control method;

[0034] Figure 2C It is a schematic diagram of a scenario example of a voice control method;

[0035] Figure 3 It is a schematic flowchart of a voice control method provided by an embodiment of the present application;

[0036] Figure 4 It is a schematic diagram of a scenario example of the voice control method provided by an embodiment of the present application;

[0037] Figure 5A It is a schematic diagram of a display interface of an electronic device;

[0038] Figure 5B It is a schematic diagram of an example process for enabling a voice control function;

[0039] Figure 6 It is a schematic flowchart of a voice control method provided by an embodiment of the present application;

[0040] Figure 7A It is a schematic flowchart of a voice control method provided by an embodiment of the present application;

[0041] Figure 7B It is a schematic diagram of a scenario example of the voice control method provided by an embodiment of the present application;

[0042] Figure 7C It is a schematic diagram of a scenario example of the voice control method provided by an embodiment of the present application;

[0043] Figure 8Schematic diagram for determining the simulated click area provided by the embodiments of the present application;

[0044] Figure 9A Schematic flow chart of a voice control method provided by the embodiments of the present application;

[0045] Figure 9B Schematic diagram of a scenario example of a voice control method provided by the embodiments of the present application;

[0046] Figure 10 Schematic diagram of the software architecture of an electronic device provided by the embodiments of the present application;

[0047] Figure 11 Schematic diagram of the interaction of each module of the electronic device when implementing the voice control method provided by the embodiments of the present application;

[0048] Figure 12 Schematic diagram of the interaction of each module of the electronic device when implementing the voice control method provided by the embodiments of the present application;

[0049] Figure 13 Schematic flow chart of a voice control method provided by the embodiments of the present application;

[0050] Figure 14 Schematic diagram of the software architecture of an electronic device provided by the embodiments of the present application;

[0051] Figure 15 Schematic diagram of the interaction of each module of the electronic device when implementing the voice control method provided by the embodiments of the present application;

[0052] Figure 16 Schematic diagram of the interaction of each module of the electronic device when implementing the voice control method provided by the embodiments of the present application;

[0053] Figure 17 Schematic flow chart of a voice control method provided by the embodiments of the present application;

[0054] Figure 18 Schematic diagram of the hardware structure of an electronic device provided by the embodiments of the present application;

[0055] Figure 19 Schematic diagram of the structure of a chip system provided by the embodiments of the present application. Detailed implementation manners

[0056] To facilitate a clear description of the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:

[0057] The goal of automatic speech recognition (ASR) is to convert the lexical content in a user's speech into computer-readable input, such as keystrokes, binary codes, or character sequences.

[0058] Natural language understanding (NLU) is a general term for all method models or tasks that support machines to understand the content of text.

[0059] Dialogue management (DM) is used to control the process of human-computer dialogue and determine the response to the user at this moment based on the dialogue history information.

[0060] An activity is one of the four major components of the Android system and is a visual interface for user operations; it provides a window for users to complete operation instructions.

[0061] An intention is an idea of hoping to achieve a certain purpose. In the field of voice control, intention recognition is an important technology. By accurately identifying and understanding the needs and intentions of users, more accurate instructions can be executed in response to the user's speech, thus meeting the user's needs. Taking the user's input speech "query today's weather" as an example, the electronic device can perform intention recognition on this speech, extract the entity content in this speech: "query", "weather", and thus determine that the user's intention is to query the weather. Another example is that the user's input speech is "return to the previous level", and the electronic device can perform intention recognition on this speech, extract the entity content in this speech: "return", "previous level", and thus determine that the user's intention is to display the previous level page. Another example is that the user's input speech "open the calendar", and the electronic device can extract the entity content in this speech: "open", "calendar", and thus determine that the user's intention is to open the calendar application.

[0062] Voice control function:

[0063] Many electronic devices support the voice control function. The electronic device collects the user's input speech through a microphone, parses and recognizes the speech, and executes the instruction corresponding to the speech, realizing the control of the electronic device by the user through speech.

[0064] Generally speaking, in order to save power consumption of electronic devices and avoid false triggering, the voice control function needs to be turned on before it can be used. For example, the user inputs a preset word (called a wake-up word) to the electronic device through voice to wake up the electronic device. After the electronic device is awakened, it can execute the command corresponding to the voice, that is, the voice control function is turned on. For example, the user can turn on or off the voice control function by turning on or off the preset switch in the human-computer interaction interface of the electronic device.

[0065] In different electronic devices, the voice control function may have different names, such as "voice control", "intelligent voice", "voice assistant", "see and speak", "voice command", "free command", "intelligent AI", etc. The specific implementation of voice control functions with different names may also be different.

[0066] Several different implementations of the voice control function are described below by way of example.

[0067] Voice Assistant:

[0068] Before the user uses a voice assistant to control an electronic device, the electronic device needs to be woken up first. In one example, the wake-up word input by the user into the electronic device is detected, and the electronic device is awakened. In another example, the electronic device is awakened by detecting the user's long press of the power button. In another example, the breath generated when the user inputs voice to the electronic device is detected, and the electronic device is awakened. Generally speaking, before the electronic device is awakened, the microphone of the electronic device works in a power-saving mode (such as searching for signals at a lower power) to pick up sounds from the surrounding environment. The voice collected by the microphone is only detected for wake-up words at the kernel layer, and the corresponding recording channel of the voice assistant is not started in the system and driver of the electronic device.

[0069] The electronic device is awakened in response to a user operation (for example, receiving a wake-up word input by voice), and the corresponding recording channel of the voice assistant is started in the system and driver. After the electronic device is awakened, the voice (audio stream) collected by the microphone is sent to the voice assistant application for processing through the recording channel corresponding to the voice assistant. In this way, the electronic device can execute the instructions corresponding to the voice, enabling the user to control the electronic device through voice; it can also realize functions such as dialogue with the user.

[0070] Taking the electronic device as a mobile phone 100 as an example, illustratively, Figure 1 FIG. 1 shows a schematic diagram of a scenario in which a user uses a voice assistant to control a mobile phone 100. Figure 1As shown, the mobile phone 100 displays the desktop interface, and the user inputs the voice "Hello YOYO" into the mobile phone 100. In response to receiving the wake-up word "Hello YOYO", the mobile phone 100 is awakened. Exemplarily, after the mobile phone 100 is awakened, it plays the voice "I'm here" to prompt the user that the mobile phone 100 has been awakened. After the mobile phone 100 is awakened, the user can control the mobile phone 100 by voice. Exemplarily, as Figure 1 shown, the user inputs the voice "Open the video" into the mobile phone. The mobile phone 100 parses and recognizes the voice input by the user and executes the instruction corresponding to the voice "Open the video". Exemplarily, in response to receiving the voice "Open the video", the mobile phone 100 starts the video application.

[0071] In some other examples, the electronic device can also be awakened by the voice assistant in response to receiving the user's operation of long pressing the power button.

[0072] In some implementation manners, after the voice assistant of the electronic device is awakened, the user can issue an instruction to the electronic device by inputting voice to the electronic device, and the electronic device executes the instruction corresponding to the voice. After the electronic device executes an instruction, or, within a period of time (such as within 8 seconds) after the voice assistant of the electronic device is awakened, if no instruction is received from the user through voice, the electronic device no longer responds to the instruction issued through voice. For example, the electronic device will close the recording channel corresponding to the voice assistant. The user needs to input the wake-up word to the electronic device again to wake up the electronic device before being able to issue an instruction to the electronic device by inputting voice again. That is to say, after the voice assistant of the electronic device is awakened, it enters a "short voice reception" state and can respond to the instruction issued by the user through voice within a relatively short period of time (such as within 8 seconds).

[0073] In some implementation manners, when the electronic device is connected to the network, it supports entering the continuous conversation scenario after being awakened, and the user can have a continuous conversation with the electronic device. After each broadcast by the electronic device, it will continue to pick up the voice and does not need to be awakened repeatedly until the user exits the continuous conversation through instructions such as "Exit".

[0074] See and speak:

[0075] See and speak is implemented locally by the electronic device without the need to connect to the network.

[0076] In some examples, see and speak is controlled by a preset switch. The user can turn on the preset switch to enable the see and speak function, or turn off the preset switch to disable the see and speak function.

[0077] Exemplarily, as Figure 2AAs shown, the user can open the settings function of the mobile phone 100; for example, the user clicks on the application icon of the "Settings" application on the desktop. In response to the user's click operation on the application icon of the "Settings" application, the mobile phone 100 displays the "Settings" interface 101. The "Settings" interface 101 includes a "Smart Voice" option 102, and the "Smart Voice" option 102 is used to set the smart voice function. Exemplarily, in response to the user's click operation on the "Smart Voice" option 102, the mobile phone 100 displays the "Smart Voice" interface 103, and the "Smart Voice" interface 103 includes a "See and Speak" option 104. The user can click on the "See and Speak" option 104 to set the options related to the see and speak function. Exemplarily, referring to Figure 2A , in response to the user's click operation on the "See and Speak" option 104, the mobile phone 100 displays the "See and Speak" interface 105. Optionally, the "See and Speak" interface 105 includes a prompt message 106 for prompting the user about the usage method of the see and speak function. The "See and Speak" interface 105 also includes a "See and Speak" switch 107 (i.e., the above-mentioned preset switch). The user can click on the "See and Speak" switch 107 to turn on or off the "See and Speak" switch. In one example, in response to receiving the user's click operation on the "See and Speak" switch 107, the "See and Speak" switch of the mobile phone 100 is turned on, and the see and speak function is enabled. Optionally, the "See and Speak" interface 105 displays a prompt message 108 for prompting the user that the see and speak function has been successfully enabled.

[0078] In one implementation, after the see and speak function is enabled, the mobile phone 100 displays a first recording icon, and this first recording icon indicates that the see and speak function has been enabled. Exemplarily, as Figure 2A shown, after the "See and Speak" switch 107 is turned on, the status bar of the interface displayed by the mobile phone 100 shows a recording icon 10a, indicating that the see and speak function has been enabled.

[0079] In one scenario, the preset switch corresponding to see and speak on the electronic device is not turned on, and the microphone of the electronic device is not enabled. When the preset switch corresponding to see and speak is turned on, the electronic device starts the microphone and starts the corresponding recording channel for see and speak in the system and the driver. In this way, the voice (audio stream) collected by the microphone can be sent to the see and speak application through the corresponding recording channel for see and speak for processing, and the user can control the electronic device through voice.

[0080] In another scenario, when the corresponding preset switch for "visible speech" is not turned on, the microphone of the electronic device operates in a power-saving mode (for example, searching for signals with a lower power) to pick up the surrounding environment sounds. When the corresponding preset switch for "visible speech" is turned on, the corresponding recording channel for "visible speech" is activated in the system and drivers of the electronic device. In this way, the voice (audio stream) collected by the microphone can be sent to the "visible speech" application through the corresponding recording channel for processing, enabling the user to control the electronic device by voice.

[0081] After "visible speech" is enabled, the corresponding recording channel for "visible speech" is activated in the system and drivers of the electronic device, and the electronic device enters a "long recording" state, continuously collecting surrounding sounds. The user can issue commands to the electronic device by voice at any time without the need to input a wake word to wake up the electronic device.

[0082] In one implementation, after any function on the electronic device enables the voice input function of the electronic device (turns on the microphone and activates the recording channel), the electronic device will send a prompt message to the user to indicate that the electronic device is in a voice collection state. In this way, the user's privacy can be prevented from being leaked. For example, after the "visible speech" function is enabled, the electronic device enters a continuous voice collection state, and a second recording icon is displayed on the display interface of the electronic device. This second recording icon indicates that the recording channel is open, used to prompt the user that the microphone is collecting voice. Exemplarily, as Figure 2A shown, the status bar of the display interface of the mobile phone 100 shows a recording icon 10b, indicating that the recording channel is open.

[0083] After the "visible speech" function is enabled, the electronic device activates the recording channel and continuously collects voice through the microphone. The user can input voice to the electronic device at any time. The electronic device parses and recognizes the voice input by the user and executes the instruction corresponding to the voice.

[0084] The instructions that the "visible speech" function supports the user to input by voice can include: system instructions, such as swiping left, swiping right, swiping up, returning to the desktop, going back, turning up the volume, turning down the volume; video application instructions, such as playing, pausing, stopping, fast-forwarding, rewinding; and e-book playback application instructions, such as going to the previous page, going to the next page, going to the table of contents, going to the next chapter.

[0085] The instructions that the "visible speech" function of the electronic device supports the user to input by voice can be divided into multiple vertical categories, and one of the vertical categories is the operation vertical category. For operation vertical category instructions, when the electronic device responds to the voice and executes the instruction corresponding to the voice, it will simulate the user's operation.

[0086] Hot words:

[0087] The electronic device can set some hot words. After the visible and speakable function is turned on, if the received voice matches at least one of the set hot words, the electronic device executes the instruction corresponding to the voice.

[0088] The hot words can be pre-configured on the electronic device, can be obtained according to the content in the display interface of the electronic device, or can be input by the user, etc. According to the different usage scopes and sources of the set hot words, the hot words can be divided into system-level hot words, scenario-level hot words, and interface hot words.

[0089] Among them, the system-level hot words are pre-configured and applicable to any application on the electronic device. When any application on the electronic device is running in the foreground, if the voice input by the user matches at least one system-level hot word, the electronic device executes the instruction corresponding to the user's voice. The system-level hot words are global and do not depend on the application or the application interface. Exemplarily, the system-level hot words can include: swipe left, swipe right, swipe up, go back to the desktop, return, etc.

[0090] The scenario-level hot words are applicable to all applications within a scenario. In one implementation, multiple scenarios are pre-set in the electronic device, and each scenario corresponds to at least one application. Exemplarily, the pre-set scenarios can include audio-video scenarios, e-book scenarios, home screen scenarios, and permission pop-up scenarios, etc. Different scenarios can be pre-set with different scenario-level hot words. Exemplarily, the scenario-level hot words corresponding to the audio-video scenario can include: play, pause, stop, fast forward, and rewind, etc. The scenario-level hot words corresponding to the e-book scenario can include: previous page, next page, table of contents, and next chapter, etc. The scenario-level hot words corresponding to the home screen scenario can include: application list information, settings menu item, etc. The scenario-level hot words corresponding to the permission pop-up scenario can include: allow, confirm, deny, and I got it, etc.

[0091] The interface hot words are applicable to a certain interface. In some embodiments, the electronic device obtains the hot words from the display interface of the foreground application, that is, obtains the interface hot words. Exemplarily, refer to Figure 2B, the mobile phone 100 displays the "My" interface 110 of the video application. The text information in the "My" interface 110 includes "5G", "8:00", "Login / Register", "My Downloads", "Following & Favorites", "My Purchases", "My Scenes", "History", "Coupon Package", "Settings", "Feedback", "Customer Service", "Home", "Membership", "Short Videos", and "My", etc. The mobile phone 100 sets the text in the currently displayed interface of the foreground application as the interface hot words, that is, the interface hot words include "5G", "8:00", "Login / Register", "My Downloads", "Following & Favorites", "My Purchases", "My Scenes", "History", "Coupon Package", "Settings", "Feedback", "Customer Service", "Home", "Membership", "Short Videos", and "My", etc. When the electronic device monitors the switching of the displayed interface, after the interface is switched, it can re-obtain the text content of the currently displayed interface and update the interface hot words.

[0092] In some scenarios, the electronic device can display multiple windows simultaneously. For example, in the scenario where the electronic device displays a floating window, the electronic device displays two windows simultaneously. In this scenario, the interface hot words obtained and saved by the electronic device are specifically the interface hot words displayed in the focused window. The focused window refers to the selected window among multiple windows, that is, the currently operated window. In this scenario, the global operations performed by the user act on this focused window.

[0093] Exemplarily, the electronic device simultaneously displays the settings interface and displays the calculator interface in the floating window. The focused window is the window corresponding to the settings interface. At this time, the interface hot words obtained and saved by the electronic device are the hot words obtained and saved from the settings interface.

[0094] Different hot words can correspond to different instructions. Since the interface hot words are set according to the displayed text of the currently displayed interface, in some embodiments, the instructions corresponding to the interface hot words are control click instructions. Specifically, when the electronic device determines that the received voice matches at least one interface hot word, it can execute the control click instruction on the position of the text corresponding to the interface hot word on the interface. In some examples, when the electronic device executes the control click instruction, it can specifically execute the click event corresponding to the control click instruction, that is, perform a simulated click operation at the position of the corresponding text.

[0095] In Figure 2BIn the example shown, the mobile phone 100 receives the voice "Open History" input by the user. The voice "Open History" matches the interface hot word "History", and the mobile phone 100 executes the instruction corresponding to the voice "Open History". Specifically, the mobile phone 100 can perform a simulated click operation within the area corresponding to the "History" control corresponding to "History", thereby realizing the user's intention to open the history. As Figure 2B shown, after the mobile phone 100 performs the simulated click operation, the history interface 113 can be displayed.

[0096] Among them, when the mobile phone 100 executes the instruction corresponding to the voice "Open History", it is necessary to determine the area corresponding to the "History" control corresponding to "History". Then, according to the area corresponding to the "History" control, a simulated click position is determined. In one example, the electronic device can determine the center position of the area corresponding to the "History" control as the simulated click position. As Figure 2B shown, the center position of the area corresponding to the "History" control is position 112, which can be determined as the simulated click position.

[0097] When the mobile phone 100 displays the "My" interface of the video application, the user can also input other voices, such as "Coupon Package". The mobile phone 100 receives the voice "Coupon Package" input by the user, determines that the voice "Coupon Package" matches the interface hot word "Coupon Package", and the mobile phone 100 will execute the instruction corresponding to the voice "Coupon Package". As Figure 2C shown, the mobile phone 100 performs a simulated click operation in the area corresponding to the "Coupon Package" control 114 corresponding to the interface hot word "Coupon Package".

[0098] However, in some embodiments, the click attribute of the "Coupon Package" control 114 is non-clickable. Then, after the mobile phone 100 performs a simulated click operation on the "Coupon Package" control 114 in response to the voice "Coupon Package", the mobile phone 100 cannot open the coupon package interface. In this way, the mobile phone 100 cannot realize the user's intention.

[0099] Based on this, an embodiment of the present application proposes a voice control method, which can be applied to an electronic device supporting voice input. After the electronic device responds to the user's operation and activates the voice control function, it can receive voice and execute the instruction corresponding to the voice. When the electronic device receives the first voice containing a control click intention on the second interface, it searches for a control (denoted as the first control) on the second interface that matches the first voice. If the first control is not clickable, then the electronic device searches for a second control on the second interface that belongs to the same control group as the first control. If the second control is clickable, a simulated click area is determined based on the second control, and a simulated click operation is performed within the simulated click area, thereby realizing the intention of the first voice. In this way, when the voice intention is accurately recognized, the possibility that the electronic device cannot realize the true intention of the voice input by the user is reduced, and the possibility of performing an accurate simulated click operation in response to the voice input by the user is increased.

[0100] Exemplarily, the above-mentioned electronic device may be a mobile phone, a tablet computer, a laptop computer, a personal computer (PC), an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a smart home device (such as a smart TV, a smart screen, a large screen, a smart speaker, a smart air conditioner, etc.), a personal digital assistant (PDA), a wearable device (such as a smart watch, a smart bracelet, etc.), a vehicle-mounted device, a virtual reality device, etc. The embodiments of the present application do not make any restrictions on this.

[0101] The following will specifically describe the implementation manner of the voice control method proposed in the embodiments of the present application with reference to the accompanying drawings. Figure 3 The flowchart of the voice control method in some embodiments is shown.

[0102] S200. Display a first interface, and the first interface includes a first switch.

[0103] The first switch is used to turn on or off the voice control function. The user can turn on or off the voice control function through the first switch on the first interface. In some embodiments, the first interface may be Figure 2A The "visible and audible" interface 105 shown in the figure; the first switch may be the "visible and audible" switch 107.

[0104] S201. Receive an operation to turn on the first switch.

[0105] S202. In response to the operation of turning on the first switch, activate the voice control function.

[0106] In some embodiments, after the voice control function is enabled, the electronic device activates the recording channel corresponding to the voice control function. This recording channel is used to obtain the voice collected by the microphone. Moreover, after the voice control function is enabled, the electronic device is in a long recording state, continuously collecting surrounding sounds. The user can issue commands to the electronic device via voice at any time without having to input a wake word to wake up the electronic device.

[0107] In addition, when the electronic device enables the voice control function, it may specifically further include: displaying a first recording icon and a second recording icon. The first recording icon indicates that the voice control function is enabled, and the second recording icon indicates that the recording channel is enabled. By using the recording icons, the user is informed that the voice control function has been enabled and the electronic device is currently recording. In this way, the user can quickly obtain the status information of the current electronic device.

[0108] S203. Display a second interface.

[0109] The second interface can be any interface of the electronic device. After the voice control function of the electronic device is enabled, on any display interface, the electronic device can receive the voice input by the user and execute the command corresponding to the voice. In some embodiments, the second interface can also be the first interface, that is, after the electronic device enables the voice control function, it can receive the voice input by the user on the first interface and execute the command corresponding to the voice.

[0110] In some embodiments, after the electronic device enables the voice control function, the electronic device can, after each display interface is switched, parse the switched display interface, obtain the control information of each control in the display interface, and save it. The control information may include: control attributes (text control / picture control), the area where the control is located (for a rectangular control, it can be represented by the upper left corner point coordinates and the lower right corner point coordinates), and whether the control is clickable. In this embodiment, after the electronic device displays the second interface, it can also parse the second interface to obtain the control information of each control in the second interface. In this way, it is convenient to make an accurate response to the received voice subsequently.

[0111] Since the electronic device has enabled the voice control function, the user can input voice to the electronic device through the voice control function on the second interface. Correspondingly, the electronic device receives the voice input by the user, as in S204.

[0112] S204. Receive a first voice.

[0113] S205. Determine the user intention corresponding to the first voice.

[0114] In some embodiments, the first voice matches at least one hotword preset in the electronic device. In this way, the electronic device can respond to the first voice and execute the instruction corresponding to the first voice, which can be denoted as the first voice instruction.

[0115] In some embodiments, after S204, the electronic device can perform voice parsing on the first voice to obtain the parsed text corresponding to the first voice. Then, the electronic device can perform intent recognition on the obtained parsed text. In this embodiment, S205 above can specifically include: parsing the first voice to obtain the parsed text corresponding to the first voice. Then, performing intent recognition on the parsed text corresponding to the first voice to obtain the user intent corresponding to the first voice.

[0116] As can be seen from the above description, after receiving the voice, the electronic device can match the voice with the hotwords. If it is determined that the voice matches at least one hotword set in the electronic device, the electronic device can execute the instruction corresponding to the voice. Therefore, in some embodiments, performing intent recognition on the parsed text corresponding to the first voice to obtain the user intent corresponding to the first voice can specifically include: matching the parsed text corresponding to the first voice with the hotwords set in the electronic device. If the parsed text corresponding to the first voice matches at least one hotword of the electronic device, the user intent can be determined according to the instruction corresponding to the hotword.

[0117] For example, if the parsed text corresponding to the first voice matches the interface hotword, the electronic device can determine that the instruction corresponding to the first voice is a control click instruction. Furthermore, it can be determined that the user intent of the first voice is a control click intent. For another example, if the hotword that the parsed text corresponding to the first voice matches belongs to "swipe left" in the system-level hotwords, the user intent can be determined to be swiping left. Or, if the hotword that the parsed text corresponding to the first voice matches belongs to "pause" in the scenario-level hotwords, the user intent can be determined to be controlling the electronic device to pause the playback.

[0118] As can be seen from the above description, the interface hotwords set in the electronic device are obtained from the current focused window. In embodiments where the electronic device displays multiple windows simultaneously, the first voice input by the user may be the content in a non-focused window. In this case, when the electronic device matches the parsed text corresponding to the first voice with the interface hotwords, it will not be able to obtain an interface hotword that matches the parsed text corresponding to the first voice. In some embodiments, in this case, the electronic device will recognize the first voice as an invalid instruction.

[0119] In some other embodiments, the electronic device may also preset multiple valid instructions in advance. When the voice input by the user matches the valid instructions set by the electronic device, the electronic device may execute the corresponding instructions in response to the voice. Exemplarily, the valid instructions may include: "Return", "Go back to the desktop", "Play / Pause", "Previous page / Next page", "Swipe left / Swipe right", and "Open [application name]", etc. Different valid instructions may correspond to different intents. In this embodiment, S205 above may specifically include: finding a valid instruction that matches the first voice. According to the valid instruction that matches the first voice, determining the user intent corresponding to the first voice. Among them, to find a valid instruction that matches the first voice, the first voice may be first subjected to voice parsing, and then the obtained parsed text is compared one by one with the valid instructions stored in the electronic device to determine whether the parsed text matches at least one valid instruction.

[0120] In the above embodiments, the situation where the electronic device finds a matching hot word or a matching valid instruction according to the parsed text corresponding to the first voice is described. In some other embodiments, it is possible that no matching hot word or valid instruction can be found for the first voice input by the user. In this case, the electronic device cannot determine the user intent and thus cannot execute the corresponding instruction in response to the voice.

[0121] When it is determined that the user intent corresponding to the first voice is a control click intent, the instruction corresponding to the first voice may be recorded as a voice click instruction.

[0122] In other embodiments, the electronic device may also determine the user intent corresponding to the first voice in other ways.

[0123] S206. Determine whether the user intent is a control click intent.

[0124] After determining the user intent, it can be determined whether the user intent is a control click intent. In some embodiments, if the first voice matches at least one interface hot word set by the electronic device, it can be determined that the user intent is a control click intent.

[0125] If the judgment result of S206 is negative, it means that the user intent corresponding to the first voice is not a control click intent. In this case, the electronic device may directly execute the instruction corresponding to the first voice to implement the user intent. It should be noted that the case where the judgment result of S206 is negative is not shown in Figure 3 is not shown.

[0126] When the judgment result of S206 is positive, it means that the user intent corresponding to the first voice is a control click intent. After that, the electronic device may determine the control to be clicked corresponding to the first voice.

[0127] S207. Search for a first control that matches the first voice in the second interface.

[0128] If it is determined according to the first voice that the user's intention is a control click intention, then the electronic device needs to determine the control to be clicked. Then, the area where the control to be clicked is located is obtained before the control click instruction can be executed. In some embodiments, the control found to match the voice is a text control on the current display interface of the electronic device.

[0129] In some embodiments, when determining the user intention corresponding to the first voice, the first voice is parsed to obtain a parsed text, and then the hot words that match the first voice are determined according to the parsed text. In this embodiment, the above S207 may specifically include: searching for a first control that matches the first voice according to the parsed text corresponding to the first voice. As Figure 2B In the example shown, the user inputs the voice "Open History", and the electronic device determines that this voice matches the interface hot word "History". After that, the electronic device can find the corresponding "History" control according to this interface hot word "History". As Figure 2C In the example shown, the user inputs the voice "Coupon Package", and the electronic device determines that this voice matches the interface hot word "Coupon Package". After that, the electronic device can find the corresponding "Coupon Package" control according to this interface hot word "Coupon Package". In the embodiment where the first voice matches at least one interface hot word, the first control that matches the first voice is a text control.

[0130] In other embodiments, the first control that matches the first voice may also be a picture control; such as Figure 2A The picture control 109 corresponding to the return icon shown.

[0131] In the embodiment where the electronic device simultaneously displays multiple windows, the second interface is the interface displayed in the focus window.

[0132] When the electronic device executes the simulated click instruction corresponding to the first voice, in order to avoid the situation where the simulated click is unsuccessful, it can first determine whether the area to be clicked is clickable, such as S208.

[0133] S208. Determine whether the first control is clickable.

[0134] It should be noted that what is judged in S208 is whether the first control itself has the property of being clickable.

[0135] As can be seen from the above embodiments, in some embodiments, after the electronic device displays the second interface, it can parse the display content of the second interface to obtain and save the control information on the second interface. In this embodiment, the above S208 may specifically include: obtaining the control information of the first control, and determining whether the first control is clickable according to the control information of the target text control.

[0136] If the judgment result of S208 is yes, it means that the first control is clickable. At this time, a simulated click operation can be directly performed within the area corresponding to the first control. S213 and S214 can be executed.

[0137] If the judgment result of S208 is no, it means that the first control is not clickable. At this time, S209 can be executed.

[0138] S209. Determine whether there is a second control belonging to the same control group as the first control.

[0139] In some embodiments, the control information obtained by the electronic device through interface parsing may further include: the control group to which the control belongs. In this embodiment, the electronic device can determine whether there is a peer control belonging to the same control group as the first control by querying the control information; and the type of the peer control and whether it is clickable.

[0140] Among them, the second control can be any type of control, such as a text control or a picture control.

[0141] In the embodiments of the present application, the above S209 may specifically include: searching for whether there is a peer control belonging to the same control group as the first control. In one example, if the second interface includes a peer control belonging to the same control group as the first control, after S209, the control information such as the type of the peer control and whether it is clickable can also be obtained.

[0142] If the judgment result of S209 is no, it means that the first control does not have a peer control belonging to the same control group. In this case, the electronic device may no longer execute the instruction corresponding to the first voice. Further, in some examples, the electronic device can issue a prompt message (such as displaying a prompt message on the display screen) to prompt the user that it cannot be clicked.

[0143] If the judgment result of S209 is yes, it means that the first control has a peer control belonging to the same control group. In this way, the instruction corresponding to the first voice can be executed based on this peer control. Before that, the electronic device also needs to judge whether the peer control is clickable, such as S210.

[0144] The interface displayed by the electronic device includes a control group with two or more controls, usually including an image control and a text control. In one scenario, the text control in the same control group is non-clickable, and the image control is clickable. In some embodiments, the above S209 may specifically include: determining whether there is a sibling image control corresponding to the first control. In this embodiment, if there is a sibling image control corresponding to the first control, then S210 is executed.

[0145] S210. Determine whether the second control is clickable.

[0146] For the specific implementation process of determining whether the second control is clickable, reference can be made to the description of determining whether the first control is clickable.

[0147] If the judgment result of S210 is negative, it means that the second control is non-clickable. In this case, the electronic device may no longer execute the instruction corresponding to the first voice. Further, the electronic device may issue a prompt message, which is used to prompt the user that it is non-clickable, such as S215. Exemplarily, the electronic device issues a prompt message, which may specifically be by displaying a text prompt message or an image prompt message on the display screen; and / or by emitting a voice prompt message through a speaker; and / or, by emitting a prompt message in a vibration mode.

[0148] If the judgment result of S210 is positive, it means that the second control is clickable. In this way, the electronic device may execute the instruction corresponding to the first voice based on this second control.

[0149] S211. Obtain the area corresponding to the second control.

[0150] In some embodiments, the electronic device may obtain the area corresponding to the second control by obtaining the control information of the second control. Obtaining the area corresponding to the control may specifically mean obtaining the coordinate position of the area corresponding to the control on the display interface.

[0151] The control may be presented in various shapes on the display interface of the electronic device, and the most common is a rectangular control. Taking the rectangular control as an example, the following method may be used to determine the coordinate position of the area corresponding to the control on the display interface: obtain the coordinates of two corner points on a diagonal line of the control (such as the coordinates of the upper left corner point and the lower right corner point, or the lower left corner point and the upper right corner point). If the second control is a rectangular control, then the above S211 may specifically obtain the coordinates of two corner points on a diagonal line of the second control, such as the coordinates of the upper left corner point and the lower right corner point.

[0152] Among them, the coordinates of a point on the display interface of the electronic device can be represented by the coordinates in the screen coordinate system. In some examples, the upper left corner of the screen is the origin coordinate (0, 0) of the screen coordinate system. The positive direction of the X-axis extends to the right from the origin, and the positive direction of the Y-axis extends downward from the origin.

[0153] If the control is of other shapes, for example, the control is circular, the coordinate position of the area corresponding to the control on the display interface can be determined by obtaining the center point coordinate and radius of the control. In other embodiments, when the control is in the shape of an ellipse, polygon, etc., the coordinate position of the area corresponding to the control on the display interface can also be determined by other means.

[0154] S212. Perform a simulated click operation based on the area corresponding to the second control.

[0155] In some embodiments, the above S212 may specifically include: performing a simulated click operation within the area corresponding to the second control. Further, performing a simulated click operation within the area corresponding to the second control may specifically include: obtaining the center position of the area corresponding to the second control as the simulated click position, and performing a simulated click operation at the simulated click position.

[0156] In other embodiments, the above S212 may specifically include: splicing the areas corresponding to the first control and the second control respectively to obtain a spliced area. Then, perform a simulated click operation within the spliced area. In some examples, performing a simulated click operation within the spliced area may specifically include: obtaining the center position of the spliced area as the simulated click position; and then performing a simulated click operation at the simulated click position.

[0157] The specific implementation process of the electronic device performing a simulated click operation at the simulated click position may refer to the description in related technologies and will not be elaborated in the embodiments of the present application.

[0158] Please refer to Figure 4 , in some examples, the mobile phone 100 displays the "My" interface 301 of the video application, and the user inputs the voice "coupon package" to the electronic device. In response to receiving the voice "coupon package", the mobile phone 100 searches for the text control that matches the voice "coupon package" in the "My" interface 301, that is, the text control 302. Then, the mobile phone 100 determines whether the text control 302 is clickable. After determining that the text control 302 is not clickable, the mobile phone 100 can search for the second control in the "My" interface 301 that belongs to the same control group as the text control 302. In one example, the mobile phone 100 finds that the second control belonging to the same control group as the text control 302 is the picture control 303. After the mobile phone 100 determines that the picture control 303 is clickable, it can execute the instruction corresponding to the voice "coupon package" based on the picture control 303, that is, the simulated click instruction.

[0159] In some embodiments, the mobile phone 100 may obtain the area corresponding to the picture control 303 and execute a simulated click instruction within the area corresponding to the picture control 303.

[0160] In other embodiments, the mobile phone 100 may splice the area corresponding to the picture control 303 and the area corresponding to the text control 302, and then execute a simulated click instruction based on the spliced area obtained by splicing; that is, execute a simulated click operation within the spliced area.

[0161] Exemplarily, after the mobile phone 100 executes the simulated click instruction corresponding to the voice "coupon package" based on the picture control 303, it may open the coupon package and display the coupon package interface 304.

[0162] In the technical solution proposed in the embodiments of the present application, after the electronic device receives the first voice containing the control click intention, if it is determined that the first control corresponding to the control click intention is not clickable, it searches for a second control in the current display interface that belongs to the same control group as the first control. If there is a clickable second control in the current display interface that belongs to the same control group as the first control, the instruction corresponding to the first voice may be executed based on the second control. In this way, when the recognition of the received voice intention is accurate, the possibility that the electronic device cannot realize the true intention of the voice input by the user can be reduced, and the possibility of executing an accurate simulated click operation in response to the voice input by the user can be improved.

[0163] In the case where the judgment result of S208 is yes, the electronic device may directly perform a simulated click operation on the first control to realize the user intention. As Figure 3 shown in S213 and S214.

[0164] S213. Obtain the area corresponding to the first control.

[0165] S214. Perform a simulated click operation in the area corresponding to the first control.

[0166] As Figure 2B shown in the example, the first control corresponding to the voice "open history" input by the user is the "history" control 111. The "history" control 111 is clickable. At this time, the mobile phone 100 may directly respond to the voice input by the user and perform a simulated click operation on the "history" control 111. In this way, the user intention can also be realized.

[0167] When the judgment result of S209 is no or the judgment result of S210 is no, the electronic device cannot perform a simulated click operation on the first control. At this time, the electronic device may issue a prompt message, such as S215.

[0168] S215. Send a prompt message.

[0169] This prompt message is used to prompt the user that the first control is not clickable. In this way, the user can be reminded of the response result of the electronic device to the first voice.

[0170] In addition, there are many forms of controls displayed on the electronic device, and some control groups may include more than three controls. In some other embodiments, when the determination result in S209 above is yes, it is possible that there are more than two second controls belonging to the same control group as the first control. In this case, the electronic device cannot determine which of the peer controls needs to perform a simulated click operation. If one of the peer controls is selected to perform a simulated click operation, it may not conform to the user's intention. Therefore, in some embodiments, after S209 and before S210, the method may further include: determining whether the number of second controls is 1. In this embodiment, when the electronic device determines that the number of second controls is 1, it then executes S210 and the subsequent processes. That is, only when the first control includes only one second control, will it be determined whether the second control is clickable. And, when the second control is clickable, the electronic device will perform a simulated click operation based on the second control.

[0171] In some other embodiments, if the electronic device determines that there are more than two second controls belonging to the same control group as the first control, the electronic device may not respond to the first voice. Further, the electronic device may send a prompt message, such as Figure 3 shown in S215.

[0172] In the technical solution proposed in the embodiments of the present application, when the first control matching the first voice is not clickable, and it is found that there is only one second control belonging to the same control group as the first control, a simulated click operation is then performed based on the second control. In this way, it can be ensured that the electronic device performs a simulated click operation at an accurate position, and the execution result is more in line with the user's intention.

[0173] In some scenarios, the electronic device can display multiple windows simultaneously in a stacked manner. In the scenario where the electronic device performs a simulated click operation in response to the user's voice, the last clicked position may be blocked by a window such as a floating window. If this position is blocked, then if the electronic device performs a simulated click operation, it may click on other windows, causing a problem of incorrect simulated clicks.

[0174] Such as Figure 5AAs shown, the mobile phone 100 displays a settings interface 401 and a calculator floating window 402 displayed as a floating window. The settings interface 401 includes a WLAN control 403, a Bluetooth control 404, and a mobile network control 405. Among them, some controls on the settings interface 401 are blocked by the calculator floating window 402. Specifically, the WLAN control 403 in the settings interface 401 is completely blocked; the Bluetooth control 404 is partially blocked; the mobile network control 405 is not blocked.

[0175] For Figure 5A the WLAN control 403 and the Bluetooth control 404 in the shown settings interface 401, if the electronic device performs a simulated click operation on these two positions, it may click on the calculator floating window 402. As Figure 5B shown, when the user inputs the voice "Bluetooth" and the mobile phone 100 executes the instruction corresponding to this language, it performs a simulated click operation at the position 404a. At this time, the mobile phone 100 will click on the control corresponding to the number "0" in the calculator floating window 402. Thus, the calculator floating window 402 of the mobile phone 100 is updated to 402a, which displays the selected number "0". This will cause a problem that the electronic device simulates a click error.

[0176] Or, Figure 4 the picture control 303 shown may also be blocked by a floating window or the like, resulting in the electronic device being unable to accurately perform a simulated click operation on the picture control 303.

[0177] Based on this, an embodiment of the present application further proposes a voice control method, which can also be applied to an electronic device that supports voice input. Figures 6 - 8 Shows the specific implementation manner of this voice control method.

[0178] Figure 6 Is a flowchart of the voice control method. In this embodiment, the method includes:

[0179] S500. Display a first interface, and the first interface includes a first switch.

[0180] S501. Receive an operation to turn on the first switch.

[0181] S502. In response to the operation of turning on the first switch, enable the voice control function.

[0182] S503. Display a second interface.

[0183] S504. Receive a first voice.

[0184] S505. Determine the user intention corresponding to the first voice.

[0185] S506. Determine whether the user's intention is a control click intention.

[0186] S507. Search for the control to be clicked that matches the first voice in the second interface.

[0187] In an embodiment where the electronic device simultaneously displays multiple windows, the second interface is the interface displayed by the focused window.

[0188] S508. Determine whether the control to be clicked is blocked.

[0189] The controls displayed by the electronic device on the display interface may be blocked by windows such as floating windows and pop-up windows, resulting in the electronic device being unable to perform a simulated click operation on the control. Combining Figure 5A As shown in the example, the controls displayed by the electronic device may be blocked by a floating window. After the electronic device determines the control to be clicked that matches the first voice, it can obtain whether there is a floating window on the electronic device.

[0190] In some other embodiments, the electronic device can also analyze the display content of the current display interface to determine whether the control to be clicked is blocked.

[0191] Combining Figure 5A As shown in the example, there are two cases where the control is blocked. One case is being completely blocked, and the other is being partially blocked. It can be understood that when the control is completely blocked, the electronic device cannot perform a simulated click operation on the control. When the control is partially blocked, the electronic device can perform a simulated click operation on the control, but there is a possibility of clicking on other windows, resulting in an incorrect simulated click. Therefore, the electronic device can also determine whether the control to be clicked is completely blocked, such as S509.

[0192] S509. Determine whether the control to be clicked is completely blocked.

[0193] The control being completely blocked means that the area corresponding to the control is completely covered by other windows. In Figure 5A As shown in the example, the WLAN control 403 is completely blocked; the Bluetooth control 404 is not completely blocked (i.e., partially blocked).

[0194] If the judgment result of S509 is yes, it means that the control to be clicked is completely blocked. In this case, the electronic device cannot perform a simulated click operation on the control to be clicked. At this time, the electronic device can execute S510.

[0195] S510. Send a prompt message.

[0196] This prompt message is used to prompt the user that the control to be clicked cannot be clicked. In this way, the user can be reminded of the response result to the first voice.

[0197] If the result of the determination in S509 is negative, it indicates that the control to be clicked is not completely blocked. In this case, the electronic device can perform a simulated click operation in the unblocked area of the control to be clicked. To avoid errors in the simulated click, the electronic device can perform the simulated click operation in the unblocked area.

[0198] S511. Redetermine the simulated click area.

[0199] Since the control to be clicked is not completely blocked, it means that there is still a part of the area of the control to be clicked where the simulated click operation can be performed. In some embodiments, the above S511 may specifically include: obtaining the area of the unblocked part of the control to be clicked as the simulated click area.

[0200] In one example, the above obtaining the area of the unblocked part of the control to be clicked may specifically include: determining the overlapping area between the occlusion window and the control to be clicked. Comparing the area corresponding to the control to be clicked with the overlapping area to determine the area of the unblocked part of the control to be clicked. In this way, the simulated click area can be quickly determined.

[0201] When determining the simulated click area, the border of the simulated click area can be determined. Taking the control to be clicked and the occlusion window as rectangles as an example, the determined simulated click area should also be a rectangle. In this embodiment, to determine the simulated click area, the coordinates of two corner points on a diagonal of the simulated click area can be specifically determined. Exemplarily, the above S511 can specifically determine the coordinates of the upper left corner point and the lower right corner point of the simulated click area, or determine the coordinates of the upper right corner point and the lower left corner point of the simulated click area. In other embodiments, the electronic device can also determine the border of the simulated click area by other means.

[0202] S512. Perform a simulated click operation based on the simulated click area.

[0203] In some embodiments, the above S512 may specifically include: obtaining the central position of the simulated click area and performing a simulated click operation at the central position of the simulated click area.

[0204] Taking the example that the electronic device performs a simulated click operation at the center position of the area corresponding to the control to be clicked, in the above embodiment, when the judgment result of S509 is no, S511 and S512 are directly executed. However, in actual situations, one case where the judgment result of S509 is no is that the control to be clicked is partially blocked, and the center position of the area corresponding to the control to be clicked is not blocked. At this time, the electronic device does not need to re-determine the simulated click area, but can still directly perform the simulated click operation at the center position of the area corresponding to the control to be clicked. Since the center position of the area corresponding to the control to be clicked is not blocked, when the electronic device performs the simulated click operation, there will be no problem of simulated click error.

[0205] Therefore, in some other embodiments, when the judgment result of S509 is no, before S511 and S512, the above method may further include: determining whether the center position of the control to be clicked is blocked. In this embodiment, the electronic device may execute S511 and S512 when the center position of the control to be clicked is blocked. In another example, when the judgment result of S509 is no and the center position of the control to be clicked is not blocked, the electronic device may directly perform the simulated click operation at the center position of the control to be clicked.

[0206] In addition, if the judgment result of S508 is no, it means that the control to be clicked is not blocked. In this case, the electronic device may directly perform a simulated click on the control to be clicked, such as S513 and S514.

[0207] S513. Obtain the area corresponding to the control to be clicked.

[0208] S514. Perform a simulated click operation in the area corresponding to the control to be clicked.

[0209] It should be noted that for the specific implementation processes of some steps in the above S501 - S514, reference may be made to the descriptions of the corresponding steps of S201 - S214 in the above embodiments; details are not repeated here.

[0210] In the technical solution proposed in the embodiments of the present application, if the electronic device detects that the control to be clicked that matches the first voice is completely blocked, the simulated click operation is no longer performed, avoiding the problem of incorrect simulated click caused by clicking on other windows during the simulated click operation. And if the electronic device detects that the control to be clicked is not completely blocked, the simulated click area is re-determined and then the simulated click area is executed. In this way, it can be ensured that when performing the simulated click, the control to be clicked is accurately clicked, avoiding the problem of incorrect simulated click in this scenario. Thus, the electronic device can perform the simulated click operation at an accurate position in response to the voice containing the control click intention input by the user, thereby accurately realizing the user's intention.

[0211] Taking the floating window that may block the control to be clicked as an example, as Figure 7A shown, the above S508 may include S601 - S604:

[0212] S601. Query whether there is a floating window.

[0213] In some embodiments, the application corresponding to the voice control function may register a floating window monitor. When the electronic device displays a floating window, the application corresponding to the voice control function may be notified through this floating window monitor. In this way, when the application corresponding to the voice control function of the electronic device needs to perform a simulated click operation in response to the voice input by the user, it can query whether there is a floating window in the current display interface of the electronic device.

[0214] In this embodiment, if the judgment result of S601 is negative, it means that there is no floating window on the electronic device currently. That is to say, the control to be clicked is not blocked by the floating window. In this case, the electronic device may execute S513 and S514.

[0215] If the judgment result of S601 is positive, it means that there is a floating window. Next, the electronic device may determine whether the control to be clicked is blocked by the floating window according to the area corresponding to the floating window and the area corresponding to the control to be clicked, that is, S602 - S604.

[0216] S602. Determine whether the floating window is the focus window.

[0217] In some embodiments, the electronic device may obtain the window identifier of the current focus window, and then determine whether the floating window is the focus window according to the window identifier of this focus window.

[0218] In some examples, the window identifier of the focus window may also be the border of this focus window. In this embodiment, the electronic device may compare the border of the floating window with the border of the focus window to determine whether the floating window is the focus window. In this way, it is convenient to confirm whether the control to be clicked is blocked by the floating window.

[0219] In some other examples, the window identifier of the focus window may specifically be the application name of the application displayed in this focus window or the activity name of the activity. In this embodiment, the electronic device may compare the application name or activity name corresponding to the floating window with the focus window to determine whether the floating window is the focus window. In this way, it is convenient to confirm whether the control to be clicked is blocked by the floating window.

[0220] If the floating window is the focused window, it means that the control to be clicked is a control on the floating window. In this case, the electronic device can directly perform a simulated click operation on the control to be clicked. That is, if the judgment result of S602 is yes, the electronic device can execute S513 and S514.

[0221] If the floating window is not the focused window, it means that the control to be clicked is not a control on the floating window, and the control to be clicked may be blocked by the floating window. In this case, the electronic device needs to further determine whether the control to be clicked is blocked by the floating window. Exemplarily, the electronic device can execute S603 and S604.

[0222] S603. Obtain area 1 corresponding to the control to be clicked and area 2 corresponding to the floating window.

[0223] For the specific implementation of obtaining the area corresponding to the control and the area corresponding to the floating window, reference can be made to the descriptions in the above embodiments and related technologies.

[0224] S604. Compare area 1 and area 2 to determine whether the control to be clicked is blocked by the floating window.

[0225] In this embodiment, the display interface of the electronic device includes a floating window. The electronic device can determine whether the control to be clicked is blocked by the floating window according to whether there is an overlapping area between the area corresponding to the floating window and the area corresponding to the control to be clicked. The floating window is usually floating and displayed on the background interface. If there is an overlapping area between the area corresponding to the floating window and the area corresponding to the control to be clicked, it means that the control to be clicked is blocked by the floating window.

[0226] If the judgment result of S604 is no, it means that although there is a floating window on the current display interface of the electronic device, the control to be clicked is not blocked by the floating window. In this case, the electronic device can also directly perform a simulated click operation on the target text, that is, execute S513 and S514.

[0227] If the judgment result of S604 is yes, the electronic device can continue to execute S509 to determine whether the control to be clicked is completely blocked.

[0228] Figure 7B This is a schematic diagram of a scenario example of the voice control method according to the embodiment of the present application. The mobile phone 100 displays the setting interface 401. In response to the voice "Bluetooth" input by the user, the mobile phone 100 performs a simulated click operation at position 402b. After that, the mobile phone 100 opens the Bluetooth sub-menu interface 402c. In this way, the user's intention is accurately realized, and the problem of incorrect simulated clicks is avoided.

[0229] Figure 7CSchematic diagram of a scenario example of the voice control method according to an embodiment of the present application. The mobile phone 100 displays a setting interface 401. In response to the voice "WLAN" input by the user, the mobile phone 100 detects that the WLAN control 403 is completely blocked by suspension. Therefore, the mobile phone 100 can display a prompt message 403a.

[0230] Figure 8 Shows Figure 5A In the example of, the positional relationship between the partially blocked Bluetooth control 404 and the calculator floating window 402. In some embodiments, in combination with Figure 8 Taking the to-be-clicked control and the floating window as rectangles as an example, the above S511 can be specifically implemented in the following manner: In this embodiment, the upper left corner of the screen is used as the origin coordinate (0, 0); the positive direction of the X-axis extends to the right from the origin, and the positive direction of the Y-axis extends downward from the origin, and the established coordinate system is denoted as the screen coordinate system.

[0231] In the quadrant where x > 0 and y > 0, the area corresponding to the to-be-clicked control (i.e., the Bluetooth control 404) is represented by coordinates: the upper left corner point A(x1, y1), and the lower right corner point B(x2, y2). The area corresponding to the floating window (i.e., the calculator floating window 402) is represented by coordinates: the upper left corner point C(x3, y3), and the lower right corner point D(x4, y4).

[0232] There is an overlapping part (the shaded area 404-1 in Figure 8 ) between the area corresponding to the to-be-clicked control and the area corresponding to the floating window, and the to-be-clicked control is not completely blocked. The simulated click area to be determined is the coordinates of the upper left corner point and the lower right corner point of the largest rectangle area where the to-be-clicked control is not blocked. In some examples, the simulated click area can be represented by coordinates:

[0233] The upper left corner point M(max(x1, x3), max(y1, y3)), and the coordinates of the lower right corner point N(min(x2, x4), min(y2, y4)). In Figure 8 the example shown, the calculated simulated click area (i.e., the area 404-2 in Figure 8 ) has coordinates M(x1, y4), N(x2, y2).

[0234] It should be noted that Figure 8 the method for determining the simulated click area shown is only one example. In other embodiments, the electronic device can also determine the simulated click area by other means.

[0235] In the technical solution proposed in the embodiments of the present application, the electronic device determines whether there is a floating window in the current display interface, whether the floating window is the focus window, and in the case where the floating window is not the focus window, compares the area corresponding to the floating window with the area corresponding to the control to be clicked, so as to determine whether the control to be clicked is blocked by the floating window. In this way, it can quickly and conveniently determine whether the control to be clicked is blocked, which is convenient for the electronic device to perform a simulated click operation in response to the first voice.

[0236] Figure 9A The figure shows the flow of another voice control method proposed in the embodiments of the present application.

[0237] In Figure 9A In the shown embodiment, after the electronic device responds to the operation of the user to turn on the first switch, the voice control function is enabled. When the electronic device determines that the user intention corresponding to the received voice is a control click intention, it can first determine whether the first control is clickable. The process of the electronic device determining whether the first control is clickable can refer to the processes of S207 - S211 and S215 in the Figure 3 shown method. When it is determined that the first control is clickable, or there is a second clickable control in the first control, it can continue to determine whether the first control (or the second control) is blocked, that is, Figure 6 the processes of S506 - S510 in the shown method. Finally, if the first control (or the second control) (i.e., the control to be clicked) is not blocked, or not completely blocked, the electronic device can perform a simulated click operation in response to the first voice, that is, Figure 6 the processes of S511 - S512 or S513 - S514 in the shown process. In another example, when the first control is not clickable and there is no sibling control, or the first control (or the second control) is completely blocked, the electronic device can send a prompt message to prompt that it is not clickable.

[0238] In the technical solution proposed in the embodiments of the present application, when receiving a voice including a control click intention, the electronic device first determines whether the first control matching the voice is clickable, and then determines whether it is blocked. In this way, when the voice intention is accurately recognized, the possibility that the electronic device cannot implement the true intention of the voice input by the user can be reduced, and the possibility of performing an accurate simulated click operation in response to the voice input by the user can be improved.

[0239] In some other embodiments, in Figure 9AIn the voice control method shown, when it is determined that the first control matching the first voice is clickable, the electronic device will perform an occlusion judgment on the first control. Further, if the electronic device determines that the first control is completely occluded, it means that the electronic device cannot perform a simulated click operation of the voice click instruction for the first control. In this case, the electronic device can also re-determine the control to be clicked. Specifically, the electronic device can find a second control belonging to the same control group as the first control and determine the second control as the new control to be clicked. Then, the electronic device determines whether the second control is occluded by the floating window.

[0240] If the second control is also completely occluded, it means that the electronic device cannot respond to the first voice to perform a simulated click operation.

[0241] If the second control is not occluded, it means that the electronic device can perform a simulated click operation on the second control in response to the first voice.

[0242] If the second control is partially occluded, the electronic device can re-determine the simulated click area based on the areas corresponding to the floating window and the second control respectively. Finally, the electronic device performs a simulated click operation within the re-determined simulated click area.

[0243] In this embodiment, the specific implementation process of the electronic device finding the second control belonging to the same control group as the first control can refer to the description in the above embodiment.

[0244] Take Figure 9B as an example. The mobile phone 100 displays the settings interface 406, and also includes a calculator floating window 407. In this example, the text control 408 in the WLAN settings option is the above-mentioned first control. When the mobile phone 100 judges the click attribute of the text control 408, it determines that the click attribute of the text control 408 is clickable. Then, the mobile phone 100 performs an occlusion judgment on the text control 408. It can be determined that the text control 408 is completely occluded by the calculator floating window 407. In this case, the mobile phone 100 can find a second control belonging to the same control group as the text control 408, such as Figure 9B the picture control 409 in the WLAN settings option shown. Then, the mobile phone 100 performs an occlusion judgment on the picture control 409. In one example, as Figure 9B shown, the picture control 409 is not occluded by the floating window. At this time, the mobile phone 100 can perform a simulated click operation on the picture control 409, and can also achieve the user's intention to open the WLAN settings option and display the WLAN settings sub-menu.

[0245] It can be understood that in another example, if the result obtained by the mobile phone 100 after performing the occlusion judgment on the picture control 409 is that the picture control 409 is partially occluded, then the mobile phone 100 can re-determine the simulated click area based on the areas corresponding to the calculator floating window 407 and the picture control 409 respectively, and then perform the simulated click operation on the re-determined simulated click area.

[0246] In another example, if the result obtained by the mobile phone 100 after performing the occlusion judgment on the picture control 409 is that the picture control 409 is completely occluded, then the mobile phone 100 can send a prompt message. This prompt message is used to prompt that it is not clickable.

[0247] In the technical solution proposed in the embodiments of the present application, when the electronic device performs the simulated click operation in response to the voice, it simultaneously considers whether the control to be clicked is clickable and whether it is occluded. In this way, it can be further ensured that the control to be clicked can execute the click event. Thus, when the electronic device accurately recognizes the voice intention, the possibility that the electronic device cannot realize the true intention of the voice input by the user is reduced, and the possibility of performing the accurate simulated click operation in response to the voice input by the user is increased.

[0248] Figures 10 - 13 Shows the implementation Figures 3 - 4 The software architecture of the electronic device implementing the voice control method shown, and the interaction process of each module of the electronic device when implementing the voice control method.

[0249] Figure 10 Is the software architecture of the electronic device in some embodiments. In some embodiments, the software system of the electronic device can adopt a layered architecture, an event-driven architecture, a microkernel architecture or a cloud architecture. The embodiments of the present application take the layered architecture System as an example to exemplarily illustrate the software structure of the electronic device.

[0250] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the System is divided into four layers, from top to bottom are the application layer, the application framework layer, the Android runtime and the system library, and the kernel layer.

[0251] The application layer may include a series of application packages (APKs). For example, applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc. In the embodiments of the present application, the application layer includes a voice control APK for providing the voice control function of the electronic device. The voice control APK includes an interface monitoring module, an interface content acquisition module (content sensor), an interface parsing module (view fetcher), and an application interaction module, etc. Among them, the interface monitoring module is used to monitor processes such as the startup, exit, and switching of the application interface. The interface content acquisition module is used to acquire the top-level activity. The interface parsing module is used to parse the top-level activity to obtain the content of the current display page (including pictures and / or text). The application interaction module is used to manage the process of the application executing instructions according to voice.

[0252] The application layer further includes a voice processing engine for parsing, recognizing, and processing voices. Among them, the voice parsing module is used to convert voice into text; the voice parsing module may belong to the ASR engine. The voice recognition module is used to understand and recognize the semantics of the text and determine whether it matches the interface text; the voice recognition module may belong to the NLU engine. The instruction mapping module is used to convert the recognized semantics into machine-executable instructions; the instruction mapping module may belong to the DM engine.

[0253] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0254] As Figure 11 shown, the application framework layer may include a window manager, a content provider, a view system, a resource manager, an activity manager service (AMS), a package manager service (PMS), and a multi-mode control module, etc.

[0255] AMS is mainly responsible for the startup, switching, scheduling of the four major components in the system and the management and scheduling of application processes, etc. Its responsibilities are similar to the process management and scheduling module in the operating system. When a process startup or component startup is initiated, the request will be passed to AMS through the inter-process communication (binder) mechanism, and AMS will then make unified processing.

[0256] The multi-mode control module is used to manage the execution of voice instructions. For example, sending the instruction to the application to make the application execute the instruction.

[0257] The system library can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.

[0258] Android runtime is responsible for the scheduling and management of the Android system. Android runtime includes a core library and a virtual machine.

[0259] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0260] The kernel layer is the layer between hardware and software. The kernel layer can include display drivers, sensor drivers, microphone drivers, Wi-Fi drivers, etc.

[0261] Combined Figure 10 , Figure 11 shows an interaction diagram of each module of the electronic device when implementing the voice control method provided in the embodiment of the present application.

[0262] AMS can monitor the lifecycle of the activity corresponding to each interface of each application on the electronic device.

[0263] In some examples, the interface content acquisition module can send a request to AMS. In response to this request, AMS returns the top-level activity to the interface content acquisition module. In some embodiments, the interface content acquisition module sending a request to AMS can specifically include: the interface content acquisition module sending a request to AMS through activitymanagerEx.requestContentNode.

[0264] In other examples, the interface content acquisition module can register a listener with AMS. When AMS senses a change in the lifecycle of the application, it can notify the top-level activity to the interface content acquisition module through a callback.

[0265] After receiving the top-level activity sent by AMS, the interface content acquisition module sends the top-level activity to the interface parsing module. The interface parsing module parses the top-level activity to obtain the display content of the electronic device, and the text controls and / or picture controls included therein.

[0266] Further, the interface parsing module may save the text controls and / or image controls obtained by parsing the top-level activity to a cache. In some embodiments, the top-level activity obtained by the interface parsing module may include: the control information of each control group. Among them, the control information of the control group may specifically include: the control information of each control included in the control group. Taking a control group including an image control and a text control as an example, the control information of the control group may include: control group [coordinate 1, coordinate 2]: image control [coordinate 3, coordinate 4], text control [coordinate 5, coordinate 6].

[0267] In one example, the control information of the image control in the control group may at least include the content of Table 1.

[0268] Table 1

[0269]

[0270] The control information of the text control in the control group may at least include the content of Table 2.

[0271] Table 2

[0272]

[0273] Among them, the label is used to indicate the serial number of the control in the control group.

[0274] The interface parsing content parses the above top-level activity, and the control information of each control or control group on the currently displayed interface can be obtained. In some embodiments, the interface parsing module may save the control information of the text control and / or image control to a cache. Among them, the control information may include: the area where the control is located (a rectangular control can be represented by the coordinates of the upper left corner point and the lower right corner point), whether the control is clickable, and the control group to which the control belongs.

[0275] After the user inputs voice to the electronic device, the voice is collected by the microphone of the electronic device. The microphone inputs the collected voice to the application interaction module of the voice control APK. The application interaction module is responsible for distributing the received voice; in some examples, the application interaction module may distribute the voice to the voice parsing module in the voice processing engine, and the voice parsing module parses the voice input by the user.

[0276] The voice parsing module parses the voice, and the parsed text corresponding to the voice can be obtained. Then, the voice parsing module sends the text corresponding to the parsed voice to the voice recognition module.

[0277] The speech recognition module performs semantic understanding and recognition on the text corresponding to the speech to obtain the intent corresponding to the speech; that is, the user's intent, which is the operation that the user hopes to control the electronic device to perform through the speech. After that, the speech recognition module can send the intent corresponding to the speech to the instruction mapping module.

[0278] The instruction mapping module maps and converts the intent corresponding to the speech into an instruction executable by the machine (denoted as instruction a).

[0279] In some embodiments, after converting the user's intent into instruction a, the instruction mapping module can determine whether instruction a is a control click instruction. If the obtained instruction a is not a control click instruction, the instruction mapping module can directly send instruction a to the corresponding application through the multimode control module to notify the application to execute instruction a.

[0280] If the obtained instruction a is a control click instruction, the instruction mapping module can send the control click instruction to the interface parsing module. Among them, the control click instruction carries the text that matches the parsed text corresponding to the first speech (i.e., the hit text).

[0281] In other embodiments, after receiving the user's intent sent by the speech recognition module, the instruction mapping module can also judge the user's intent to determine whether the user's intent is a control click intent. And, when the instruction mapping module determines that the user's intent is a control click intent, after converting the user's intent into instruction a, it can send instruction a to the interface parsing module. If it is determined that the user's intent is not a control click intent, after converting the user's intent into instruction a, the instruction mapping module can directly send it to the corresponding application through the multimode control module to notify the application to execute instruction a.

[0282] Taking the user input speech as a control click intent, that is, instruction a is a control click instruction as an example for illustration. In the embodiments of the present application, after receiving the control click instruction sent by the instruction mapping module, the interface parsing module can, according to the hit text carried in the control click instruction, search in the cache for the control information of the control corresponding to the hit text (i.e., the first control, hereinafter denoted as the target text control).

[0283] In the embodiments of the present application, the interface parsing module obtains the control information of the target text control from the cache and judges whether the target text control is clickable. In some examples, if the target text control is not clickable, the interface parsing module can search in the cache for a second control (hereinafter denoted as the target sibling control) that belongs to the same control group as the target text control.

[0284] If the interface parsing module finds the target sibling control corresponding to the target text control, it can also query the number of target sibling controls, the area corresponding to the target sibling control, and whether the target sibling control is clickable.

[0285] In some embodiments, when the interface parsing module determines that there is a target sibling control belonging to the same control group as the target text control and it is clickable, it can execute instruction a based on the area corresponding to the found target sibling control. In another example, if the interface parsing module determines that there is no target sibling control belonging to the same control group as the target text control, or all sibling controls belonging to the same control group as the target text control are not clickable, it determines that the control click instruction cannot be executed.

[0286] In some other embodiments, when the interface parsing module determines that there is a target sibling control belonging to the same control group as the target text control, there is only one target sibling control, and it is clickable, it can execute instruction a based on the area corresponding to the found target sibling control. In this embodiment, if the interface parsing module determines that there are multiple sibling controls belonging to the same control group as the target text control, or there is only one sibling control belonging to the same control group as the target text control and it is not clickable, it determines that instruction a cannot be executed.

[0287] Further, after the interface parsing module determines that instruction a can be executed, it can notify the corresponding application to execute instruction a through the multi-mode control module.

[0288] After the interface parsing module determines that the control click instruction cannot be executed, it can send a notification message to the application interaction module. In response to receiving this notification message, the application interaction module can issue a prompt message indicating non-clickability.

[0289] The following combines Figure 12 the module interaction to introduce the specific process of the interface parsing module for determining whether the target text control is clickable. In this embodiment, the area corresponding to the control is the border of the control.

[0290] In this embodiment, the interface parsing module includes a clickability judgment sub-module and a sibling control query sub-module. Among them, the clickability judgment sub-module is used to query and obtain whether the target text control is clickable from the cache after receiving the control click instruction sent by the instruction mapping module. The sibling control query sub-module can be used to query the control information corresponding to the target text control.

[0291] If the target text control is clickable, the interface parsing module sends a control click instruction to the multi-module control module through the clickability judgment sub-module, and this control click instruction carries the border of the target text control; that is Figure 12 the first case shown in

[0292] If the target text control is not clickable, the clickability judgment module sends a notification message to the sibling control query sub-module. In response to this notification message, the sibling control query sub-module queries and obtains from the cache the sibling controls that belong to the same control group as the target text control, as well as the click attributes (i.e., whether they are clickable) and borders of the sibling controls.

[0293] Based on the query results, the sibling control query sub-module can determine whether there are sibling controls for the target text control.

[0294] If there is one sibling control for the target text control and this sibling control is clickable, the interface parsing module sends a control click instruction to the multi-module control module through the clickability judgment sub-module, and this control click instruction carries the border of the sibling control; that is Figure 12 the second case shown.

[0295] If the target text control has no sibling controls, the interface parsing module returns a non-clickable notification message to the application interaction module through the clickability judgment sub-module; that is Figure 12 the third case shown.

[0296] If the target text control has no sibling controls, the interface parsing module returns a non-clickable notification message to the application interaction module through the clickability judgment sub-module; that is Figure 12 the fourth case shown.

[0297] If the target text control has multiple sibling controls, the interface parsing module returns a non-clickable notification message to the application interaction module through the clickability judgment sub-module; that is Figure 12 the fifth case shown.

[0298] In response to receiving the notification message sent by the clickability judgment sub-module, the application interaction module can issue a prompt message.

[0299] Figure 13 Shows the timing interaction diagram among the modules during the process of the voice control method proposed in the embodiment of the present application. In this embodiment, taking the first voice input by the user containing the control click intention as an example for illustration. The specific implementation manners of the instructions corresponding to the voice input by the user in other embodiments being other instructions can be referred to Figure 13 the process shown and will not be elaborated in the embodiments of the present application.

[0300] The user turns on the first switch. In response to receiving the operation of the user turning on the first switch, the electronic device enables the voice control function.

[0301] The interface content acquisition module registers a listener with the AMS. It should be noted that the interface content acquisition module can register the listener when the electronic device enables the voice control function. When the AMS senses a change in the application's life cycle, it can notify the top-level activity to the interface content acquisition module through a callback.

[0302] Then, the interface content acquisition module can send the top-level activity to the interface parsing module, notifying the interface parsing module to perform parsing. The interface parsing module parses the top-level activity to obtain the controls and control information of the currently displayed interface. Then, it saves the control information to the cache. Moreover, the interface parsing module sends the text of the currently displayed interface, that is, the interface hot words, to the speech recognition module.

[0303] The user inputs a first voice to the electronic device. Specifically, the microphone collects the voice input by the user. Since the current voice control function is enabled and the recording channel corresponding to the first voice control function is opened, the microphone sends the first voice to the application interaction module of the voice control APK.

[0304] The application interaction module distributes the received first voice to the speech processing engine for processing. Specifically, the application interaction module sends it to the speech parsing module in the speech processing engine. After receiving the first voice, the speech parsing module can parse the first voice to obtain the parsed text. Then, the speech parsing module sends the parsed text to the speech recognition module for intent recognition. The speech recognition module recognizes the intent of the received parsed text to obtain the user intent corresponding to the parsed text. The user intent recognized by the speech recognition module is a control click intent. After that, the speech recognition module can send the control click intent to the instruction mapping module for instruction mapping. After receiving the control click intent, the instruction mapping module maps the control click intent to an instruction that the electronic device can execute, that is, a control click instruction.

[0305] Then, the instruction mapping module sends the control click instruction to the interface parsing module, and the control click instruction carries the hit text.

[0306] After receiving the control click instruction, the interface parsing module sends a query request to the cache. The cache returns the click attributes and borders of the target text control corresponding to the hit text to the interface parsing module. Then, the interface parsing module determines whether the target text control is clickable.

[0307] If the target text control is clickable, the interface parsing module can send the control click instruction to the multi-mode control module, and the control click instruction carries the border of the target text control. The multi-mode control module can notify the corresponding application to execute the control click instruction, that is, perform a simulated click operation within the border of the target text control.

[0308] If the target text control is not clickable, the interface parsing module can cache and send a query request. The cache returns sibling controls, click attributes, and borders to the interface parsing module. Then, the interface parsing module determines whether there are sibling controls. If there are sibling controls (i.e., the above-mentioned target sibling controls), the interface parsing module continues to determine whether the number of sibling controls is 1. If the number of sibling controls is 1, the interface parsing module continues to determine whether the sibling control is clickable. Finally, if the only sibling control is clickable, the interface can send a control click instruction to the multimode control module; the control click instruction carries the border of the sibling control. The multimode control module can notify the corresponding application to execute the control click instruction.

[0309] If the target text control has no sibling controls, or the number of target sibling controls is greater than 1, or the number of target sibling controls is 1 and the sibling control is not clickable, the interface parsing module notifies the application interaction module that it is not clickable. The application interaction module can issue a prompt message, which is used to prompt that it is not clickable.

[0310] Figures 14 - 17 Shows the implementation Figures 6 - 8 The software architecture of an electronic device implementing the voice control method shown, and the interaction process of each module of the electronic device when implementing the voice control method.

[0311] Figure 14 It is the software architecture of an electronic device in some embodiments. In this embodiment, the voice control APK of the electronic device further includes a floating window monitoring module and an occlusion judgment module. The floating window monitoring module is used to monitor whether the electronic device displays a floating window and obtain the border information of the floating window. The occlusion judgment module is used to judge whether the control to be clicked is occluded.

[0312] Combined with Figure 14 , Figure 15 Shows the interaction schematic diagram of each module of the electronic device when implementing the voice control method provided in the embodiments of the present application.

[0313] The floating window monitoring module registers a floating window monitor with the AMS. After the AMS detects that the electronic device displays a floating window, it notifies the floating window monitoring module of the border of the floating window through a callback. In some examples, the floating window monitoring module registers a floating window monitor with the AMS when the electronic device enables the voice control function.

[0314] The specific implementation process in which the user inputs voice, and the application interaction module of the voice control APK distributes the voice to the voice processing engine, obtains a control click instruction, and returns the control click instruction to the interface parsing module refers to Figure 11 the description process of the corresponding steps in

[0315] In this embodiment, after the interface parsing module queries the border of the control to be clicked that matches the hit text, it sends the border of the control to be clicked to the occlusion judgment module.

[0316] The occlusion judgment module can query whether there is a floating window on the electronic device through the floating window monitoring module. At the same time, the occlusion judgment module can also query the current focused window through AMS. Then, the occlusion judgment module can judge whether there is a floating window, and when there is a floating window, it can judge whether the control to be clicked is completely occluded in combination with the current focused window. In the case where there is a floating window and the control to be clicked is completely occluded, the occlusion judgment module notifies the application interaction module that it cannot be clicked. The application interaction module can send a prompt message, which is used to prompt that it cannot be clicked. In other cases, the occlusion judgment module sends a control click instruction to the multimode control module, and the control click instruction carries the simulated click area. The multimode control module can notify the corresponding application to execute the control click instruction, that is, perform a simulated click operation in the simulated click area.

[0317] Figure 16 It shows the specific process of the occlusion judgment module judging whether there is a floating window, and when there is a floating window, judging whether the control to be clicked is completely occluded in combination with the current focused window. In this embodiment, the occlusion judgment module includes a border comparison sub-module and a border calculation sub-module. Among them, the border comparison sub-module is used to determine whether there is a floating window, and when there is a floating window, compare the control to be clicked and the floating window to determine whether the control to be clicked is completely occluded by the floating window. The border calculation sub-module is used to re-determine the simulated click area when the control to be clicked is not completely occluded.

[0318] After receiving the border of the control to be clicked sent by the interface parsing module, the border comparison sub-module obtains the border of the floating window from the floating window monitoring module. In one example, if there is no floating window currently, the floating window monitoring module returns null to the border comparison sub-module.

[0319] If the border comparison sub-module obtains null from the floating window monitoring module, it is the first case where there is no floating window. The border sub-module can send a control click instruction to the multimode control module, and the control click instruction carries the border of the control to be clicked.

[0320] After the border comparison sub-module obtains the border of the floating window, it obtains the focused window from AMS and judges whether the floating window is the focused window. If there is a floating window and the floating window is the focused window, it means that the control to be clicked is a control in the floating window, and there is no possibility that the control to be clicked is occluded by the floating window.

[0321] In some other examples, if there is a floating window and the floating window is not the focus window, there is a possibility that the control to be clicked is blocked by the floating window. At this time, the border comparison sub-module can compare the borders of the floating window and the control to be clicked to determine whether the control to be clicked is blocked by the floating window.

[0322] If the border comparison sub-module determines that the control to be clicked is not blocked by the floating window, it is the second case. The border comparison sub-control can also send a control click instruction to the multi-mode control module, and the control click instruction carries the border of the control to be clicked.

[0323] If the border comparison sub-module determines that the control to be clicked is not completely blocked by the floating window, it is the third case. At this time, the border comparison sub-module can notify the border calculation sub-module to re-determine the simulated click area. Specifically, the border comparison sub-module sends the borders of the floating window and the control to be clicked to the border calculation sub-module. The border calculation sub-module calculates and determines the border of the simulated click area based on the borders of the floating window and the control to be clicked. After that, the border calculation sub-module can send a control click instruction to the multi-mode control module, and the control click instruction carries the border of the simulated click area.

[0324] If the border comparison sub-module determines that the control to be clicked is completely blocked by the floating window, it is the fourth case. In this case, the border comparison sub-module notifies the application interaction module that it cannot be clicked. The application interaction module can issue a prompt message.

[0325] Among them, in the above first, second, and third cases, after receiving the control click instruction, the multi-mode control module can notify the corresponding application to execute the control click instruction.

[0326] Figure 17 Shows the timing interaction diagram among the modules in the process of the voice control method proposed in the embodiment of the present application. In this embodiment, it is described by taking the first voice input by the user containing the control click intention as an example. In other embodiments, the specific implementation manner of the instruction corresponding to the voice input by the user being other instructions can refer to Figure 13 the shown process and will not be elaborated in the embodiment of the present application.

[0327] The user turns on the first switch. In response to receiving the operation of the user turning on the first switch, the electronic device enables the voice control function.

[0328] The interface content acquisition module registers a listener with the AMS. It should be noted that the interface content acquisition module can register a listener when the electronic device enables the voice control function. When the AMS senses the change in the application life cycle, it can notify the top-level activity to the interface content acquisition module through a callback.

[0329] Then, the interface content acquisition module can send the top-level activity to the interface parsing module, notifying the interface parsing module to perform parsing. The interface parsing module parses the top-level activity to obtain the controls of the currently displayed interface. Then, the control information is saved to the cache. Also, the interface parsing module sends the text of the currently displayed interface to the speech recognition module.

[0330] In addition, the floating window monitoring module registers a floating window monitor with the AMS, and the AMS can notify the floating window monitoring module of the floating window's border through a callback.

[0331] The user inputs a first voice to the electronic device. Specifically, the microphone collects the voice input by the user. Since the current voice control function is enabled and the recording channel corresponding to the first voice control function is opened, the microphone sends the first voice to the application interaction module of the voice control APK.

[0332] The application interaction module distributes the received first voice to the speech processing engine for processing. Specifically, the application interaction module sends it to the speech parsing module in the speech processing engine. After receiving the first voice, the speech parsing module can parse the first voice to obtain the parsed text. Then, the speech parsing module sends the parsed text to the speech recognition module for intent recognition. The speech recognition module recognizes the intent of the received parsed text to obtain the user intent corresponding to the parsed text. The user intent recognized by the speech recognition module is a control click intent. After that, the speech recognition module can send the control click intent to the instruction mapping module for instruction mapping. After receiving the control click intent, the instruction mapping module maps the control click intent to an instruction that the electronic device can execute, that is, a control click instruction.

[0333] Then, the instruction mapping module sends the control click instruction to the interface parsing module, and the control click instruction carries the hit text.

[0334] After receiving the control click instruction, the interface parsing module sends a query request to the cache. The cache returns the border of the control to be clicked that matches the hit text to the interface parsing module. Then, the interface parsing module sends the border of the control to be clicked to the occlusion judgment module.

[0335] The occlusion judgment module can send a query request to the floating window monitoring module. The floating window monitoring module returns the border of the floating window / empty to the occlusion judgment module. Then, the occlusion judgment module determines whether there is a floating window. If there is no floating window, the occlusion judgment module can directly send the control click instruction to the multi-mode control module, and the control click instruction carries the border of the control to be clicked.

[0336] If there is a floating window, the occlusion determination module sends a query request to the AMS. In response to the query request, the AMS returns the current focus window of the electronic device to the occlusion determination module. Thereafter, the occlusion determination module may first determine whether the floating window is the focus window. If the floating window is the focus window, the occlusion determination module directly sends a control click instruction to the multi-mode control module, and the control click instruction carries the border of the control to be clicked.

[0337] If the floating window is not the focus window, the occlusion determination module may continue to determine whether the control to be clicked is occluded by the floating window. If the control to be clicked is not occluded by the floating window, the occlusion determination module directly sends a control click instruction to the multi-mode control module, and the control click instruction carries the border of the control to be clicked.

[0338] If the control to be clicked is occluded by the floating window, the occlusion determination module may continue to determine whether the control to be clicked is completely occluded. If the control to be clicked is not completely occluded, the occlusion determination module combines the border of the floating window and the border of the control to be clicked to calculate a new simulated click area. Then, the occlusion determination module sends a control click instruction to the multi-mode control module, and the control click instruction carries the border of the new simulated click area.

[0339] If the control to be clicked is completely occluded, the occlusion determination module notifies the application interaction module that it cannot be clicked. The application interaction module may issue a prompt message for prompting that it cannot be clicked.

[0340] In some embodiments, an electronic device having Figure 14 and Figure 15 the software architecture shown may also be used to implement the voice control method shown in FIG. 9. In this embodiment, Figure 15 in the module interaction shown, the control to be clicked is the first control that matches the first voice determined by the interface parsing module, or the second control that belongs to the same control group as the first control.

[0341] It should be noted that in the embodiments of the present application, in the process of the electronic device recognizing the received voice and executing the instruction corresponding to the voice, there is the following corresponding relationship: The electronic device parses the display interface to obtain: text controls and picture controls; wherein, the text contained in the text controls is saved as the interface hot words; that is, the text of the text controls in the display interface corresponds one-to-one with the interface hot words of the display interface. The electronic device recognizes the voice input by the user to obtain: parsed text - user intention - user intention is converted into an instruction - the interface hot word is hit - the text control corresponding to the hit interface hot word (denoted as the first control). That is, the voice corresponds one-to-one with the parsed text, user intention, hit interface hot word, first control, and instruction. Some text controls and other controls jointly belong to a control group. There is a one-to-one correspondence between the text controls and other controls under the same control group.

[0342] Figure 18 Shows a schematic structural diagram of an electronic device 700 provided in an embodiment of the present application.

[0343] The electronic device 700 may include a processor 710, an external memory interface 720, an internal memory 721, a universal serial bus (USB) interface 730, a charging management module 740, a power management module 741, a battery 742, an antenna 1, an antenna 2, a mobile communication module 750, a wireless communication module 760, an audio module 770, a speaker 770A, a receiver 770B, a microphone 770C, a headphone interface 770D, a sensor module 780, a button 790, a motor 791, a camera 792, a display screen 793, and a subscriber identification module (SIM) card interface 794, etc. Among them, the sensor module 780 may include a pressure sensor 780A, a touch sensor 780B, etc.

[0344] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 700. In other embodiments of the present application, the electronic device 700 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0345] The processor 710 may include one or more processing units. For example, the processor 710 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. For example, the processor 710 is used to execute the voice control method in the embodiments of the present application.

[0346] Among them, the controller may be the nerve center and command center of the electronic device 700. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.

[0347] A memory can also be set in the processor 710 for storing instructions and data. In some embodiments, the memory in the processor 710 is a cache memory. This memory can save the instructions or data that the processor 710 has just used or recycled. If the processor 710 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 710, and thus improves the efficiency of the system.

[0348] The USB interface 730 is an interface that conforms to the USB standard specification. Specifically, it can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 730 can be used to connect a charger to charge the electronic device 700, and can also be used to transfer data between the electronic device 700 and peripheral devices.

[0349] The external memory interface 720 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 700. The external memory card communicates with the processor 710 through the external memory interface 720 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0350] The internal memory 721 can be used to store computer-executable program code. The executable program code includes instructions. The processor 710 executes various functional applications and data processing of the electronic device 700 by running the instructions stored in the internal memory 721. The internal memory 721 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function (such as a sound playback function, an image playback function, etc.).

[0351] In addition, the internal memory 721 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0352] The charge management module 740 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charge management module 740 can receive the charging input of the wired charger through the USB interface 730.

[0353] The power management module 741 is used to connect the battery 742, the charge management module 740, and the processor 710. The power management module 741 receives the inputs of the battery 742 and / or the charge management module 740 to supply power to the processor 710, the internal memory 721, the external memory, the display screen 793, the camera 792, the wireless communication module 760, etc.

[0354] In some other embodiments, the power management module 741 may also be disposed in the processor 710. In some other embodiments, the power management module 741 and the charging management module 740 may also be disposed in the same device.

[0355] The wireless communication function of the electronic device 700 may be implemented by the antenna 1, the antenna 2, the mobile communication module 750, the wireless communication module 760, the modulation and demodulation processor, and the baseband processor, etc.

[0356] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 700 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: the antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0357] The mobile communication module 750 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 700. The mobile communication module 750 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 750 can receive electromagnetic waves through the antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 750 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves through the antenna 1 and radiate it out.

[0358] The wireless communication module 760 can provide solutions for wireless communications including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied to the electronic device 700. The wireless communication module 760 may be one or more devices integrating at least one communication processing module. The wireless communication module 760 receives electromagnetic waves through the antenna 2, frequency-modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 710. The wireless communication module 760 can also receive the signal to be transmitted from the processor 710, frequency-modulate and amplify it, and convert it into electromagnetic waves through the antenna 2 and radiate it out.

[0359] In some embodiments, antenna 1 of electronic device 700 is coupled to mobile communication module 750, and antenna 2 is coupled to wireless communication module 760, so that electronic device 700 can communicate with the network and other devices through wireless communication technologies.

[0360] Electronic device 100 can implement audio functions through audio module 770, speaker 770A, receiver 770B, microphone 770C, headphone jack 770D, and application processor, etc. For example, music playback, recording, etc.

[0361] Audio module 770 is used to convert digital audio information into analog audio signals for output, and is also used to convert analog audio inputs into digital audio signals. Audio module 770 can also be used for encoding and decoding audio signals. In some embodiments, audio module 770 can be disposed in processor 710, or some functional modules of audio module 770 can be disposed in processor 710. Speaker 770A, also known as a "loudspeaker", is used to convert audio electrical signals into sound signals. Receiver 770B, also known as a "handset", is used to convert audio electrical signals into sound signals. Microphone 770C, also known as a "microphone", "transducer", is used to convert sound signals into electrical signals. Headphone jack 770D is used to connect a wired headphone. Headphone jack 770D can be a USB interface 730, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a Cellular Telecommunications Industry Association of the USA (CTIA) standard interface. In the embodiments of the present application, after the voice control function is turned on, electronic device 100 continuously collects surrounding sounds through microphone 770C.

[0362] Pressure sensor 780A is used to sense pressure signals and can convert pressure signals into electrical signals. In some embodiments, pressure sensor 780A can be disposed on display screen 793. There are many types of pressure sensors 780A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on pressure sensor 780A, the capacitance between the electrodes changes. Electronic device 700 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on display screen 793, electronic device 700 detects the intensity of the touch operation according to pressure sensor 780A. Electronic device 700 can also calculate the position of the touch according to the detection signal of pressure sensor 780A.

[0363] The touch sensor 780B, also known as the "touch panel". The touch sensor 780B can be disposed on the display screen 793, and together with the display screen 793, they form a touch screen, also known as the "touch control screen". The touch sensor 780B is used to detect touch operations acting on it or nearby. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 793. In some other embodiments, the touch sensor 780B can also be disposed on the surface of the electronic device 700, at a different position from that of the display screen 793.

[0364] The keys 790 include a power-on key, volume keys, etc. The keys 790 can be mechanical keys or touch keys. The electronic device 700 can receive key inputs and generate key signal inputs related to the user settings and function control of the electronic device 700.

[0365] The motor 791 can generate vibration prompts. The motor 791 can be used for incoming call vibration prompts or touch vibration feedback.

[0366] The camera 792 is used to capture still images or videos. In some embodiments, the electronic device 700 can include one or N cameras 792, where N is a positive integer greater than 1.

[0367] The electronic device 700 realizes the display function through the GPU, the display screen 793, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 793 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 710 can include one or more GPUs, which execute program instructions to generate or change display information.

[0368] The display screen 793 is used to display images, videos, etc. In some embodiments, the electronic device 700 can include one or N display screens 793, where N is a positive integer greater than 1.

[0369] The SIM card interface 794 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 794 to achieve contact and separation with the electronic device 700. The electronic device 700 can support one or N SIM card interfaces, where N is a positive integer greater than 1.

[0370] The voice control methods introduced in the above embodiments can all be executed in an electronic device with the above hardware structure.

[0371] Some other embodiments of the present application provide an electronic device (such as mobile phone 100). The electronic device may include: a microphone, a display screen, a processor, and a memory. The microphone, the display screen, and the memory are respectively coupled to the processor. The microphone is used to collect voices; the display screen is used to display the interface of the electronic device; the memory is coupled to the processor. The memory is further used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device can perform each function or step that the mobile phone 100 performs in the above method embodiments. The structure of the electronic device may refer to Figure 18 the structure of the electronic device 700 shown.

[0372] Embodiments of the present application further provide a chip system, as Figure 19 shown, the chip system 1000 includes at least one processor 1001 and at least one interface circuit 1002. The processor 1001 and the interface circuit 1002 can be interconnected through a line. For example, the interface circuit 1002 can be used to receive signals from other devices (such as the memory of a computer). Again, for example, the interface circuit 1002 can be used to send signals to other devices (such as the processor 1001). Exemplarily, the interface circuit 1002 can read the instructions stored in the memory and send the instructions to the processor 1001. When the instructions are executed by the processor 1001, the computer can perform each step in the above embodiments. Of course, the chip system can also include other discrete devices, and the embodiments of the present application do not make specific limitations on this.

[0373] Embodiments of the present application further provide a computer-readable storage medium, the computer-readable storage medium includes computer instructions, when the computer instructions run on the above electronic device (such as mobile phone 100), the electronic device is enabled to perform each function or step that the mobile phone 100 performs in the above method embodiments.

[0374] Embodiments of the present application further provide a computer program product, when the computer program product runs on a computer, the computer is enabled to perform each function or step that the computer performs in the above method embodiments. Among them, the computer can be an electronic device, such as mobile phone 100.

[0375] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0376] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0377] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0378] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0379] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes.

[0380] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A voice control method, characterized in that, Applied to an electronic device, the electronic device includes a microphone, and the method includes: Displaying a first interface, the first interface including a first switch; In response to a user's operation of turning on the first switch, enabling a voice control function; the enabling of the voice control function includes: enabling a recording channel corresponding to the voice control function, the recording channel being used to obtain the voice collected by the microphone; In response to a voice click instruction collected by the microphone, searching for a control to be clicked that matches the voice click instruction in the current display interface of the electronic device; In the case where the control to be clicked is partially blocked by a floating window, determining a first simulated click area based on the area corresponding to the floating window and the area corresponding to the control to be clicked; Executing a click event corresponding to the voice click instruction within the first simulated click area.

2. The method according to claim 1, characterized in that, The situation where the control to be clicked is partially blocked by a floating window includes: The current display interface of the electronic device includes a floating window, the floating window is not the focus window, and the area corresponding to the floating window and the area corresponding to the control to be clicked include an overlapping area; the overlapping area is smaller than the area corresponding to the control to be clicked.

3. The method according to claim 2, wherein The determining of the first simulated click area based on the area corresponding to the floating window and the area corresponding to the control to be clicked includes: Subtracting the overlapping area from the area corresponding to the control to be clicked to obtain the first simulated click area.

4. The method according to claim 2, wherein The method further includes: In the case where the current display interface of the electronic device includes a floating window, obtaining the area corresponding to the current focus window of the electronic device; Comparing the area corresponding to the focus window with the area corresponding to the floating window; If the area corresponding to the focus window is exactly the same as the area corresponding to the floating window, determining the floating window as the focus window.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: In the case where the control to be clicked is not blocked, executing a click event corresponding to the voice click instruction within a second simulated click area; the second simulated click area includes the area corresponding to the control to be clicked; Wherein, the situation where the control to be clicked is not blocked includes: the current display interface of the electronic device does not include a floating window; Or, the current display interface of the electronic device includes a floating window, and the floating window is the focus window; Or, the current display interface of the electronic device includes a floating window, the floating window is not the focus window, and the area corresponding to the control to be clicked and the area corresponding to the floating window do not have an overlapping area.

6. The method according to any one of claims 1-5, characterized in that The method further includes: In the case where the control to be clicked is completely blocked by the floating window, sending a prompt message, the prompt message being used to indicate that it is not clickable.

7. The method according to any one of claims 1-6, characterized in that, The searching for a control to be clicked that matches the voice click instruction in the current display interface of the electronic device includes: Searching for a first control that matches the voice click instruction in the current display interface of the electronic device; Obtaining the click attribute of the first control; If the click attribute of the first control is clickable, determining the first control as the control to be clicked.

8. The method according to claim 7, characterized in that The voice click instruction matches at least one interface hot word set by the at least one electronic device, and the first control is a text control.

9. The method according to claim 7 or 8, characterized in that, The finding of the control to be clicked that matches the voice click instruction in the current display interface of the electronic device further includes: If the click attribute of the first control is non-clickable, then a second control belonging to the same control group as the first control will be found; Obtain the click attribute of the second control; If the click attribute of the second control is clickable, then the second control is determined as the control to be clicked.

10. The method according to claim 7 or 8, characterized in that, The method further includes: If the click attribute of the first control is clickable and the first control is completely blocked by a floating window, find a second control belonging to the same control group as the first control; Obtain the click attribute of the second control; If the click attribute of the second control is clickable, then the second control is determined as the new control to be clicked, and an occlusion determination is made on the new control to be clicked.

11. The method according to any one of claims 2-9, characterized in that, The electronic device includes a voice control application package APK, a voice processing engine, and a window activity manager AMS; the method further includes: The voice control APK receives the current display interface of the electronic device returned by the AMS, and parses the current display interface to obtain and save the text controls and picture controls of the current display interface; The voice control APK uses the text included in the text controls of the current display interface as interface hot words and sends them to the voice processing engine; The voice control APK distributes the voice click instruction to the voice processing engine in response to receiving the voice click instruction; The voice processing engine finds a target interface hot word that matches the voice click instruction; The voice processing engine returns the target interface hot word to the voice control APK; The voice control APK queries whether the electronic device includes a floating window, and when the electronic device includes a floating window, determines whether the floating window is a focus window, and determines whether there is an overlapping area between the area corresponding to the floating window and the area corresponding to the control to be clicked.

12. An electronic device, characterized in that, The electronic device includes: a microphone, a display screen, a memory, and a processor; the microphone, the display screen, and the memory are respectively coupled to the processor; The microphone is used to collect language, the display screen is used to display the interface of the electronic device; the memory is used to store computer instructions; when the computer instructions are executed by the processor, the electronic device executes the method according to any one of claims 1-11.

13. A computer-readable storage medium, characterized in that, It includes computer instructions, and when the computer instructions are executed by the processor of the electronic device, the electronic device executes the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Method and device for game picture display, storage medium and electronic equipment

    CN107469351A

  • Voice control method and electronic equipment

    CN109584879A

  • Floating window control method and related product

    CN112654957A

  • Voice assistant standby method and device, equipment and storage medium

    CN116129897A

  • Voice control method and device and electronic equipment

    CN116560611A

Cited By

  • Knowable window control method, control device, electronic equipment and vehicle

    CN121070303A