Voice control method and electronic equipment

By identifying unclickable controls in the voice control electronic device and switching to the clickable controls of the same control group for simulated clicks, the problem of intent that cannot be achieved due to unclickable controls is solved, and the accuracy of voice control is improved.

CN120279901APending Publication Date: 2025-07-08HONOR DEVICE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202311868263.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

During the voice control process, when the unclickable text control cannot respond to the user's voice command, the user's click intention cannot be realized.

Method used

By setting up voice control functions in electronic devices, collecting voice commands using microphones, identifying unclickable controls, finding clickable controls in the same control group, and performing simulated click operations in their corresponding areas to ensure that they respond to user intentions.

Benefits of technology

Improves the accuracy of voice control, reduces the possibility that intentions cannot be achieved due to unclickable controls, and ensures that click operations are performed in the exact location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279901A_ABST
    Figure CN120279901A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice control method and electronic equipment, relates to the technical field of voice processing, and is used for realizing a click intention of a user in a scene that a text control matched with voice cannot be clicked. The method is applied to the electronic equipment, the electronic equipment comprises a microphone, and the method comprises the following steps: displaying a first interface comprising a first switch; and in response to the operation of turning on the first switch, turning on the voice control function. Specifically, a recording channel corresponding to a voice control function is opened, and the recording channel is used for obtaining voice collected by a microphone. And in response to a voice click instruction acquired by the microphone, searching a first control matched with the first voice instruction in the current display interface. And under the condition that the click attribute of the first control is that the first control cannot be clicked, searching a second control belonging to the same control group with the first control in the current display interface. And if the click attribute of the second control is clickable, executing a click event corresponding to the voice click instruction based on an area corresponding to the second control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of speech processing technologies, and in particular, to a speech control method and an electronic device. Background Art

[0002] With the development of technologies, speech control has gradually become a commonly used human-computer interaction method. In particular, when a user's hands are occupied in scenarios such as driving, cooking, or reading, it is very convenient and fast to control an electronic device through speech.

[0003] In the process of an electronic device performing a click operation in response to speech input by a user, in related technologies, it is necessary to first find a text control to be clicked corresponding to the speech. Then, the electronic device simulates a click operation on the center position of the text control to be clicked. However, among the controls displayed by the electronic device, some text controls are non-clickable. In this case, when the electronic device responds to the speech and simulates a click operation on the text control to be clicked, it will not be able to respond, and thus the user's intention cannot be realized. Therefore, there is an urgent need for a method to realize the user's click intention in a scenario where the text control matching the speech is non-clickable. Summary of the Invention

[0004] Embodiments of the present application provide a speech control method and an electronic device for realizing the user's click intention in a scenario where the text control matching the speech is non-clickable.

[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, there is provided a method applied to an electronic device, the electronic device including a microphone. The method includes:

[0007] The electronic device displays a first interface, and the first interface includes a first switch. In response to the user's operation of turning on the first switch, the voice control function is enabled. The enabling of this voice control function is controlled by the first switch and does not require waking up the electronic device. For example, this voice control function is a visible-and-speak function. After the voice control function is enabled, when a voice click instruction collected by the microphone is received, the first control that matches the first voice instruction can be searched for on the current display interface of the electronic device. Then, the click attribute of the first control is judged. If the click attribute of the first control is non-clickable, it is searched whether there is a second control in the current display interface that belongs to the same control group as the first control. If there is a second control and its click attribute is clickable, the electronic device executes the click event corresponding to the voice click instruction based on the area corresponding to the second control. In this way, when the voice intention recognition is accurate, it is ensured that the electronic device can respond to the voice input by the user and execute the click instruction at the accurate position. Furthermore, the possibility that the electronic device cannot realize the true intention of the voice input by the user is reduced, and the possibility of executing an accurate simulated click operation in response to the voice input by the user is increased.

[0008] Among them, enabling the voice control function includes: enabling the recording channel corresponding to the voice control function, and the recording channel is used to obtain the voice collected by the microphone. After the recording channel corresponding to the voice control function is enabled, the voice control function can obtain the voice (audio stream) collected by the microphone through this recording channel.

[0009] In a possible implementation manner of the first aspect, the above-mentioned execution of the click event corresponding to the voice click instruction based on the area corresponding to the second control may specifically include: determining a simulated click area based on the area corresponding to the second control. Then, determining a simulated click position within the simulated click area. Finally, the electronic device executes a simulated click operation at the simulated click position. In this way, it is ensured that the electronic device can respond to the voice input by the user and execute the click instruction at the accurate position.

[0010] In a possible implementation manner of the first aspect, the above-mentioned simulated click area is the area corresponding to the second control. Since the click attribute of the second control is clickable, therefore, the electronic device can execute the click instruction in the area corresponding to the second control.

[0011] In a possible implementation of the first aspect, before determining the simulated click area based on the area corresponding to the second control, the above method further includes: obtaining the area corresponding to the first control. In this implementation, determining the simulated click area based on the area corresponding to the second control may specifically include: splicing the area corresponding to the second control with the area corresponding to the first control to obtain a spliced area, and the simulated click area includes the spliced area. In this way, the first control and the found second control are spliced, and then the simulated click position is determined from the spliced area. Usually, the proportion of the clickable control area in the same control group displayed on the electronic device is greater than that of the non-clickable control area. Therefore, when splicing the clickable control area and the non-clickable control area, the determined simulated click position is also clickable. Therefore, a clickable simulated click position can be obtained to execute the click instruction.

[0012] In a possible implementation of the first aspect, determining the simulated click position within the simulated click area may specifically include: determining the center position of the simulated click area as the simulated click position. In this way, the accuracy of the electronic device when executing the click instruction can be improved, and the possibility of clicking on other controls can be reduced.

[0013] In a possible implementation of the first aspect, after finding the second control that belongs to the same control group as the first control in the current display interface, the above method may further include: obtaining the number of second controls. There may be multiple second controls at the same level that belong to the same control group as the first control, or there may be none. It can be understood that if the first control has no second controls at the same level, the click instruction cannot be executed. When the first control includes multiple second controls at the same level, since it is impossible to accurately determine the control that the user needs to click, in this case, the electronic device may also not execute the click instruction. In this implementation, when the number of second controls is 1, the electronic device further determines whether the second control is clickable. And if the click attribute of the second control is clickable, based on the area corresponding to the second control, the click event corresponding to the voice click instruction is executed. In this way, it can be ensured that the electronic device can execute the click instruction at the accurate position in response to the voice click instruction; the problem that the simulated click operation is executed at the wrong position, resulting in the click instruction not conforming to the user's intention, can be avoided.

[0014] In a possible implementation of the first aspect, the above method further includes: when no second control that belongs to the same control group as the first control is found, sending a prompt message. The prompt message is used to indicate non-clickability. In this way, the user can be quickly informed of non-clickability, which is convenient for prompting the user to re-enter the voice instruction.

[0015] In a possible implementation of the first aspect, the above method further includes: when a second control belonging to the same control group as the first control is found and the click attribute of the second control is non-clickable, a prompt message is issued. The prompt message is used to indicate non-clickability. In this way, the user can be quickly informed of non-clickability, which is convenient for prompting the user to re-enter a voice command.

[0016] In a possible implementation of the first aspect, the above method further includes: when a second control belonging to the same control group as the first control is found and the number of second controls is greater than 1, a prompt message is issued. The prompt message is used to indicate non-clickability. In this way, it can be ensured that the electronic device can execute the click command at the accurate position in response to the voice click command; the problem that the click command is executed at the wrong position, resulting in the click command not conforming to the user's intention, can be avoided.

[0017] In a possible implementation of the first aspect, the electronic device includes a voice control application package APK, a voice processing engine, and an activity manager service AMS. The above method further includes: the voice control APK receives the top-level activity of the electronic device returned by the AMS, parses the top-level activity, obtains and saves the controls included in the current display interface of the electronic device, and the controls include text controls. The voice control APK sends the text included in the text controls of the current display interface to the voice processing engine. The voice control APK distributes the voice click command to the voice processing engine in response to receiving the voice click command. The voice processing engine searches for the target text that matches the voice click command. The voice processing engine maps the voice click command to a control click command and returns the control click command to the voice control APK, and the control click command carries the target text. The voice control APK searches for the first control corresponding to the target text from the text controls of the current display interface and obtains the click attribute of the first control.

[0018] In a possible implementation of the first aspect, the above method further includes: when the click attribute of the first control is non-clickable, the voice control APK searches for a second control belonging to the same control group as the first control and obtains the click attribute of the second control.

[0019] In a possible implementation of the first aspect, the voice control APK includes an interface content acquisition module and an interface parsing module; the above method further includes: the interface content acquisition module sends an acquisition request to the AMS. In response to the acquisition request, the AMS sends the top-level activity of the electronic device to the interface content acquisition module. The interface content acquisition module sends the top-level activity to the interface parsing module. The interface parsing module parses the top-level activity to obtain the display content of the current display interface of the electronic device.

[0020] In a possible implementation of the first aspect, the voice click instruction matches at least one interface hot word of the electronic device. The interface hot word of the electronic device is set according to the text content in the current display interface of the electronic device.

[0021] In a possible implementation of the first aspect, in response to the user's operation of closing the first switch, the electronic device turns off the voice control function. Among them, turning off the voice control function includes: turning off the recording channel corresponding to the voice control function.

[0022] In another possible implementation of the first aspect, turning on the voice control function further includes: displaying a first recording icon and a second recording icon, where the first recording icon indicates that the voice control function is turned on, and the second recording icon indicates that the recording channel is turned on.

[0023] In a second aspect, the present application further provides an electronic device. The electronic device may include: a microphone, a display screen, a processor, and a memory. The microphone, the display screen, and the memory are respectively coupled to the processor. The microphone is used to collect language, the display screen is used to display the interface of the electronic device. The memory is used to store computer execution instructions. When the electronic device runs, the processor executes the computer execution instructions stored in the memory, so that the electronic device executes the voice control method according to any one of the above first aspects.

[0024] In a third aspect, the present application provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed by the processor of the electronic device, the electronic device can execute the voice control method according to any one of the above first aspects.

[0025] In a fourth aspect, a computer program product including instructions is provided. When it runs on an electronic device, the electronic device can execute the voice control method according to any one of the above first aspects.

[0026] In a fifth aspect, a device (for example, the device may be a chip system) is provided. The device includes a processor for supporting the electronic device to implement the functions involved in the above first aspect. In a possible design, the device further includes a memory for storing necessary program instructions and data of the electronic device. When the device is a chip system, it may be composed of chips or may include chips and other discrete devices.

[0027] Among them, the technical effects brought by any one of the design methods in the second aspect to the fifth aspect can refer to the technical effects brought by different design methods in the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of a scenario example of a voice control method;

[0029] Figure 2A It is a schematic diagram of an example process for enabling a voice control function;

[0030] Figure 2B It is a schematic diagram of a scenario example of a voice control method;

[0031] Figure 2C It is a schematic diagram of a scenario example of a voice control method;

[0032] Figure 3 It is a schematic flowchart of a voice control method provided by an embodiment of the present application;

[0033] Figure 4 It is a schematic diagram of a scenario example of the voice control method provided by an embodiment of the present application;

[0034] Figure 5A It is a schematic diagram of the display interface of an electronic device;

[0035] Figure 5B It is a schematic diagram of an example process for enabling a voice control function;

[0036] Figure 6 It is a schematic flowchart of a voice control method provided by an embodiment of the present application;

[0037] Figure 7A It is a schematic flowchart of a voice control method provided by an embodiment of the present application;

[0038] Figure 7B It is a schematic diagram of a scenario example of the voice control method provided by an embodiment of the present application;

[0039] Figure 7C It is a schematic diagram of a scenario example of the voice control method provided by an embodiment of the present application;

[0040] Figure 8 It is a schematic diagram for determining an analog click area provided by an embodiment of the present application;

[0041] Figure 9A It is a schematic flowchart of a voice control method provided by an embodiment of the present application;

[0042] Figure 9B It is a schematic diagram of a scenario example of a voice control method provided by an embodiment of the present application;

[0043] Figure 10 It is a schematic diagram of the software architecture of an electronic device provided by an embodiment of the present application;

[0044] Figure 11 It is a schematic diagram of the interaction of each module of an electronic device when implementing the voice control method provided by an embodiment of the present application;

[0045] Figure 12 Schematic diagram of interactions among various modules of an electronic device when implementing the voice control method provided in an embodiment of the present application;

[0046] Figure 13 Flow chart of a voice control method provided in an embodiment of the present application;

[0047] Figure 14 Schematic diagram of the software architecture of an electronic device provided in an embodiment of the present application;

[0048] Figure 15 Schematic diagram of interactions among various modules of an electronic device when implementing the voice control method provided in an embodiment of the present application;

[0049] Figure 16 Schematic diagram of interactions among various modules of an electronic device when implementing the voice control method provided in an embodiment of the present application;

[0050] Figure 17 Flow chart of a voice control method provided in an embodiment of the present application;

[0051] Figure 18 Schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0052] Figure 19 Schematic diagram of the structure of a chip system provided in an embodiment of the present application. Detailed implementation manners

[0053] To facilitate a clear description of the technical solutions in the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:

[0054] The goal of automatic speech recognition (ASR) is to convert the lexical content in the user's speech into computer-readable input, such as keystrokes, binary codes, or character sequences.

[0055] Natural language understanding (NLU) is a general term for all method models or tasks that support machines to understand the content of text.

[0056] Dialogue management (DM) is used to control the process of human-machine dialogue and decide the response to the user at this moment according to the dialogue history information.

[0057] Activity is the Android system One of the four major components of the system, it is a visual interface for user operations; it provides users with a window to complete operation instructions.

[0058] Intent is the idea of ​​wanting to achieve a certain purpose. In the field of voice control, intent recognition is an important technology. Only by accurately identifying and understanding the needs and intentions of users can more accurate instructions be executed in response to the user's voice to meet the needs of users. Taking the voice input by the user as "query today's weather" as an example, the electronic device can perform intent recognition on the voice and extract the entity content in the voice: "query", "weather", so as to determine that the user's intention is to query the weather. For another example, the voice input by the user is "return to the previous level", the electronic device can perform intent recognition on the voice and extract the entity content in the voice: "return", "previous level", so as to determine that the user's intention is to display the previous level page. For another example, the voice input by the user is "open the calendar", the electronic device can extract the entity content in the voice: "open", "calendar", so as to determine that the user's intention is to open the calendar application.

[0059] Voice control function:

[0060] Many electronic devices support voice control. Electronic devices collect user input voice through microphones, analyze and recognize the voice, and execute commands corresponding to the voice, so that users can control the electronic devices through voice.

[0061] Generally speaking, in order to save power consumption of electronic devices and avoid false triggering, the voice control function needs to be turned on before it can be used. For example, the user inputs a preset word (called a wake-up word) to the electronic device through voice to wake up the electronic device. After the electronic device is awakened, it can execute the command corresponding to the voice, that is, the voice control function is turned on. For example, the user can turn on or off the voice control function by turning on or off the preset switch in the human-computer interaction interface of the electronic device.

[0062] In different electronic devices, the voice control function may have different names, such as "voice control", "intelligent voice", "voice assistant", "see and speak", "voice command", "free command", "intelligent AI", etc. The specific implementation of voice control functions with different names may also be different.

[0063] Several different implementations of the voice control function are described below by way of example.

[0064] Voice Assistant:

[0065] Before a user uses a voice assistant to control an electronic device, the electronic device needs to be woken up first. In one example, when the wake-up word spoken by the user to the electronic device is detected, the electronic device is woken up. In another example, when the operation of the user long pressing the power button is detected, the electronic device is woken up. In yet another example, when the breath generated when the user inputs voice to the electronic device is detected, the electronic device is woken up. Generally speaking, before the electronic device is woken up, the microphone of the electronic device works in a power-saving mode (such as searching for signals with a lower power) to pick up the sound of the surrounding environment. The voice collected by the microphone only performs wake-up word detection at the kernel layer, and the corresponding recording channels of the voice assistant are not started in the system and driver of the electronic device.

[0066] The electronic device is woken up in response to a user operation (such as receiving the wake-up word spoken as input). The corresponding recording channels of the voice assistant are started in the system and driver. After the electronic device is woken up, the voice (audio stream) collected by the microphone is sent to the voice assistant application for processing through the corresponding recording channels of the voice assistant. In this way, the electronic device can execute the instructions corresponding to the voice, enabling the user to control the electronic device through voice; it can also implement functions such as having a conversation with the user.

[0067] Taking the electronic device as the mobile phone 100 as an example, exemplarily, Figure 1 shows a schematic diagram of a scenario example where a user uses a voice assistant to control the mobile phone 100. As Figure 1 shown, the mobile phone 100 displays the desktop interface, and the user inputs the voice "Hello YOYO" to the mobile phone 100. In response to receiving the wake-up word "Hello YOYO", the mobile phone 100 is woken up. Exemplarily, after the mobile phone 100 is woken up, it plays the voice "I'm here" to prompt the user that the mobile phone 100 has been woken up. After the mobile phone 100 is woken up, the user can control the mobile phone 100 through voice. Exemplarily, as Figure 1 shown, the user inputs the voice "Open the video" to the mobile phone. The mobile phone 100 parses and recognizes the voice input by the user and executes the instructions corresponding to the voice "Open the video". Exemplarily, in response to receiving the voice "Open the video", the mobile phone 100 starts the video application.

[0068] In some other examples, the electronic device can also wake up the voice assistant in response to receiving the operation of the user long pressing the power button.

[0069] In some implementations, after the voice assistant of the electronic device is awakened, the user can issue an instruction to the electronic device by inputting voice to the electronic device, and the electronic device executes the instruction corresponding to the voice. After the electronic device executes an instruction, or if no instruction is received from the user through voice within a certain period of time (such as within 8 seconds) after the voice assistant of the electronic device is awakened, the electronic device no longer responds to instructions issued through voice. For example, the electronic device will close the recording channel corresponding to the voice assistant. The user needs to input the wake-up word to the electronic device again to wake up the electronic device before they can issue instructions to the electronic device by inputting voice again. That is to say, after the voice assistant of the electronic device is awakened, it enters a "short voice reception" state and can respond to instructions issued by the user through voice within a relatively short period of time (such as within 8 seconds).

[0070] In some implementations, when the electronic device is connected to the network, it supports entering a continuous conversation scenario after being awakened, and the user can have a continuous conversation with the electronic device. After each announcement by the electronic device, it will continue to pick up sound without the need to be awakened again until the user exits the continuous conversation through instructions such as "exit".

[0071] See and speak:

[0072] See and speak is implemented locally by the electronic device without the need to connect to the network.

[0073] In some examples, see and speak is controlled by a preset switch. The user can turn on the preset switch to enable the see and speak function, or turn off the preset switch to disable the see and speak function.

[0074] Exemplarily, as Figure 2A shown, the user can open the settings function of the mobile phone 100; for example, the user clicks on the application icon of the "Settings" application on the desktop. In response to the user's click operation on the application icon of the "Settings" application, the mobile phone 100 displays the "Settings" interface 101. The "Settings" interface 101 includes a "Smart Voice" option 102, and the "Smart Voice" option 102 is used to set the smart voice function. Exemplarily, in response to the user's click operation on the "Smart Voice" option 102, the mobile phone 100 displays the "Smart Voice" interface 103, and the "Smart Voice" interface 103 includes a "See and Speak" option 104. The user can click on the "See and Speak" option 104 to set options related to the see and speak function. Exemplarily, refer to Figure 2A, in response to the user's click operation on the "Speak as You See" option 104, the mobile phone 100 displays the "Speak as You See" interface 105. Optionally, the "Speak as You See" interface 105 includes a prompt message 106 for prompting the user about the usage method of the Speak as You See function. The "Speak as You See" interface 105 also includes a "Speak as You See" switch 107 (i.e., the above-mentioned preset switch). The user can click the "Speak as You See" switch 107 to turn on or off the "Speak as You See" switch. In one example, in response to receiving the user's click operation on the "Speak as You See" switch 107, the "Speak as You See" switch of the mobile phone 100 is turned on, enabling the Speak as You See function. Optionally, the "Speak as You See" interface 105 displays a prompt message 108 for prompting the user that the Speak as You See function has been successfully enabled.

[0075] In one implementation, after the Speak as You See function is enabled, the mobile phone 100 displays a first recording icon, which indicates that the Speak as You See function has been enabled. Exemplarily, as Figure 2A shown, after the "Speak as You See" switch 107 is turned on, the status bar of the interface displayed by the mobile phone 100 shows a recording icon 10a, indicating that the Speak as You See function has been enabled.

[0076] In one scenario, the preset switch corresponding to Speak as You See on the electronic device is not turned on, and the microphone of the electronic device is not enabled. When the preset switch corresponding to Speak as You See is turned on, the electronic device activates the microphone and starts the corresponding recording channel for Speak as You See in the system and the driver. In this way, the voice (audio stream) collected by the microphone can be sent to the Speak as You See application for processing through the corresponding recording channel for Speak as You See, enabling the user to control the electronic device by voice.

[0077] In another scenario, the preset switch corresponding to Speak as You See on the electronic device is not turned on, and the microphone of the electronic device works in a power-saving mode (such as searching for signals with a lower power) to pick up the surrounding sounds. When the preset switch corresponding to Speak as You See is turned on, the corresponding recording channel for Speak as You See is started in the system and the driver of the electronic device. In this way, the voice (audio stream) collected by the microphone can be sent to the Speak as You See application for processing through the corresponding recording channel for Speak as You See, enabling the user to control the electronic device by voice.

[0078] After Speak as You See is enabled, the corresponding recording channel for Speak as You See is started in the system and the driver of the electronic device, and the electronic device enters a "long recording" state, continuously collecting surrounding sounds. The user can issue commands to the electronic device by voice at any time without the need to enter a wake-up word to wake up the electronic device.

[0079] In one implementation, after any function on the electronic device enables the voice input function of the electronic device (turns on the microphone and the recording channel), the electronic device will send a prompt message to the user to indicate that the electronic device is in a voice collection state. In this way, the privacy of the user can be avoided from being leaked. For example, after the visible and speakable function is enabled, the electronic device enters a continuous voice collection state, and a second recording icon is displayed on the display interface of the electronic device. This second recording icon indicates that the recording channel is open and is used to prompt the user that the microphone is collecting voice. Exemplarily, as Figure 2A shown, the status bar of the display interface of the mobile phone 100 displays a recording icon 10b, indicating that the recording channel is open.

[0080] After the visible and speakable function is enabled, the electronic device opens the recording channel and continuously collects voice through the microphone. The user can input voice to the electronic device at any time. The electronic device parses and recognizes the voice input by the user and executes the instruction corresponding to the voice.

[0081] The instructions supported by the visible and speakable function for the user to input through voice can include: system instructions, such as swiping left, swiping right, swiping up, returning to the desktop, going back, increasing the volume, and decreasing the volume; video application instructions, such as playing, pausing, stopping, fast forwarding, and rewinding; and e-book playback application instructions, such as going to the previous page, going to the next page, going to the table of contents, and going to the next chapter.

[0082] The visible and speakable function of the electronic device supports classifying the instructions input by the user through voice into multiple vertical categories, and one of the vertical categories is the operation vertical category. For the operation vertical category instructions, when the electronic device responds to the voice and executes the instruction corresponding to the voice, it will simulate the user operation.

[0083] Hot words:

[0084] The electronic device can set some hot words. After the visible and speakable function is enabled, if the received voice matches at least one of the set hot words, the electronic device executes the instruction corresponding to the voice.

[0085] The hot words can be pre-configured on the electronic device, can be obtained according to the content in the display interface of the electronic device, or can be input by the user, etc. According to the different usage scopes and sources of the set hot words, the hot words can be divided into system-level hot words, scenario-level hot words, and interface hot words.

[0086] Among them, the system-level hot words are pre-configured and applicable to any application on the electronic device. When any application on the electronic device is running in the foreground, if the voice input by the user matches at least one system-level hot word, the electronic device executes the instruction corresponding to the user's voice. The system-level hot words are global and do not depend on the application or the application interface. Exemplarily, the system-level hot words can include: swiping left, swiping right, swiping up, returning to the desktop, going back, etc.

[0087] Scene-level hot words are applicable to all applications within a scene. In one implementation, multiple scenes are preset in the electronic device, and each scene corresponds to at least one application. Exemplarily, the preset scenes may include an audio-video scene, an e-book scene, a main screen scene, and a permission pop-up window scene, etc. Different scenes may preset different scene-level hot words. Exemplarily, the scene-level hot words corresponding to the audio-video scene may include: play, pause, stop, fast forward, and rewind, etc. The scene-level hot words corresponding to the e-book scene may include: previous page, next page, table of contents, and next chapter, etc. The scene-level hot words corresponding to the main screen scene may include: application list information, settings menu item, etc. The scene-level hot words corresponding to the permission pop-up window scene may include: allow, confirm, deny, and I know, etc.

[0088] Interface hot words are applicable to a certain interface. In some embodiments, the electronic device obtains hot words from the display interface of the foreground application, that is, obtains interface hot words. Exemplarily, referring to Figure 2B , the mobile phone 100 displays the "My" interface 110 of the video application. The text information in the "My" interface 110 includes "5G", "8:00", "Login / Register", "My Downloads", "Following & Favorites", "My Purchases", "My Scenes", "History", "Coupon Pack", "Settings", "Feedback", "Customer Service", "Home", "Membership", "Short Videos", and "My", etc. The mobile phone 100 sets the text in the currently displayed interface of the foreground application as the interface hot words, that is, the interface hot words include "5G", "8:00", "Login / Register", "My Downloads", "Following & Favorites", "My Purchases", "My Scenes", "History", "Coupon Pack", "Settings", "Feedback", "Customer Service", "Home", "Membership", "Short Videos", and "My", etc. When the electronic device monitors the switching of the display interface, it can, after the interface is switched, re-obtain the text content of the currently displayed interface and update the interface hot words.

[0089] In some scenes, the electronic device can display multiple windows simultaneously. For example, in the scene where the electronic device displays a floating window, the electronic device displays two windows simultaneously. In this scenario, the interface hot words obtained and saved by the electronic device are specifically the interface hot words displayed in the focused window. The focused window refers to the selected window among multiple windows, that is, the currently operated window. In this scenario, the global operations performed by the user act on this focused window.

[0090] Exemplarily, the electronic device simultaneously displays the settings interface and displays the calculator interface in the floating window, and the focused window is the window corresponding to the settings interface. At this time, the interface hot words obtained and saved by the electronic device at this time are the hot words obtained and saved from the settings interface.

[0091] Different hot words can correspond to different instructions. Since the interface hot words are set according to the displayed text of the current display interface, in some embodiments, the instructions corresponding to the interface hot words are control click instructions. Specifically, when the electronic device determines that the received voice matches at least one interface hot word, it can execute the control click instruction on the position of the text corresponding to the interface hot word on the interface. In some examples, when the electronic device executes the control click instruction, it can specifically execute the click event corresponding to the control click instruction, that is, perform a simulated click operation at the position of the corresponding text.

[0092] In Figure 2B the example shown, the mobile phone 100 receives the voice "Open History" input by the user. The voice "Open History" matches the interface hot word "History", and the mobile phone 100 executes the instruction corresponding to the voice "Open History". Specifically, the mobile phone 100 can perform a simulated click operation within the area corresponding to the "History" control corresponding to "History", so as to realize the user's intention of opening the history. As Figure 2B shown, after the mobile phone 100 executes the simulated click operation, the history interface 113 can be displayed.

[0093] Among them, when the mobile phone 100 executes the instruction corresponding to the voice "Open History", it needs to determine the area corresponding to the "History" control corresponding to "History". Then, according to the area corresponding to the "History" control, a simulated click position is determined. In one example, the electronic device can determine the center position of the area corresponding to the "History" control as the simulated click position. As Figure 2B shown, the center position of the area corresponding to the "History" control is position 112, which can be determined as the simulated click position.

[0094] When the mobile phone 100 displays the "My" interface of the video application, the user can also input other voices, such as "Coupon Package". The mobile phone 100 receives the voice "Coupon Package" input by the user, determines that the voice "Coupon Package" matches the interface hot word "Coupon Package", and the mobile phone 100 will execute the instruction corresponding to the voice "Coupon Package". As Figure 2C shown, the mobile phone 100 performs a simulated click operation in the area corresponding to the "Coupon Package" control 114 corresponding to the interface hot word "Coupon Package".

[0095] However, in some embodiments, the click attribute of the "Coupon Package" control 114 is non-clickable. Then, after the mobile phone 100 responds to the voice "Coupon Package" and performs a simulated click operation on the "Coupon Package" control 114, the mobile phone 100 cannot open the coupon package interface. In this way, the mobile phone 100 cannot realize the user's intention.

[0096] Based on this, an embodiment of the present application proposes a voice control method, which can be applied to an electronic device supporting voice input. After the electronic device responds to the user's operation and activates the voice control function, it can receive voice and execute the instruction corresponding to the voice. When the electronic device receives the first voice containing the control click intention on the second interface, it searches for the control (denoted as the first control) matching the first voice on the second interface. If the first control is not clickable, then the electronic device searches for the second control belonging to the same control group as the first control on the second interface. If the second control is clickable, a simulated click area is determined based on the second control, and a simulated click operation is performed within the simulated click area, thereby realizing the intention of the first voice. In this way, when the voice intention is accurately recognized, the possibility that the electronic device cannot realize the true intention of the voice input by the user is reduced, and the possibility of performing an accurate simulated click operation in response to the voice input by the user is increased.

[0097] Exemplarily, the above-mentioned electronic device may be a mobile phone, a tablet computer, a notebook computer, a personal computer (PC), an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a smart home device (such as a smart TV, a smart screen, a large screen, a smart speaker, a smart air conditioner, etc.), a personal digital assistant (PDA), a wearable device (such as a smart watch, a smart bracelet, etc.), a vehicle-mounted device, a virtual reality device, etc. The embodiments of the present application do not make any restrictions on this.

[0098] Hereinafter, the specific implementation manner of the voice control method proposed in the embodiments of the present application will be described in detail with reference to the accompanying drawings. Figure 3 The flowchart of the voice control method in some embodiments is shown.

[0099] S200. Display a first interface, where the first interface includes a first switch.

[0100] The first switch is used to turn on or off the voice control function. The user can turn on or off the voice control function through the first switch on the first interface. In some embodiments, the first interface may be Figure 2A the "visible and audible" interface 105 shown in the figure; the first switch may be the "visible and audible" switch 107.

[0101] S201. Receive an operation to turn on the first switch.

[0102] S202. Respond to the operation of turning on the first switch and activate the voice control function.

[0103] In some embodiments, after the voice control function is enabled, the electronic device activates the recording channel corresponding to the voice control function. This recording channel is used to obtain the voice collected by the microphone. Moreover, after the voice control function is enabled, the electronic device is in a long recording state, continuously collecting the surrounding sounds. The user can issue commands to the electronic device through voice at any time, without the need to input a wake word to wake up the electronic device.

[0104] In addition, when the electronic device enables the voice control function, it may specifically further include: displaying a first recording icon and a second recording icon, where the first recording icon indicates that the voice control function is enabled, and the second recording icon indicates that the recording channel is enabled. Informing the user that the voice control function has been enabled and the electronic device is recording in the form of recording icons. In this way, the user can quickly obtain the current status information of the electronic device.

[0105] S203. Display a second interface.

[0106] The second interface can be any interface of the electronic device. After the voice control function of the electronic device is enabled, on any display interface, it can receive the voice input by the user and execute the command corresponding to the voice. In some embodiments, the second interface can also be the first interface, that is, after the electronic device enables the voice control function, it can receive the voice input by the user on the first interface and execute the command corresponding to the voice.

[0107] In some embodiments, after the electronic device enables the voice control function, it can, after each display interface is switched, parse the switched display interface, obtain the control information of each control in the display interface, and save it. The control information may include: control attributes (text control / picture control), the area where the control is located (for a rectangular control, it can be represented by the upper left corner point coordinates and the lower right corner point coordinates), and whether the control is clickable. In this embodiment, after the electronic device displays the second interface, it can also parse the second interface to obtain the control information of each control in the second interface. In this way, it is convenient to make an accurate response to the received voice subsequently.

[0108] Since the electronic device has enabled the voice control function, the user can input voice to the electronic device through the voice control function on the second interface. Correspondingly, the electronic device receives the voice input by the user, as in S204.

[0109] S204. Receive the first voice.

[0110] S205. Determine the user intention corresponding to the first voice.

[0111] In some embodiments, the first voice matches at least one hot word preset in the electronic device. In this way, the electronic device can respond to the first voice and execute the instruction corresponding to the first voice, which can be recorded as the first voice instruction.

[0112] In some embodiments, after S204, the electronic device can perform voice parsing on the first voice to obtain the parsed text corresponding to the first voice. Then, the electronic device can perform intent recognition on the obtained parsed text. In this embodiment, S205 above can specifically include: parsing the first voice to obtain the parsed text corresponding to the first voice. Then, performing intent recognition on the parsed text corresponding to the first voice to obtain the user intent corresponding to the first voice.

[0113] As can be seen from the above description, after receiving the voice, the electronic device can match the voice with the hot words. If it is determined that the voice matches at least one hot word set in the electronic device, the electronic device can execute the instruction corresponding to the voice. Therefore, in some embodiments, performing intent recognition on the parsed text corresponding to the first voice to obtain the user intent corresponding to the first voice can specifically include: matching the parsed text corresponding to the first voice with the hot words set in the electronic device. If the parsed text corresponding to the first voice matches at least one hot word of the electronic device, the user intent can be determined according to the instruction corresponding to the hot word.

[0114] For example, if the parsed text corresponding to the first voice matches the interface hot word, the electronic device can determine that the instruction corresponding to the first voice is a control click instruction. Furthermore, it can be determined that the user intent of the first voice is a control click intent. Another example is that if the hot word that the parsed text corresponding to the first voice matches belongs to "swipe left" in the system-level hot words, the user intent can be determined to be swiping left. Or, if the hot word that the parsed text corresponding to the first voice matches belongs to "pause" in the scenario-level hot words, the user intent can be determined to be controlling the electronic device to pause playback.

[0115] As can be seen from the above description, the interface hot words set by the electronic device are obtained from the current focused window. In embodiments where the electronic device displays multiple windows simultaneously, the first voice input by the user may be the content in a non-focused window. In this case, when the electronic device matches the parsed text corresponding to the first voice with the interface hot words, it will not be able to obtain an interface hot word that matches the parsed text corresponding to the first voice. In some embodiments, in this case, the electronic device will recognize the first voice as an invalid instruction.

[0116] In some other embodiments, the electronic device can also preset multiple valid instructions in advance. When the voice input by the user matches the valid instructions set by the electronic device, the electronic device can execute the corresponding instructions in response to the voice. Exemplarily, the valid instructions can include: "Return", "Go back to the desktop", "Play / Pause", "Previous page / Next page", "Swipe left / Swipe right", and "Open [app name]", etc. Different valid instructions can correspond to different intents. In this embodiment, the above S205 can specifically include: finding the valid instruction that matches the first voice. According to the valid instruction that matches the first voice, determining the user intent corresponding to the first voice. Among them, to find the valid instruction that matches the first voice, the first voice can be first subjected to voice parsing, and then the obtained parsed text is compared one by one with the valid instructions stored in the electronic device to determine whether the parsed text matches at least one valid instruction.

[0117] In the above embodiments, the cases where the electronic device finds the matching hot word or the matching valid instruction according to the parsed text corresponding to the first voice are described. In some other embodiments, it is possible that no matching hot word or valid instruction can be found for the first voice input by the user. In this case, the electronic device cannot determine the user intent and thus cannot execute the corresponding instruction in response to the voice.

[0118] When it is determined that the user intent corresponding to the first voice is a control click intent, the instruction corresponding to the first voice can be recorded as a voice click instruction.

[0119] In other embodiments, the electronic device can also determine the user intent corresponding to the first voice in other ways.

[0120] S206. Determine whether the user intent is a control click intent.

[0121] After determining the user intent, it can be determined whether the user intent is a control click intent. In some embodiments, if the first voice matches at least one interface hot word set by the electronic device, it can be determined that the user intent is a control click intent.

[0122] If the judgment result of S206 is negative, it means that the user intent corresponding to the first voice is not a control click intent. In this case, the electronic device can directly execute the instruction corresponding to the first voice to achieve the user intent. It should be noted that the case where the judgment result of S206 is negative is not shown in Figure 3 is not shown.

[0123] When the judgment result of S206 is positive, it means that the user intent corresponding to the first voice is a control click intent. After that, the electronic device can determine the control to be clicked corresponding to the first voice.

[0124] S207. Locate the first control that matches the first voice in the second interface.

[0125] If it is determined based on the first voice that the user's intention is a control click intention, then the electronic device needs to determine the control to be clicked. Then, the area where the control to be clicked is located is obtained before the control click instruction can be executed. In some embodiments, the control found to match the voice is a text control on the current display interface of the electronic device.

[0126] In some embodiments, when determining the user intention corresponding to the first voice, the first voice is parsed to obtain a parsed text, and then the hot word that matches the first voice is determined according to the parsed text. In this embodiment, the above S207 may specifically include: locating the first control that matches the first voice according to the parsed text corresponding to the first voice. As Figure 2B In the example shown, the user inputs the voice "Open History". The electronic device determines that this voice matches the interface hot word "History". After that, the electronic device can locate the corresponding "History" control according to this interface hot word "History". As Figure 2C In the example shown, the user inputs the voice "Coupon Package". The electronic device determines that this voice matches the interface hot word "Coupon Package". After that, the electronic device can locate the corresponding "Coupon Package" control according to this interface hot word "Coupon Package". In the embodiment where the first voice matches at least one interface hot word, the first control that matches the first voice is a text control.

[0127] In other embodiments, the first control that matches the first voice may also be a picture control; such as Figure 2A The picture control 109 corresponding to the return icon shown.

[0128] In the embodiment where the electronic device simultaneously displays multiple windows, the second interface is the interface displayed in the focused window.

[0129] When the electronic device executes the simulated click instruction corresponding to the first voice, in order to avoid the situation where the simulated click is unsuccessful, it can first determine whether the area to be clicked is clickable, such as S208.

[0130] S208. Determine whether the first control is clickable.

[0131] It should be noted that what is judged in S208 is whether the first control itself has the property of being clickable.

[0132] As can be seen from the above embodiments, in some embodiments, after the electronic device displays the second interface, it can parse the display content of the second interface to obtain and save the control information on the second interface. In this embodiment, the above S208 may specifically include: obtaining the control information of the first control, and determining whether the first control is clickable according to the control information of the target text control.

[0133] If the judgment result of S208 is yes, it means that the first control is clickable. At this time, a simulated click operation can be directly performed within the area corresponding to the first control. S213 and S214 can be executed.

[0134] If the judgment result of S208 is no, it means that the first control is not clickable. At this time, S209 can be executed.

[0135] S209. Determine whether there is a second control belonging to the same control group as the first control.

[0136] In some embodiments, the control information obtained by the electronic device through interface parsing may further include: the control group to which the control belongs. In this embodiment, the electronic device can determine whether there are sibling controls belonging to the same control group as the first control by querying the control information; as well as the type of the sibling control and whether it is clickable.

[0137] Among them, the second control can be any type of control, such as a text control or a picture control.

[0138] In the embodiments of the present application, the above S209 may specifically include: searching for whether there are sibling controls belonging to the same control group as the first control. In one example, if there are sibling controls belonging to the same control group as the first control in the second interface, after S209, the control information such as the type of the sibling control and whether it is clickable can also be obtained.

[0139] If the judgment result of S209 is no, it means that there are no sibling controls belonging to the same control group as the first control. In this case, the electronic device may no longer execute the instruction corresponding to the first voice. Further, in some examples, the electronic device can issue a prompt message (such as displaying a prompt message on the display screen) to prompt the user that it is not clickable.

[0140] If the judgment result of S209 is yes, it means that there are sibling controls belonging to the same control group as the first control. In this way, the instruction corresponding to the first voice can be executed based on the sibling control. Before that, the electronic device also needs to judge whether the sibling control is clickable, such as S210.

[0141] The interface displayed by the electronic device includes a control group with two or more controls, usually including an image control and a text control. In one scenario, the text control in the same control group is not clickable, and the image control is clickable. In some embodiments, the above S209 may specifically include: determining whether there is a sibling image control corresponding to the first control. In this embodiment, if there is a sibling image control corresponding to the first control, then S210 is executed.

[0142] S210. Determine whether the second control is clickable.

[0143] For the specific implementation process of determining whether the second control is clickable, reference may be made to the description of determining whether the first control is clickable.

[0144] If the judgment result of S210 is negative, it means that the second control is not clickable. In this case, the electronic device may no longer execute the instruction corresponding to the first voice. Further, the electronic device may issue a prompt message, which is used to prompt the user that it is not clickable, such as S215. Exemplarily, the electronic device issues a prompt message, which may specifically be by displaying a text prompt message or a picture prompt message on the display screen; and / or by emitting a voice prompt message through a speaker; and / or, by emitting a prompt message in a vibration mode.

[0145] If the judgment result of S210 is positive, it means that the second control is clickable. In this way, the electronic device may execute the instruction corresponding to the first voice based on this second control.

[0146] S211. Obtain the area corresponding to the second control.

[0147] In some embodiments, the electronic device may obtain the area corresponding to the second control by obtaining the control information of the second control. Obtaining the area corresponding to the control may specifically mean obtaining the coordinate position of the area corresponding to the control on the display interface.

[0148] The control may be presented in various shapes on the display interface of the electronic device, and the most common is a rectangular control. Taking the rectangular control as an example, the following method may be used to determine the coordinate position of the area corresponding to the control on the display interface: obtain the coordinates of two corner points on a diagonal of the control (such as the coordinates of the upper left corner point and the lower right corner point, or the lower left corner point and the upper right corner point). If the second control is a rectangular control, then the above S211 may specifically obtain the coordinates of two corner points on a diagonal of the second control, such as the coordinates of the upper left corner point and the lower right corner point.

[0149] Among them, the coordinates of a point on the display interface of the electronic device can be represented by the coordinates in the screen coordinate system. In some examples, the upper left corner of the screen is the origin coordinate (0, 0) of the screen coordinate system. The positive direction of the X-axis extends to the right from the origin, and the positive direction of the Y-axis extends downward from the origin.

[0150] If the control is of other shapes, such as circular, the coordinate position of the area corresponding to the control on the display interface can be determined by obtaining the center point coordinate and radius of the control. In other embodiments, if the control is in the shape of an ellipse, polygon, etc., the coordinate position of the area corresponding to the control on the display interface can also be determined by other means.

[0151] S212. Perform a simulated click operation based on the area corresponding to the second control.

[0152] In some embodiments, the above S212 may specifically include: performing a simulated click operation within the area corresponding to the second control. Further, performing a simulated click operation within the area corresponding to the second control may specifically include: obtaining the center position of the area corresponding to the second control as the simulated click position, and performing a simulated click operation at the simulated click position.

[0153] In other embodiments, the above S212 may specifically include: splicing the areas corresponding to the first control and the second control respectively to obtain a spliced area. Then, perform a simulated click operation within the spliced area. In some examples, performing a simulated click operation within the spliced area may specifically include: obtaining the center position of the spliced area as the simulated click position; and then performing a simulated click operation at the simulated click position.

[0154] The specific implementation process of the electronic device performing a simulated click operation at the simulated click position may refer to the description in related technologies and will not be elaborated in the embodiments of the present application.

[0155] Please refer to Figure 4 , in some examples, the mobile phone 100 displays the "My" interface 301 of the video application, and the user inputs the voice "Coupon Package" to the electronic device. In response to receiving the voice "Coupon Package", the mobile phone 100 searches for the text control that matches the voice "Coupon Package" in the "My" interface 301, that is, the text control 302. Then, the mobile phone 100 determines whether the text control 302 is clickable. After determining that the text control 302 is not clickable, the mobile phone 100 can search for the second control in the "My" interface 301 that belongs to the same control group as the text control 302. In one example, the mobile phone 100 finds that the second control that belongs to the same control group as the text control 302 is the picture control 303. After the mobile phone 100 determines that the picture control 303 is clickable, it can execute the instruction corresponding to the voice "Coupon Package" based on the picture control 303, that is, the simulated click instruction.

[0156] In some embodiments, the mobile phone 100 may obtain the area corresponding to the picture control 303 and execute a simulated click instruction within the area corresponding to the picture control 303.

[0157] In other embodiments, the mobile phone 100 may splice the area corresponding to the picture control 303 with the area corresponding to the text control 302, and then execute a simulated click instruction based on the spliced area obtained by splicing; that is, execute a simulated click operation within the spliced area.

[0158] Exemplarily, after the mobile phone 100 executes the simulated click instruction corresponding to the voice "coupon package" based on the picture control 303, it can open the coupon package and display the coupon package interface 304.

[0159] In the technical solution proposed in the embodiments of the present application, after the electronic device receives the first voice containing the control click intention, if it is determined that the first control corresponding to the control click intention is not clickable, it searches for a second control in the same control group as the first control on the current display interface. If there is a clickable second control in the same control group as the first control on the current display interface, the instruction corresponding to the first voice can be executed based on the second control. In this way, when the recognition of the received voice intention is accurate, the possibility that the electronic device cannot realize the true intention of the voice input by the user can be reduced, and the possibility of executing an accurate simulated click operation in response to the voice input by the user can be increased.

[0160] When the judgment result of S208 is yes, the electronic device can directly perform a simulated click operation on the first control to realize the user intention. As Figure 3 shown in S213 and S214.

[0161] S213. Obtain the area corresponding to the first control.

[0162] S214. Perform a simulated click operation within the area corresponding to the first control.

[0163] As Figure 2B shown in the example, the first control corresponding to the voice "open history" input by the user is the "history" control 111. The "history" control 111 is clickable. At this time, the mobile phone 100 can directly respond to the voice input by the user and perform a simulated click operation on the "history" control 111. In this way, the user intention can also be realized.

[0164] When the judgment result of S209 is no or the judgment result of S210 is no, the electronic device cannot perform a simulated click operation on the first control. At this time, the electronic device can issue a prompt message, such as S215.

[0165] S215. Send a prompt message.

[0166] This prompt message is used to prompt the user that the first control is not clickable. In this way, the user can be reminded of the response result of the electronic device to the first voice.

[0167] In addition, there are many forms of controls displayed on the electronic device, and some control groups may include more than three controls. In some other embodiments, when the judgment result in S209 is yes, it is possible that there are more than two second controls belonging to the same control group as the first control. In this case, the electronic device cannot determine which one of the peer controls needs to perform the simulated click operation. If one of the peer controls is selected to perform the simulated click operation, it may not conform to the user's intention. Therefore, in some embodiments, after S209 and before S210, the method may further include: judging whether the number of second controls is 1. In this embodiment, when the electronic device determines that the number of second controls is 1, it then executes S210 and the subsequent processes. That is, only when the first control includes only one second control, will it be judged whether the second control is clickable. And, when the second control is clickable, the electronic device will perform the simulated click operation based on the second control.

[0168] In some other embodiments, if the electronic device determines that there are more than two second controls belonging to the same control group as the first control, the electronic device may not respond to the first voice. Further, the electronic device may send a prompt message, such as Figure 3 shown in S215.

[0169] In the technical solution proposed in the embodiment of the present application, when the first control matching the first voice is not clickable, and only one second control belonging to the same control group as the first control is found, the simulated click operation is then performed based on the second control. In this way, it can be ensured that the electronic device performs the simulated click operation at the accurate position, and the execution result is more in line with the user's intention.

[0170] In some scenarios, the electronic device can display multiple windows simultaneously in a stacked manner. In the scenario where the electronic device performs a simulated click operation in response to the user's voice, the last clicked position may be blocked by a window such as a floating window. If this position is blocked, then if the electronic device performs the simulated click operation, it may click on other windows, causing a problem of incorrect simulated clicks.

[0171] Such as Figure 5AAs shown, the mobile phone 100 displays a settings interface 401 and a calculator floating window 402 displayed as a floating window. The settings interface 401 includes a WLAN control 403, a Bluetooth control 404, and a mobile network control 405. Among them, some controls on the settings interface 401 are blocked by the calculator floating window 402. Specifically, the WLAN control 403 in the settings interface 401 is completely blocked; the Bluetooth control 404 is partially blocked; the mobile network control 405 is not blocked.

[0172] For Figure 5A the WLAN control 403 and the Bluetooth control 404 in the shown settings interface 401, if the electronic device performs a simulated click operation on these two positions, it may click on the calculator floating window 402. As Figure 5B shown, when the user inputs the voice "Bluetooth" and the mobile phone 100 executes the instruction corresponding to this language, a simulated click operation is performed at the position 404a. At this time, the mobile phone 100 will click on the control corresponding to the number "0" in the calculator floating window 402. Thus, the calculator floating window 402 of the mobile phone 100 is updated to 402a, which displays the selected number "0". This will cause a problem that the simulated click of the electronic device goes wrong.

[0173] Or, Figure 4 the picture control 303 shown may also be blocked by a floating window or the like, resulting in the electronic device being unable to accurately perform a simulated click operation on the picture control 303.

[0174] Based on this, an embodiment of the present application also proposes a voice control method, which can also be applied to an electronic device supporting voice input. Figures 6 - 8 Shows the specific implementation manner of this voice control method.

[0175] Figure 6 Is a flowchart of the voice control method. In this embodiment, the method includes:

[0176] S500. Display a first interface, and the first interface includes a first switch.

[0177] S501. Receive an operation to turn on the first switch.

[0178] S502. In response to the operation of turning on the first switch, enable the voice control function.

[0179] S503. Display a second interface.

[0180] S504. Receive a first voice.

[0181] S505. Determine the user intention corresponding to the first voice.

[0182] S506. Determine whether the user's intention is a control click intention.

[0183] S507. Find the control to be clicked that matches the first voice in the second interface.

[0184] In an embodiment where the electronic device simultaneously displays multiple windows, the second interface is the interface displayed by the focused window.

[0185] S508. Determine whether the control to be clicked is blocked.

[0186] The controls displayed by the electronic device on the display interface may be blocked by windows such as floating windows and pop-up windows, resulting in the electronic device being unable to perform a simulated click operation on the control. Combining Figure 5A As shown in the example, the controls displayed by the electronic device may be blocked by a floating window. After the electronic device determines the control to be clicked that matches the first voice, it can obtain whether there is a floating window on the electronic device.

[0187] In some other embodiments, the electronic device can also analyze the display content of the current display interface to determine whether the control to be clicked is blocked.

[0188] Combining Figure 5A As shown in the example, there are two cases where the control is blocked. One case is that it is completely blocked, and the other is that it is partially blocked. It can be understood that when the control is completely blocked, the electronic device cannot perform a simulated click operation on the control. When the control is partially blocked, the electronic device can perform a simulated click operation on the control, but it may click on other windows, resulting in an incorrect simulated click. Therefore, the electronic device can also determine whether the control to be clicked is completely blocked, such as S509.

[0189] S509. Determine whether the control to be clicked is completely blocked.

[0190] The control being completely blocked means that the area corresponding to the control is completely covered by other windows. In Figure 5A As shown in the example, the WLAN control 403 is completely blocked; the Bluetooth control 404 is not completely blocked (i.e., partially blocked).

[0191] If the judgment result of S509 is yes, it means that the control to be clicked is completely blocked. In this case, the electronic device cannot perform a simulated click operation on the control to be clicked. At this time, the electronic device can execute S510.

[0192] S510. Send a prompt message.

[0193] This prompt message is used to prompt the user that the control to be clicked cannot be clicked. In this way, the user can be reminded of the response result of the first voice.

[0194] If the result of the determination in S509 is negative, it indicates that the control to be clicked is not completely blocked. In this case, the electronic device can perform a simulated click operation in the unblocked area of the control to be clicked. To avoid errors in the simulated click, the electronic device can perform the simulated click operation in the unblocked area.

[0195] S511. Re-determine the simulated click area.

[0196] Since the control to be clicked is not completely blocked, it means that there is still a part of the control to be clicked where the simulated click operation can be performed. In some embodiments, the above S511 may specifically include: obtaining the area of the unblocked part of the control to be clicked as the simulated click area.

[0197] In one example, the above obtaining the area of the unblocked part of the control to be clicked may specifically include: determining the overlapping area between the occlusion window and the control to be clicked. Comparing the area corresponding to the control to be clicked with the overlapping area to determine the area of the unblocked part of the control to be clicked. In this way, the simulated click area can be quickly determined.

[0198] When determining the simulated click area, the border of the simulated click area can be determined. Taking the control to be clicked and the occlusion window as rectangles as an example, the determined simulated click area should also be a rectangle. In this embodiment, to determine the simulated click area, the coordinates of two corner points on a diagonal of the simulated click area can be specifically determined. Exemplarily, the above S511 can specifically determine the coordinates of the upper left corner point and the lower right corner point of the simulated click area, or determine the coordinates of the upper right corner point and the lower left corner point of the simulated click area. In other embodiments, the electronic device can also determine the border of the simulated click area by other means.

[0199] S512. Perform a simulated click operation based on the simulated click area.

[0200] In some embodiments, the above S512 may specifically include: obtaining the center position of the simulated click area and performing a simulated click operation at the center position of the simulated click area.

[0201] Taking the example that the electronic device performs a simulated click operation at the center position of the area corresponding to the control to be clicked, in the above embodiment, when the judgment result of S509 is no, S511 and S512 are directly executed. However, in actual situations, one case where the judgment result of S509 is no is that the control to be clicked is partially blocked, and the center position of the area corresponding to the control to be clicked is not blocked. At this time, the electronic device does not need to re-determine the simulated click area, but can still directly perform the simulated click operation at the center position of the area corresponding to the control to be clicked. Since the center position of the area corresponding to the control to be clicked is not blocked, when the electronic device performs the simulated click operation, there will be no problem of simulated click error.

[0202] Therefore, in some other embodiments, when the judgment result of S509 is no, before S511 and S512, the above method may further include: determining whether the center position of the control to be clicked is blocked. In this embodiment, the electronic device may execute S511 and S512 when the center position of the control to be clicked is blocked. In another example, when the judgment result of S509 is no and the center position of the control to be clicked is not blocked, the electronic device may directly perform the simulated click operation at the center position of the control to be clicked.

[0203] In addition, if the judgment result of S508 is no, it means that the control to be clicked is not blocked. In this case, the electronic device may directly perform a simulated click on the control to be clicked, such as S513 and S514.

[0204] S513. Obtain the area corresponding to the control to be clicked.

[0205] S514. Perform a simulated click operation in the area corresponding to the control to be clicked.

[0206] It should be noted that for the specific implementation processes of some steps in the above S501 - S514, reference may be made to the descriptions of the corresponding steps of S201 - S214 in the above embodiment; details are not repeated here.

[0207] In the technical solution proposed in the embodiments of the present application, if the electronic device detects that the control to be clicked that matches the first voice is completely blocked, the simulated click operation is no longer performed, avoiding the problem of simulated click error caused by clicking on other windows during the simulated click operation. And if the electronic device detects that the control to be clicked is not completely blocked, the simulated click area is re-determined and then the simulated click area is executed. In this way, it can be ensured that when performing the simulated click, the control to be clicked is accurately clicked, avoiding the problem of simulated click error in this scenario. Thus, the electronic device can perform the simulated click operation at the accurate position in response to the voice containing the control click intention input by the user, thereby accurately implementing the user intention.

[0208] Taking the floating window that may block the control to be clicked as an example, as Figure 7A shown, the above S508 may include S601 - S604:

[0209] S601. Query whether there is a floating window.

[0210] In some embodiments, the application corresponding to the voice control function may register a floating window monitor. When the electronic device displays a floating window, the application corresponding to the voice control function may be notified through this floating window monitor. In this way, when the application corresponding to the voice control function of the electronic device needs to perform a simulated click operation in response to the voice input by the user, it can query whether there is a floating window on the current display interface of the electronic device.

[0211] In this embodiment, if the judgment result of S601 is negative, it means that there is no floating window on the electronic device currently. That is to say, the control to be clicked is not blocked by the floating window. In this case, the electronic device may execute S513 and S514.

[0212] If the judgment result of S601 is positive, it means that there is a floating window. Next, the electronic device may determine whether the control to be clicked is blocked by the floating window according to the area corresponding to the floating window and the area corresponding to the control to be clicked, that is, S602 - S604.

[0213] S602. Determine whether the floating window is the focus window.

[0214] In some embodiments, the electronic device may obtain the window identifier of the current focus window, and then determine whether the floating window is the focus window according to the window identifier of this focus window.

[0215] In some examples, the window identifier of the focus window may also be the border of this focus window. In this embodiment, the electronic device may compare the border of the floating window with the border of the focus window to determine whether the floating window is the focus window. In this way, it is convenient to confirm whether the control to be clicked is blocked by the floating window.

[0216] In some other examples, the window identifier of the focus window may specifically be the application name of the application displayed in this focus window or the activity name of the activity. In this embodiment, the electronic device may compare the application name or activity name corresponding to the floating window with the focus window to determine whether the floating window is the focus window. In this way, it is convenient to confirm whether the control to be clicked is blocked by the floating window.

[0217] If the floating window is the focus window, it means that the control to be clicked is a control on the floating window. In this case, the electronic device can directly perform a simulated click operation on the control to be clicked. That is, if the judgment result of S602 above is yes, the electronic device can execute S513 and S514.

[0218] If the floating window is not the focus window, it means that the control to be clicked is not a control on the floating window, and the control to be clicked may be blocked by the floating window. In this case, the electronic device needs to further determine whether the control to be clicked is blocked by the floating window. Exemplarily, the electronic device can execute S603 and S604.

[0219] S603. Obtain area 1 corresponding to the control to be clicked and area 2 corresponding to the floating window.

[0220] For the specific implementation of obtaining the area corresponding to the control and the area corresponding to the floating window, reference can be made to the descriptions in the above embodiments and related technologies.

[0221] S604. Compare area 1 and area 2 to determine whether the control to be clicked is blocked by the floating window.

[0222] In this embodiment, the display interface of the electronic device includes a floating window. The electronic device can determine whether the control to be clicked is blocked by the floating window according to whether there is an overlapping area between the area corresponding to the floating window and the area corresponding to the control to be clicked. The floating window is usually floating and displayed on the background interface. If there is an overlapping area between the area corresponding to the floating window and the area corresponding to the control to be clicked, it means that the control to be clicked is blocked by the floating window.

[0223] If the judgment result of S604 is no, it means that although there is a floating window on the current display interface of the electronic device, the control to be clicked is not blocked by the floating window. In this case, the electronic device can also directly perform a simulated click operation on the target text, that is, execute S513 and S514.

[0224] If the judgment result of S604 is yes, the electronic device can continue to execute S509 to determine whether the control to be clicked is completely blocked.

[0225] Figure 7B This is a schematic diagram of a scenario example of the voice control method according to an embodiment of the present application. The mobile phone 100 displays a setting interface 401. In response to the voice "Bluetooth" input by the user, the mobile phone 100 performs a simulated click operation at position 402b. After that, the mobile phone 100 opens a Bluetooth sub-menu interface 402c. In this way, the user's intention is accurately realized, and the problem of incorrect simulated clicks is avoided.

[0226] Figure 7CSchematic diagram of a scenario example of the voice control method according to an embodiment of the present application. The mobile phone 100 displays a setting interface 401. In response to the voice "WLAN" input by the user, the mobile phone 100 detects that the WLAN control 403 is completely blocked by suspension. Therefore, the mobile phone 100 can display a prompt message 403a.

[0227] Figure 8 Shows Figure 5A In the example of, the positional relationship between the partially blocked Bluetooth control 404 and the calculator floating window 402. In some embodiments, in combination with Figure 8 Taking the case where both the control to be clicked and the floating window are rectangles as an example, the above S511 can be specifically implemented in the following manner: In this embodiment, the upper left corner of the screen is used as the origin coordinate (0,0); the positive direction of the X-axis extends to the right from the origin, and the positive direction of the Y-axis extends downward from the origin, and the established coordinate system is denoted as the screen coordinate system.

[0228] In the quadrant where x>0 and y>0, the area corresponding to the control to be clicked (i.e., the Bluetooth control 404) is represented by coordinates: the upper left corner point A(x1,y1), and the lower right corner point B(x2,y2). The area corresponding to the floating window (i.e., the calculator floating window 402) is represented by coordinates: the upper left corner point C(x3,y3), and the lower right corner point D(x4,y4).

[0229] There is an overlapping part between the area corresponding to the control to be clicked and the area corresponding to the floating window ( Figure 8 The shaded area 404-1 in), and the control to be clicked is not completely blocked. The area of the simulated click to be determined is the coordinates of the upper left corner point and the lower right corner point of the largest rectangle area where the control to be clicked is not blocked. In some examples, the simulated click area can be represented by coordinates:

[0230] The upper left corner point M(max(x1,x3),max(y1,y3)), and the coordinates of the lower right corner point N(min(x2,x4),min(y2,y4)). In Figure 8 In the example shown, the coordinates of the calculated simulated click area (i.e., Figure 8 The area 404-2 in) are M(x1,y4), N(x2,y2).

[0231] It should be noted that Figure 8 The method of determining the simulated click area shown is only one example. In other embodiments, the electronic device can also determine the simulated click area by other means.

[0232] In the technical solution proposed in the embodiment of the present application, the electronic device determines whether there is a floating window in the current display interface, whether the floating window is the focus window, and in the case where the floating window is not the focus window, compares the area corresponding to the floating window with the area corresponding to the control to be clicked, so as to determine whether the control to be clicked is blocked by the floating window. In this way, it is possible to quickly and conveniently determine whether the control to be clicked is blocked, which is convenient for the electronic device to perform a simulated click operation in response to the first voice.

[0233] Figure 9A Fig. shows the flow of another voice control method proposed in the embodiment of the present application.

[0234] In Figure 9A In the shown embodiment, after the electronic device responds to the operation of receiving the user to turn on the first switch, the voice control function is turned on. When the electronic device determines that the user intention corresponding to the received voice is a control click intention, it can first determine whether the first control is clickable. The process of the electronic device determining whether the first control is clickable can refer to the processes of S207 - S211 and S215 in the method shown in Figure 3 In the case where it is determined that the first control is clickable, or there is a second clickable control in the first control, it is possible to continue to determine whether the first control (or the second control) is blocked, that is, Figure 6 the processes of S506 - S510 in the method shown in. Finally, if the first control (or the second control) (i.e., the control to be clicked) is not blocked, or not completely blocked, the electronic device can perform a simulated click operation in response to the first voice, that is, Figure 6 the processes of S511 - S512 or S513 - S514 in the shown process. In another example, when the first control is not clickable and there is no sibling control, or the first control (or the second control) is completely blocked, the electronic device can send a prompt message to prompt that it is not clickable.

[0235] In the technical solution proposed in the embodiment of the present application, when receiving a voice including a control click intention, the electronic device first determines whether the first control matching the voice is clickable, and then determines whether it is blocked. In this way, when the voice intention recognition is accurate, the possibility that the electronic device cannot realize the true intention of the voice input by the user can be reduced, and the possibility of performing an accurate simulated click operation in response to the voice input by the user can be improved.

[0236] In some other embodiments, in Figure 9AIn the voice control method shown, when it is determined that the first control matching the first voice is clickable, the electronic device will perform an occlusion judgment on the first control. Further, if the electronic device determines that the first control is completely occluded, it means that the electronic device cannot perform a simulated click operation of the voice click instruction for the first control. In this case, the electronic device can also re-determine the control to be clicked. Specifically, the electronic device can find a second control belonging to the same control group as the first control and determine the second control as the new control to be clicked. Then, the electronic device determines whether the second control is occluded by the floating window.

[0237] If the second control is also completely occluded, it means that the electronic device cannot respond to the first voice to perform a simulated click operation.

[0238] If the second control is not occluded, it means that the electronic device can perform a simulated click operation on the second control in response to the first voice.

[0239] If the second control is partially occluded, the electronic device can re-determine the simulated click area based on the areas corresponding to the floating window and the second control respectively. Finally, the electronic device performs a simulated click operation within the re-determined simulated click area.

[0240] In this embodiment, the specific implementation process of the electronic device finding the second control belonging to the same control group as the first control can refer to the description in the above embodiment.

[0241] Take Figure 9B as an example. The mobile phone 100 displays the settings interface 406 and also includes a calculator floating window 407. In this example, the text control 408 in the WLAN settings option is the above-mentioned first control. When the mobile phone 100 judges the click attribute of the text control 408, it determines that the click attribute of the text control 408 is clickable. Then, the mobile phone 100 performs an occlusion judgment on the text control 408. It can be determined that the text control 408 is completely occluded by the calculator floating window 407. In this case, the mobile phone 100 can find a second control belonging to the same control group as the text control 408, such as Figure 9B the picture control 409 in the WLAN settings option shown. Then, the mobile phone 100 performs an occlusion judgment on the picture control 409. In one example, as Figure 9B shown, the picture control 409 is not occluded by the floating window. At this time, the mobile phone 100 can perform a simulated click operation on the picture control 409, and can also achieve the user's intention to open the WLAN settings option and display the WLAN settings sub-menu.

[0242] It can be understood that in another example, if the result obtained by the mobile phone 100 after determining the occlusion of the picture control 409 is that the picture control 409 is partially occluded, the mobile phone 100 can re-determine the simulated click area based on the areas corresponding to the calculator floating window 407 and the picture control 409 respectively, and then perform a simulated click operation on the re-determined simulated click area.

[0243] In another example, if the result obtained by the mobile phone 100 after determining the occlusion of the picture control 409 is that the picture control 409 is completely occluded, the mobile phone 100 can send a prompt message. This prompt message is used to prompt that it is not clickable.

[0244] In the technical solution proposed in the embodiments of the present application, when the electronic device performs a simulated click operation in response to voice, it simultaneously considers whether the control to be clicked is clickable and whether it is occluded. In this way, it can be further ensured that the control to be clicked can execute a click event. Thus, when the electronic device accurately recognizes the voice intention, the possibility that the electronic device cannot realize the true intention of the voice input by the user is reduced, and the possibility of performing an accurate simulated click operation in response to the voice input by the user is increased.

[0245] Figures 10 - 13 Shows the implementation of Figures 3 - 4 The software architecture of the electronic device implementing the voice control method shown, as well as the interaction process of each module of the electronic device when implementing the voice control method.

[0246] Figure 10 Is the software architecture of the electronic device in some embodiments. In some embodiments, the software system of the electronic device can adopt a layered architecture, an event-driven architecture, a microkernel architecture or a cloud architecture. The embodiments of the present application take the layered architecture System as an example to exemplarily illustrate the software structure of the electronic device.

[0247] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the System is divided into four layers, from top to bottom are the application layer, the application framework layer, the Android runtime and the system library, and the kernel layer.

[0248] The application layer may include a series of application packages (APKs). For example, applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc. In the embodiments of this application, the application layer includes a voice control APK for providing the voice control function of the electronic device. The voice control APK includes an interface monitoring module, an interface content acquisition module (content sensor), an interface parsing module (view fetcher), and an application interaction module, etc. Among them, the interface monitoring module is used to monitor processes such as the startup, exit, and switching of the application interface. The interface content acquisition module is used to acquire the top-level activity. The interface parsing module is used to parse the top-level activity to obtain the content of the current display page (including pictures and / or text). The application interaction module is used to manage the process of the application executing instructions according to voice.

[0249] The application layer further includes a voice processing engine for parsing, recognizing, and processing voices. Among them, the voice parsing module is used to convert voice into text; the voice parsing module may belong to the ASR engine. The voice recognition module is used to understand and recognize the semantics of the text and determine whether it matches the interface text; the voice recognition module may belong to the NLU engine. The instruction mapping module is used to convert the recognized semantics into machine-executable instructions; the instruction mapping module may belong to the DM engine.

[0250] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0251] As Figure 11 shown, the application framework layer may include a window manager, a content provider, a view system, a resource manager, an activity manager service (AMS), a package manager service (PMS), and a multi-mode control module, etc.

[0252] AMS is mainly responsible for the startup, switching, scheduling of the four major components in the system, and the management and scheduling of application processes, etc. Its responsibilities are similar to the process management and scheduling module in the operating system. When a process startup or component startup is initiated, the request will be passed to AMS through the inter-process communication (binder) mechanism, and AMS will then make unified processing.

[0253] The multi-mode control module is used to manage the execution of voice instructions. For example, sending the instruction to the application to make the application execute the instruction.

[0254] The system library may include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing library (e.g., openGL ES), 2D graphics engine (e.g., SGL), etc.

[0255] Android runtime is responsible for the scheduling and management of the Android system. Android runtime includes core libraries and a virtual machine.

[0256] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0257] The kernel layer is the layer between hardware and software. The kernel layer may include display drivers, sensor drivers, microphone drivers, Wi-Fi drivers, etc.

[0258] Combined Figure 10 , Figure 11 shows an interaction schematic diagram of each module of the electronic device when implementing the voice control method provided in the embodiments of the present application.

[0259] AMS can monitor the lifecycle of the activity corresponding to each interface of each application on the electronic device.

[0260] In some examples, the interface content acquisition module may send a request to AMS. In response to this request, AMS returns the top-level activity to the interface content acquisition module. In some embodiments, the interface content acquisition module sending a request to AMS may specifically include: the interface content acquisition module sending a request to AMS through activitymanagerEx.requestContentNode.

[0261] In some other examples, the interface content acquisition module may register a listener with AMS. When AMS senses a change in the lifecycle of the application, it can notify the top-level activity to the interface content acquisition module through a callback.

[0262] After receiving the top-level activity sent by AMS, the interface content acquisition module sends the top-level activity to the interface parsing module. The interface parsing module parses the top-level activity to obtain the display content of the electronic device, and the text controls and / or picture controls included therein.

[0263] Further, the interface parsing module may save the text controls and / or picture controls obtained by parsing the top-level activity to a cache. In some embodiments, the top-level activity obtained by the interface parsing module may include: the control information of each control group. Specifically, the control information of the control group may include: the control information of each control included in the control group. Taking a control group including a picture control and a text control as an example, the control information of the control group may include: Control group [coordinate 1, coordinate 2]: Picture control [coordinate 3, coordinate 4], Text control [coordinate 5, coordinate 6].

[0264] In one example, the control information of the picture control in the control group may at least include the content in Table 1.

[0265] Table 1

[0266]

[0267] The control information of the text control in the control group may at least include the content in Table 2.

[0268] Table 2

[0269]

[0270] Among them, the label is used to indicate the serial number of the control in the control group.

[0271] The interface parsing content parses the above top-level activity, and the control information of each control or control group on the currently displayed interface can be obtained. In some embodiments, the interface parsing module may save the control information of the text control and / or picture control to a cache. Among them, the control information may include: the area where the control is located (a rectangular control can be represented by the coordinates of the upper left corner point and the lower right corner point), whether the control is clickable, and the control group to which the control belongs.

[0272] After the user inputs voice to the electronic device, the microphone of the electronic device collects the voice. The microphone inputs the collected voice to the application interaction module of the voice control APK. The application interaction module is responsible for distributing the received voice; in some examples, the application interaction module may distribute the voice to the voice parsing module in the voice processing engine, and the voice parsing module parses the voice input by the user.

[0273] The voice parsing module parses the voice, and the parsed text corresponding to the voice can be obtained. Then, the voice parsing module sends the text corresponding to the parsed voice to the voice recognition module.

[0274] The speech recognition module performs semantic understanding and recognition on the text corresponding to the speech to obtain the intention corresponding to the speech; that is, the user's intention, which is the operation that the user hopes to control the electronic device to perform through the speech. After that, the speech recognition module can send the intention corresponding to the speech to the instruction mapping module.

[0275] The instruction mapping module maps and converts the intention corresponding to the speech into an instruction executable by the machine (denoted as instruction a).

[0276] In some embodiments, after the instruction mapping module converts the user intention into instruction a, it can determine whether instruction a is a control click instruction. If the obtained instruction a is not a control click instruction, the instruction mapping module can directly send instruction a to the corresponding application through the multimode control module to notify the application to execute instruction a.

[0277] If the obtained instruction a is a control click instruction, the instruction mapping module can send the control click instruction to the interface parsing module. Among them, the control click instruction carries the text that matches the parsed text corresponding to the first speech (i.e., the hit text).

[0278] In some other embodiments, after receiving the user intention sent by the speech recognition module, the instruction mapping module can also judge the user intention to determine whether the user intention is a control click intention. And, when the instruction mapping module determines that the user intention is a control click intention, after converting the user intention into instruction a, it can send instruction a to the interface parsing module. If it is determined that the user intention is not a control click intention, after converting the user intention into instruction a, the instruction mapping module can directly send it to the corresponding application through the multimode control module to notify the application to execute instruction a.

[0279] Taking the user input speech as a control click intention, that is, instruction a is a control click instruction as an example for illustration. In the embodiments of the present application, after receiving the control click instruction sent by the instruction mapping module, the interface parsing module can search for the control information of the control corresponding to the hit text (i.e., the first control, hereinafter denoted as the target text control) in the cache according to the hit text carried in the control click instruction.

[0280] In the embodiments of the present application, the interface parsing module obtains the control information of the target text control from the cache and judges whether the target text control is clickable. In some examples, if the target text control is not clickable, the interface parsing module can search for a second control (hereinafter denoted as the target sibling control) in the cache that belongs to the same control group as the target text control.

[0281] If the interface parsing module finds the target sibling control corresponding to the target text control, it can also query the number of target sibling controls, the area corresponding to the target sibling control, and whether the target sibling control is clickable.

[0282] In some embodiments, when the interface parsing module determines that there is a target sibling control belonging to the same control group as the target text control and it is clickable, it can execute instruction a based on the area corresponding to the found target sibling control. In another example, if the interface parsing module determines that there is no target sibling control belonging to the same control group as the target text control, or all sibling controls belonging to the same control group as the target text control are not clickable, it determines that the control click instruction cannot be executed.

[0283] In some other embodiments, when the interface parsing module determines that there is a target sibling control belonging to the same control group as the target text control, there is only one target sibling control, and it is clickable, it can execute instruction a based on the area corresponding to the found target sibling control. In this embodiment, if the interface parsing module determines that there are multiple sibling controls belonging to the same control group as the target text control, or there is only one sibling control belonging to the same control group as the target text control and it is not clickable, it determines that instruction a cannot be executed.

[0284] Furthermore, after the interface parsing module determines that instruction a can be executed, it can notify the corresponding application to execute instruction a through the multi-mode control module.

[0285] After the interface parsing module determines that the control click instruction cannot be executed, it can send a notification message to the application interaction module. In response to receiving this notification message, the application interaction module can issue a prompt message indicating non-clickability.

[0286] The following combines Figure 12 the module interaction to introduce the specific process of the interface parsing module to determine whether the target text control is clickable. In this embodiment, the area corresponding to the control is the border of the control.

[0287] In this embodiment, the interface parsing module includes a clickability judgment sub-module and a sibling control query sub-module. Among them, the clickability judgment sub-module is used to query and obtain whether the target text control is clickable from the cache after receiving the control click instruction sent by the instruction mapping module. The sibling control query sub-module can be used to query the control information corresponding to the target text control.

[0288] If the target text control is clickable, the interface parsing module sends a control click instruction to the multi-module control module through the clickability judgment sub-module, and this control click instruction carries the border of the target text control; that is Figure 12 the first case shown in

[0289] If the target text control is not clickable, the clickability judgment module sends a notification message to the sibling control query sub-module. In response to this notification message, the sibling control query sub-module queries and obtains from the cache the sibling controls that belong to the same control group as the target text control, as well as the click attributes (i.e., whether they are clickable) and borders of the sibling controls.

[0290] Based on the query results, the sibling control query sub-module can determine whether there are sibling controls for the target text control.

[0291] If there is one sibling control for the target text control and this sibling control is clickable, the interface parsing module sends a control click instruction to the multi-module control module through the clickability judgment sub-module, and this control click instruction carries the border of the sibling control; that is Figure 12 the second case shown.

[0292] If the target text control has no sibling controls, the interface parsing module returns a non-clickable notification message to the application interaction module through the clickability judgment sub-module; that is Figure 12 the third case shown.

[0293] If the target text control has no sibling controls, the interface parsing module returns a non-clickable notification message to the application interaction module through the clickability judgment sub-module; that is Figure 12 the fourth case shown.

[0294] If the target text control has multiple sibling controls, the interface parsing module returns a non-clickable notification message to the application interaction module through the clickability judgment sub-module; that is Figure 12 the fifth case shown.

[0295] In response to receiving the notification message sent by the clickability judgment sub-module, the application interaction module can issue a prompt message.

[0296] Figure 13 The timing interaction diagram among the modules is shown during the process of the voice control method proposed in the embodiment of the present application. In this embodiment, it is described by taking the first voice input by the user containing the intention of clicking on a control as an example. In other embodiments, the specific implementation manner of the instruction corresponding to the voice input by the user being other instructions can be referred to Figure 13 the process shown and will not be elaborated in the embodiment of the present application.

[0297] The user turns on the first switch. In response to receiving the operation of the user turning on the first switch, the electronic device enables the voice control function.

[0298] The interface content acquisition module registers a listener with the AMS. It should be noted that the interface content acquisition module can register the listener when the voice control function of the electronic device is turned on. When the AMS senses a change in the application's life cycle, it can notify the top-level activity to the interface content acquisition module through a callback.

[0299] Then, the interface content acquisition module can send the top-level activity to the interface parsing module, notifying the interface parsing module to perform parsing. The interface parsing module parses the top-level activity to obtain the controls and control information of the currently displayed interface. Then, it saves the control information to the cache. Moreover, the interface parsing module sends the text of the currently displayed interface, that is, the interface hot words, to the speech recognition module.

[0300] The user inputs a first voice to the electronic device. Specifically, the microphone collects the voice input by the user. Since the current voice control function is turned on and the recording channel corresponding to the first voice control function is opened, the microphone sends the first voice to the application interaction module of the voice control APK.

[0301] The application interaction module distributes the received first voice to the speech processing engine for processing. Specifically, the application interaction module sends it to the speech parsing module in the speech processing engine. After receiving the first voice, the speech parsing module can parse the first voice to obtain the parsed text. Then, the speech parsing module sends the parsed text to the speech recognition module for intent recognition. The speech recognition module recognizes the intent of the received parsed text to obtain the user intent corresponding to the parsed text. The user intent recognized by the speech recognition module is a control click intent. After that, the speech recognition module can send the control click intent to the instruction mapping module for instruction mapping. After receiving the control click intent, the instruction mapping module maps the control click intent to an instruction that the electronic device can execute, that is, a control click instruction.

[0302] Then, the instruction mapping module sends the control click instruction to the interface parsing module, and the control click instruction carries the hit text.

[0303] After receiving the control click instruction, the interface parsing module sends a query request to the cache. The cache returns the click attributes and borders of the target text control corresponding to the hit text to the interface parsing module. Then, the interface parsing module determines whether the target text control is clickable.

[0304] If the target text control is clickable, the interface parsing module can send the control click instruction to the multi-mode control module, and the control click instruction carries the border of the target text control. The multi-mode control module can notify the corresponding application to execute the control click instruction, that is, perform a simulated click operation within the border of the target text control.

[0305] If the target text control is not clickable, the interface parsing module can cache and send a query request. The cache returns sibling controls, click attributes, and borders to the interface parsing module. Then, the interface parsing module determines whether there are sibling controls. If there are sibling controls (i.e., the above-mentioned target sibling controls), the interface parsing module continues to determine whether the number of sibling controls is 1. If the number of sibling controls is 1, the interface parsing module continues to determine whether this sibling control is clickable. Finally, if this unique sibling control is clickable, the interface can send a control click instruction to the multimode control module; the control click instruction carries the border of this sibling control. The multimode control module can notify the corresponding application to execute the control click instruction.

[0306] If the target text control has no sibling controls, or the number of target sibling controls is greater than 1, or the number of target sibling controls is 1 and this sibling control is not clickable, the interface parsing module notifies the application interaction module that it is not clickable. The application interaction module can issue a prompt message, which is used to prompt that it is not clickable.

[0307] Figures 14 - 17 Shows the implementation Figures 6 - 8 The software architecture of the electronic device implementing the voice control method shown, and the interaction process of each module of the electronic device when implementing the voice control method.

[0308] Figure 14 It is the software architecture of the electronic device in some embodiments. In this embodiment, the voice control APK of the electronic device further includes a floating window monitoring module and an occlusion judgment module. The floating window monitoring module is used to monitor whether the electronic device displays a floating window and obtain the border information of the floating window. The occlusion judgment module is used to judge whether the control to be clicked is occluded.

[0309] Combined with Figure 14 , Figure 15 Shows the interaction schematic diagram of each module of the electronic device when implementing the voice control method provided in the embodiments of the present application.

[0310] The floating window monitoring module registers floating window monitoring with the AMS. After the AMS detects that the electronic device displays a floating window, it notifies the floating window monitoring module of the border of the floating window through a callback. In some examples, the floating window monitoring module registers floating window monitoring with the AMS when the electronic device enables the voice control function.

[0311] The specific implementation process of the user inputting voice, the application interaction module of the voice control APK distributing the voice to the voice processing engine, obtaining a control click instruction, and returning the control click instruction to the interface parsing module refers to Figure 11 the description process of the corresponding steps in

[0312] In this embodiment, after the interface parsing module queries the border of the control to be clicked that matches the hit text, it sends the border of the control to be clicked to the occlusion judgment module.

[0313] The occlusion judgment module can query whether there is a floating window on the electronic device currently through the floating window monitoring module. At the same time, the occlusion judgment module can also query the current focused window through the AMS. Then, the occlusion judgment module can judge whether there is a floating window, and when there is a floating window, combine the current focused window to determine whether the control to be clicked is completely occluded. In the case where there is a floating window and the control to be clicked is completely occluded, the occlusion judgment module notifies the application interaction module that it cannot be clicked. The application interaction module can send a prompt message, and this prompt message is used to prompt that it cannot be clicked. In other cases, the occlusion judgment module sends a control click instruction to the multimode control module, and this control click instruction carries the simulated click area. The multimode control module can notify the corresponding application to execute the control click instruction, that is, perform a simulated click operation in the simulated click area.

[0314] Figure 16 Shows the specific process of the occlusion judgment module judging whether there is a floating window, and when there is a floating window, combining the current focused window to determine whether the control to be clicked is completely occluded. In this embodiment, the occlusion judgment module includes a border comparison sub-module and a border calculation sub-module. Among them, the border comparison sub-module is used to determine whether there is a floating window, and when there is a floating window, compare the control to be clicked and the floating window to determine whether the control to be clicked is completely occluded by the floating window. The border calculation sub-module is used to re-determine the simulated click area when the control to be clicked is not completely occluded.

[0315] After receiving the border of the control to be clicked sent by the interface parsing module, the border comparison sub-module obtains the border of the floating window from the floating window monitoring module. In one example, if there is no floating window currently, the floating window monitoring module returns null to the border comparison sub-module.

[0316] If the border comparison sub-module obtains null from the floating window monitoring module, it is the first case, that is, there is no floating window. The border sub-module can send a control click instruction to the multimode control module, and this control click instruction carries the border of the control to be clicked.

[0317] After obtaining the border of the floating window, the border comparison sub-module obtains the focused window from the AMS and judges whether the floating window is the focused window. If there is a floating window and the floating window is the focused window, it means that the control to be clicked is a control in the floating window, and there is no possibility that the control to be clicked is occluded by the floating window.

[0318] In some other examples, if there is a floating window and the floating window is not the focus window, there is a possibility that the control to be clicked is blocked by the floating window. At this time, the border comparison sub-module can compare the borders of the floating window and the control to be clicked to determine whether the control to be clicked is blocked by the floating window.

[0319] If the border comparison sub-module determines that the control to be clicked is not blocked by the floating window, it is the second case. The border comparison sub-control can also send a control click instruction to the multi-mode control module, and the control click instruction carries the border of the control to be clicked.

[0320] If the border comparison sub-module determines that the control to be clicked is not completely blocked by the floating window, it is the third case. At this time, the border comparison sub-module can notify the border calculation sub-module to re-determine the simulated click area. Specifically, the border comparison sub-module sends the borders of the floating window and the control to be clicked to the border calculation sub-module. The border calculation sub-module calculates and determines the border of the simulated click area according to the borders of the floating window and the control to be clicked. Then, the border calculation sub-module can send a control click instruction to the multi-mode control module, and the control click instruction carries the border of the simulated click area.

[0321] If the border comparison sub-module determines that the control to be clicked is completely blocked by the floating window, it is the fourth case. In this case, the border comparison sub-module notifies the application interaction module that it cannot be clicked. The application interaction module can issue a prompt message.

[0322] Among them, in the first, second, and third cases above, after receiving the control click instruction, the multi-mode control module can notify the corresponding application to execute the control click instruction.

[0323] Figure 17 Shows the timing interaction diagram among the modules during the process of the voice control method proposed in the embodiment of the present application. In this embodiment, the example where the first voice input by the user includes the intention of clicking a control is used for illustration. In other embodiments, the specific implementation manner of the instruction corresponding to the voice input by the user being other instructions can be referred to Figure 13 the shown process and will not be elaborated in the embodiments of the present application.

[0324] The user turns on the first switch. The electronic device responds to the operation of receiving the user turning on the first switch and enables the voice control function.

[0325] The interface content acquisition module registers a listener with the AMS. It should be noted that the interface content acquisition module registering a listener with the AMS can be registered when the electronic device enables the voice control function. When the AMS senses the change in the application life cycle, it can notify the top-level activity to the interface content acquisition module through a callback.

[0326] Then, the interface content acquisition module can send the top-level activity to the interface parsing module, notifying the interface parsing module to perform parsing. The interface parsing module parses the top-level activity to obtain the controls on the currently displayed interface. Then, it saves the control information to the cache. Moreover, the interface parsing module sends the text of the currently displayed interface to the speech recognition module.

[0327] In addition, the floating window monitoring module registers a floating window monitor with the AMS, and the AMS can notify the floating window monitoring module of the floating window's border through a callback.

[0328] The user inputs a first voice to the electronic device. Specifically, the microphone collects the voice input by the user. Since the current voice control function is enabled and the recording channel corresponding to the first voice control function is opened, the microphone sends the first voice to the application interaction module of the voice control APK.

[0329] The application interaction module distributes the received first voice to the speech processing engine for processing. Specifically, the application interaction module sends it to the speech parsing module in the speech processing engine. After receiving the first voice, the speech parsing module can parse the first voice to obtain the parsed text. Then, the speech parsing module sends the parsed text to the speech recognition module for intent recognition. The speech recognition module recognizes the intent of the received parsed text to obtain the user intent corresponding to the parsed text. The user intent recognized by the speech recognition module is a control click intent. After that, the speech recognition module can send the control click intent to the instruction mapping module for instruction mapping. After receiving the control click intent, the instruction mapping module maps the control click intent to an instruction that the electronic device can execute, that is, a control click instruction.

[0330] Then, the instruction mapping module sends the control click instruction to the interface parsing module, and the control click instruction carries the hit text.

[0331] After receiving the control click instruction, the interface parsing module sends a query request to the cache. The cache returns the border of the control to be clicked that matches the hit text to the interface parsing module. Then, the interface parsing module sends the border of the control to be clicked to the occlusion judgment module.

[0332] The occlusion judgment module can send a query request to the floating window monitoring module. The floating window monitoring module returns the border of the floating window / empty to the occlusion judgment module. Then, the occlusion judgment module determines whether there is a floating window. If there is no floating window, the occlusion judgment module can directly send the control click instruction to the multi-mode control module, and the control click instruction carries the border of the control to be clicked.

[0333] If there is a floating window, the occlusion determination module sends a query request to the AMS. In response to the query request, the AMS returns the current focus window of the electronic device to the occlusion determination module. Subsequently, the occlusion determination module can first determine whether the floating window is the focus window. If the floating window is the focus window, the occlusion determination module directly sends a control click instruction to the multimode control module, and the control click instruction carries the border of the control to be clicked.

[0334] If the floating window is not the focus window, the occlusion determination module can continue to determine whether the control to be clicked is occluded by the floating window. If the control to be clicked is not occluded by the floating window, the occlusion determination module directly sends a control click instruction to the multimode control module, and the control click instruction carries the border of the control to be clicked.

[0335] If the control to be clicked is occluded by the floating window, the occlusion determination module can continue to determine whether the control to be clicked is completely occluded. If the control to be clicked is not completely occluded, the occlusion determination module combines the border of the floating window and the border of the control to be clicked to calculate a new simulated click area. Then, the occlusion determination module sends a control click instruction to the multimode control module, and the control click instruction carries the border of the new simulated click area.

[0336] If the control to be clicked is completely occluded, the occlusion determination module notifies the application interaction module that it cannot be clicked. The application interaction module can issue a prompt message for prompting that it cannot be clicked.

[0337] In some embodiments, an electronic device having Figure 14 and Figure 15 the software architecture shown can also be used to implement the voice control method shown in FIG. 9. In this embodiment, Figure 15 in the module interaction shown, the control to be clicked is the first control determined by the interface parsing module to match the first voice, or the second control belonging to the same control group as the first control.

[0338] It should be noted that in the embodiments of the present application, during the process of the electronic device recognizing the received voice and executing the instruction corresponding to the voice, there is the following corresponding relationship: the electronic device parses the display interface to obtain: text controls and picture controls; among them, the text contained in the text controls is saved as the interface hot words; that is, the text of the text controls in the display interface corresponds one-to-one with the interface hot words of the display interface. The electronic device recognizes the voice input by the user to obtain: parsed text - user intention - user intention is converted into an instruction - hits the interface hot word - the text control corresponding to the hit interface hot word (denoted as the first control). That is, the voice corresponds one-to-one with the parsed text, user intention, hit interface hot word, first control, and instruction. Some text controls and other controls jointly belong to a control group. There is a one-to-one correspondence between the text controls and other controls under the same control group.

[0339] Figure 18 Shows a schematic structural diagram of an electronic device 700 provided by an embodiment of the present application.

[0340] The electronic device 700 may include a processor 710, an external memory interface 720, an internal memory 721, a universal serial bus (USB) interface 730, a charging management module 740, a power management module 741, a battery 742, an antenna 1, an antenna 2, a mobile communication module 750, a wireless communication module 760, an audio module 770, a speaker 770A, a receiver 770B, a microphone 770C, a headphone interface 770D, a sensor module 780, a key 790, a motor 791, a camera 792, a display screen 793, and a subscriber identification module (SIM) card interface 794, etc. Among them, the sensor module 780 may include a pressure sensor 780A, a touch sensor 780B, etc.

[0341] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 700. In other embodiments of the present application, the electronic device 700 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0342] The processor 710 may include one or more processing units. For example, the processor 710 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. For example, the processor 710 is used to execute the voice control method in the embodiments of the present application.

[0343] Among them, the controller may be the nerve center and command center of the electronic device 700. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.

[0344] A memory can also be provided in the processor 710 for storing instructions and data. In some embodiments, the memory in the processor 710 is a cache memory. This memory can hold the instructions or data that the processor 710 has just used or recycled. If the processor 710 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses and reduces the waiting time of the processor 710, thus improving the efficiency of the system.

[0345] The USB interface 730 is an interface that conforms to the USB standard specification. Specifically, it can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 730 can be used to connect a charger to charge the electronic device 700, or to transfer data between the electronic device 700 and peripheral devices.

[0346] The external memory interface 720 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 700. The external memory card communicates with the processor 710 through the external memory interface 720 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0347] The internal memory 721 can be used to store computer-executable program code. The executable program code includes instructions. The processor 710 executes various functional applications and data processing of the electronic device 700 by running the instructions stored in the internal memory 721. The internal memory 721 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function (such as a sound playback function, an image playback function, etc.).

[0348] In addition, the internal memory 721 can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0349] The charge management module 740 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charge management module 740 can receive the charging input of the wired charger through the USB interface 730.

[0350] The power management module 741 is used to connect the battery 742, the charge management module 740, and the processor 710. The power management module 741 receives the inputs from the battery 742 and / or the charge management module 740 to supply power to the processor 710, the internal memory 721, the external memory, the display screen 793, the camera 792, and the wireless communication module 760, etc.

[0351] In some other embodiments, the power management module 741 may also be disposed in the processor 710. In some other embodiments, the power management module 741 and the charging management module 740 may also be disposed in the same device.

[0352] The wireless communication function of the electronic device 700 may be implemented by the antenna 1, the antenna 2, the mobile communication module 750, the wireless communication module 760, the modulation and demodulation processor, and the baseband processor, etc.

[0353] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 700 can be used to cover single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0354] The mobile communication module 750 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 700. The mobile communication module 750 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 750 can receive electromagnetic waves through the antenna 1, filter and amplify the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 750 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation.

[0355] The wireless communication module 760 can provide solutions for wireless communications applied to the electronic device 700, including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 760 may be one or more devices integrating at least one communication processing module. The wireless communication module 760 receives electromagnetic waves through the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 710. The wireless communication module 760 can also receive the signal to be transmitted from the processor 710, perform frequency modulation and amplification on it, and convert it into electromagnetic waves through the antenna 2 for radiation.

[0356] In some embodiments, antenna 1 of the electronic device 700 is coupled to the mobile communication module 750, and antenna 2 is coupled to the wireless communication module 760, enabling the electronic device 700 to communicate with the network and other devices through wireless communication technologies.

[0357] The electronic device 100 can implement audio functions through the audio module 770, speaker 770A, receiver 770B, microphone 770C, headphone jack 770D, and the application processor, etc. For example, music playback, recording, etc.

[0358] The audio module 770 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into a digital audio signal. The audio module 770 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 770 can be disposed in the processor 710, or some functional modules of the audio module 770 can be disposed in the processor 710. The speaker 770A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The receiver 770B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. The microphone 770C, also known as the "microphone", "transmitter", is used to convert a sound signal into an electrical signal. The headphone jack 770D is used to connect a wired headphone. The headphone jack 770D can be a USB interface 730, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface. In the embodiments of the present application, after the voice control function is turned on, the electronic device 100 continuously collects the surrounding sounds through the microphone 770C.

[0359] The pressure sensor 780A is used to sense a pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 780A can be disposed on the display screen 793. There are many types of pressure sensors 780A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 780A, the capacitance between the electrodes changes. The electronic device 700 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 793, the electronic device 700 detects the intensity of the touch operation according to the pressure sensor 780A. The electronic device 700 can also calculate the position of the touch according to the detection signal of the pressure sensor 780A.

[0360] The touch sensor 780B, also known as the "touch panel". The touch sensor 780B can be disposed on the display screen 793. The touch sensor 780B and the display screen 793 form a touch screen, also known as the "touch display screen". The touch sensor 780B is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 793. In some other embodiments, the touch sensor 780B can also be disposed on the surface of the electronic device 700, at a different position from that of the display screen 793.

[0361] The keys 790 include a power-on key, volume keys, etc. The keys 790 can be mechanical keys. They can also be touch keys. The electronic device 700 can receive key inputs and generate key signal inputs related to the user settings and function controls of the electronic device 700.

[0362] The motor 791 can generate vibration prompts. The motor 791 can be used for incoming call vibration prompts and can also be used for touch vibration feedback.

[0363] The camera 792 is used to capture still images or videos. In some embodiments, the electronic device 700 can include one or N cameras 792, where N is a positive integer greater than 1.

[0364] The electronic device 700 realizes the display function through the GPU, the display screen 793, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 793 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 710 can include one or more GPUs, which execute program instructions to generate or change display information.

[0365] The display screen 793 is used to display images, videos, etc. In some embodiments, the electronic device 700 can include one or N display screens 793, where N is a positive integer greater than 1.

[0366] The SIM card interface 794 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 794 to achieve contact and separation with the electronic device 700. The electronic device 700 can support one or N SIM card interfaces, where N is a positive integer greater than 1.

[0367] The voice control methods introduced in the above embodiments can all be executed in an electronic device having the above hardware structure.

[0368] Some other embodiments of the present application provide an electronic device (such as mobile phone 100). The electronic device may include: a microphone, a display screen, a processor, and a memory. The microphone, the display screen, and the memory are respectively coupled to the processor. The microphone is used to collect voices; the display screen is used to display the interface of the electronic device; the memory is coupled to the processor. The memory is further used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device can perform each function or step that the mobile phone 100 performs in the above method embodiments. The structure of the electronic device may refer to Figure 18 the structure of the electronic device 700 shown.

[0369] Embodiments of the present application further provide a chip system, as Figure 19 shown. The chip system 1000 includes at least one processor 1001 and at least one interface circuit 1002. The processor 1001 and the interface circuit 1002 can be interconnected through a line. For example, the interface circuit 1002 can be used to receive signals from other devices (such as the memory of a computer). For another example, the interface circuit 1002 can be used to send signals to other devices (such as the processor 1001). Exemplarily, the interface circuit 1002 can read the instructions stored in the memory and send the instructions to the processor 1001. When the instructions are executed by the processor 1001, the computer can perform each step in the above embodiments. Of course, the chip system can also include other discrete devices, and the embodiments of the present application do not make specific limitations thereto.

[0370] Embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium includes computer instructions. When the computer instructions run on the above electronic device (such as mobile phone 100), the electronic device is caused to perform each function or step that the mobile phone 100 performs in the above method embodiments.

[0371] Embodiments of the present application further provide a computer program product. When the computer program product runs on a computer, the computer is caused to perform each function or step that the computer performs in the above method embodiments. Wherein, the computer may be an electronic device, such as mobile phone 100.

[0372] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0373] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0374] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0375] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0376] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes.

[0377] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A voice control method, characterized in that, Applied to an electronic device, the electronic device includes a microphone, and the method includes: Displaying a first interface, the first interface including a first switch; In response to a user's operation of turning on the first switch, enabling a voice control function; the enabling of the voice control function includes: enabling a recording channel corresponding to the voice control function, the recording channel being used to acquire the voice collected by the microphone; In response to a voice click instruction collected by the microphone, searching for a first control that matches the first voice instruction on the current display interface of the electronic device; In the case where the click attribute of the first control is non-clickable, searching for a second control that belongs to the same control group as the first control on the current display interface; If the click attribute of the second control is clickable, based on the area corresponding to the second control, execute the click event corresponding to the voice click instruction.

2. The method according to claim 1, characterized in that, The executing the click event corresponding to the voice click instruction based on the area corresponding to the second control includes: Determining a simulated click area based on the area corresponding to the second control; Determining a simulated click position within the simulated click area; At the simulated click position, execute a simulated click operation.

3. The method according to claim 2, wherein The simulated click area is the area corresponding to the second control.

4. The method according to claim 2, wherein Before determining the simulated click area based on the area corresponding to the second control, the method further includes: acquiring the area corresponding to the first control; The determining the simulated click area based on the area corresponding to the second control includes: Splicing the area corresponding to the second control with the area corresponding to the first control to obtain a spliced area, and the simulated click area includes the spliced area.

5. The method according to any one of claims 1-3, characterized in that, The voice click instruction matches at least one interface hot word set in the electronic device, and the first control is a text control.

6. The method according to any one of claims 1-5, characterized in that, After searching for a second control that belongs to the same control group as the first control on the current display interface, the method further includes: Acquiring the number of the second controls; The if the click attribute of the second control is clickable, based on the area corresponding to the second control, execute the click event corresponding to the voice click instruction includes: In the case where the number of the second controls is 1, if the click attribute of the second control is clickable, based on the area corresponding to the second control, execute the click event corresponding to the voice click instruction.

7. The method according to claim 6, characterized in that, The method further includes: In any of the following cases, sending a prompt message: No second control that belongs to the same control group as the first control is found; or, A second control that belongs to the same control group as the first control is found, and the click attribute of the second control is non-clickable; or, A second control that belongs to the same control group as the first control is found, and the number of the second controls is greater than 1; wherein, the prompt message is used to indicate non-clickability.

8. The method according to any one of claims 1 to 7, wherein The electronic device includes a voice control application package APK, a voice processing engine, and an activity manager service AMS; the method further includes: The voice control APK receives the top-level activity of the electronic device returned by the AMS, parses the top-level activity, obtains and saves the controls included in the current display interface of the electronic device, and the controls include text controls; The voice control APK sends the text included in the text control of the current display interface to the voice processing engine; The voice control APK distributes the voice click instruction to the voice processing engine in response to receiving the voice click instruction; The voice processing engine searches for the target text that matches the voice click instruction; The voice processing engine maps the voice click instruction to a control click instruction and returns the control click instruction to the voice control APK, and the control click instruction carries the target text; The voice control APK searches for the first control corresponding to the target text from the text controls of the current display interface and obtains the click attribute of the first control.

9. The method according to claim 8, wherein The method further includes: When the click attribute of the first control is non-clickable, the voice control APK searches for the second control belonging to the same control group as the first control and obtains the click attribute of the second control.

10. An electronic device, characterized in that, The electronic device includes: a microphone, a display screen, a memory, and a processor; the microphone, the display screen, and the memory are respectively coupled to the processor; The microphone is used to collect language, the display screen is used to display the interface of the electronic device; the memory is used to store computer instructions; when the computer instructions are executed by the processor, the electronic device executes the method according to any one of claims 1-9.

11. A computer-readable storage medium, characterized in that It includes computer instructions, and when the computer instructions are executed by the processor of the electronic device, the electronic device executes the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Voice control method, device thereof, ventilator and storage medium

    CN109166584A

  • Voice control method and electronic equipment

    CN113794800A

  • Control method and control device for user interface

    CN115729418A

  • Voice control method and device, electronic equipment and computer readable storage medium

    CN115798469A

  • Click control method and system based on image recognition and voice recognition

    CN116088992A