Voice control method and electronic device

By enabling voice control, the electronic device determines whether the control to be clicked is obscured and determines the simulated click area in the area corresponding to the floating window and the area of ​​the control to be clicked. This solves the problem of inaccurate execution of click intent caused by the obscuration of the control and realizes accurate clicking in obscured scenarios.

CN120279903BActive Publication Date: 2026-05-29HONOR DEVICE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2023-12-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

When multiple windows are displayed on an electronic device, the controls under voice control are obscured, making it impossible to accurately execute the user's click intentions.

Method used

After the voice control function is enabled, the electronic device determines whether the control to be clicked is obscured, and determines the simulated click area in the area corresponding to the floating window and the area of ​​the control to be clicked, and executes the click event corresponding to the voice click command to avoid simulated click errors.

Benefits of technology

Even when controls are obscured, ensure that electronic devices can accurately execute the user's click intent, reducing the possibility of simulated click errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279903B_ABST
    Figure CN120279903B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a voice control method and an electronic device, and relate to the technical field of voice processing, which is used to accurately execute the click intention of a user in a scenario where a control matched with the voice is blocked. The method is applied to an electronic device including a microphone, and includes: displaying a first interface including a first switch. In response to an operation of opening the first switch, a voice control function is started. Specifically, a voice control function corresponding recording channel is started, and the recording channel is used to acquire a voice collected by the microphone. In response to a voice click instruction collected by the microphone, a to-be-clicked control matched with the voice click instruction is searched in a current display interface of the electronic device. In a case where the to-be-clicked control is partially blocked by a floating window, a first simulated click area is determined based on an area corresponding to the floating window and an area corresponding to the to-be-clicked control. In the first simulated click area, a click event corresponding to the voice click instruction is executed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice processing technology, and in particular to a voice control method and electronic device. Background Technology

[0002] With technological advancements, voice control has gradually become a common human-computer interaction method. Especially in scenarios where users' hands are occupied, such as driving, cooking, or reading, controlling electronic devices via voice is extremely convenient and quick.

[0003] In the process of an electronic device responding to a user's voice input to perform a click operation, related technologies first require locating the control to be clicked that matches the voice. Then, the electronic device simulates a click operation at the center of the text control to be clicked. However, in some scenarios, the electronic device displays multiple windows simultaneously, and the control to be clicked may be obscured. In this case, the electronic device may not be able to accurately click the control to be clicked in response to the voice, thus failing to realize the user's intention. Therefore, there is an urgent need for a method that can accurately execute the user's click intention even when the control matching the voice is obscured. Summary of the Invention

[0004] This application provides a voice control method and electronic device for accurately executing a user's click intent in a scenario where the control matching the voice is obscured.

[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0006] Firstly, a voice control method is provided, which is applied to an electronic device, including a microphone.

[0007] The method includes:

[0008] The electronic device displays a first interface, which includes a first switch. In response to the user turning on the first switch, a voice control function is activated. This voice control function is activated via the first switch and does not require waking the electronic device. For example, this voice control function is a "see-and-speak" function. After the voice control function is activated, it receives a voice click command from the microphone and searches for a clickable control matching the first voice command on the current display screen. Then, the electronic device determines whether the clickable control is obscured. If the clickable control is partially obscured by a floating window, the electronic device first determines a first simulated click area based on the area corresponding to the floating window and the area corresponding to the clickable control. Then, within the first simulated click area, it executes the click event corresponding to the voice click command. This avoids the problem of the electronic device accidentally clicking the clickable control when executing the click event corresponding to the voice click command, thus preventing simulation click errors in such scenarios. Therefore, the electronic device, responding to the user's voice input containing the intention to click a control, can execute a simulated click operation at the accurate location, thereby accurately realizing the user's intention.

[0009] Enabling the voice control function includes: activating the corresponding recording channel for the voice control function. This recording channel is used to acquire the voice captured by the microphone. Once the recording channel for the voice control function is activated, the voice control function can acquire the voice (audio stream) captured by the microphone through this recording channel.

[0010] In one possible implementation of the first aspect, the control to be clicked is a control in the focus window of the electronic device.

[0011] In one possible implementation of the first aspect, the clickable control is partially obscured by the floating window. Specifically, this may include: the current display interface of the electronic device includes a floating window, the floating window is not the focus window, and the area corresponding to the floating window and the area corresponding to the clickable control overlap. The overlapping area is smaller than the area corresponding to the clickable control, meaning the clickable control is not completely obscured.

[0012] When the floating window is the focused window, and the control to be clicked is a control on the floating window, the clickable control will not be obscured by the floating window. Therefore, if the current display interface includes a floating window, and the floating window is not the focused window, the clickable control may be obscured. In this case, by comparing the areas corresponding to the floating window and the clickable control, it can be determined whether there is an overlapping area, thus confirming whether the clickable control is obscured by the floating window. In this solution, the electronic device can quickly determine whether the clickable control is obscured.

[0013] In one possible implementation of the first aspect, determining the first simulated click area based on the area corresponding to the floating window and the area corresponding to the control to be clicked can specifically include: subtracting the overlapping area from the area corresponding to the control to be clicked to obtain the first simulated click area. The first simulated click area is the portion of the area corresponding to the control to be clicked that is not obscured by the floating window. Executing a click event within this first simulated click area ensures that the electronic device will not click on other windows, thereby accurately executing voice click commands.

[0014] In one possible implementation of the first aspect, the method may further include: when it is detected that the current display interface of the electronic device includes a floating window, obtaining the area corresponding to the currently focused window of the electronic device. Then, comparing the area corresponding to the focused window with the area corresponding to the floating window. If the area corresponding to the focused window and the area corresponding to the floating window are completely identical, then the floating window is determined to be the focused window. In this way, it is possible to quickly determine whether the floating window is the focused window, which is convenient for determining whether the clickable control is obscured by the floating window.

[0015] In one possible implementation of the first aspect, the method may further include: executing a click event corresponding to a voice click command within a second simulated click area, provided the control to be clicked is not obscured. The second simulated click area includes the area corresponding to the control to be clicked. If the control to be clicked is not obscured, the click event can be executed directly within the area corresponding to the control.

[0016] Specifically, the requirement that the clickable control is not obscured can include: the current display interface of the electronic device does not include floating windows.

[0017] Alternatively, the control to be clicked is not obscured, which may specifically include: the current display interface of the electronic device includes a floating window, and the floating window is the focus window.

[0018] Alternatively, the fact that the control to be clicked is not obscured can specifically include: the current display interface of the electronic device includes a floating window, the floating window is not the focus window, and the area corresponding to the control to be clicked does not overlap with the area corresponding to the floating window.

[0019] In one possible implementation of the first aspect, the method may further include: issuing a prompt message when the control to be clicked is completely obscured by the floating window. The prompt message indicates that the control is not clickable. This serves to inform the user of the result of the voice command.

[0020] In one possible implementation of the first aspect, searching for a clickable control matching the voice click command on the current display screen of the electronic device includes: searching for a first control matching the voice click command on the current display screen of the electronic device; obtaining the click attribute of the first control; and determining the first control as the clickable control if the click attribute of the first control is clickable. Before performing an occlusion judgment on the clickable control, the electronic device also judges the click attribute of the clickable control. This further ensures that the clickable control can execute a click event. This ensures that, while accurately recognizing the voice intent, the electronic device reduces the possibility of failing to realize the true intent of the user's voice input, and increases the likelihood of responding to the user's voice input and executing an accurate simulated click operation.

[0021] In one possible implementation of the first aspect, the voice click command is matched with an interface hotword set by at least one electronic device; in this implementation, the first control is a text control.

[0022] In one possible implementation of the first aspect, searching for a clickable control matching the voice click command on the current display interface of the electronic device may further include: if the click attribute of the first control is not clickable, then searching for a second control belonging to the same control group as the first control. The click attribute of the second control is obtained. If the click attribute of the second control is clickable, then the second control is determined as the clickable control. This further ensures that the clickable control is capable of executing a click event. This guarantees that, while accurately recognizing the user's voice intent, the possibility of the electronic device failing to realize the true intent of the user's voice input is reduced, increasing the likelihood of responding to the user's voice input and executing accurate simulated click operations.

[0023] In one possible implementation of the first aspect, searching for a clickable control matching the voice click command on the current display interface of the electronic device may further include: if the click attribute of the first control is clickable, and the first control is completely obscured by a floating window, searching for a second control belonging to the same control group as the first control. Obtaining the click attribute of the second control. If the click attribute of the second control is clickable, then determining the second control as the clickable control.

[0024] In one possible implementation of the first aspect, the electronic device includes a voice control application package (APK), a voice processing engine, and a window activity manager (AMS). The method may further include: the voice control APK receiving the current display interface of the electronic device returned by the AMS, parsing the current display interface, and obtaining and saving the text controls and image controls of the current display interface. The voice control APK sends the text contained in the text controls of the current display interface as interface hot words to the voice processing engine. In response to receiving a voice click command, the voice control APK distributes the voice click command to the voice processing engine. The voice processing engine searches for target interface hot words that match the voice click command. The voice processing engine returns the target interface hot words to the voice control APK. The voice control APK queries whether the electronic device includes a floating window, and if the electronic device includes a floating window, determines whether the floating window is the focused window, and determines whether the area corresponding to the floating window and the area corresponding to the clickable control overlap.

[0025] In another possible implementation of the first aspect, within the first simulated click area, executing the click event corresponding to the voice click command includes: determining the simulated click position within the first simulated click area; and executing the simulated click operation at the simulated click position.

[0026] Secondly, this application also provides an electronic device. The electronic device may include a microphone, a display screen, a processor, and a memory. The microphone, display screen, and memory are each coupled to the processor. The microphone is used to capture speech, and the display screen is used to display the interface of the electronic device. The memory is used to store computer execution instructions. When the electronic device is running, the processor executes the computer execution instructions stored in the memory to cause the electronic device to perform the voice control method as described in any of the first aspects above.

[0027] Thirdly, this application provides a computer-readable storage medium including computer instructions that, when executed by a processor of an electronic device, enable the electronic device to perform the voice control method of any of the first aspects described above.

[0028] Fourthly, a computer program product containing instructions is provided, which, when run on an electronic device, enables the electronic device to execute any of the voice control methods described in the first aspect above.

[0029] Fifthly, an apparatus (e.g., a system-on-a-chip) is provided, comprising a processor for supporting an electronic device in performing the functions described in the first aspect above. In one possible design, the apparatus further comprises a memory for storing program instructions and data necessary for the electronic device. When the apparatus is a system-on-a-chip, it may be composed of chips or may include chips and other discrete devices.

[0030] The technical effects of any of the design methods in aspects two through five can be found in the technical effects of different design methods in aspect one, and will not be repeated here. Attached Figure Description

[0031] Figure 1 This is a schematic diagram illustrating a scenario example of a voice control method.

[0032] Figure 2A A schematic diagram illustrating the activation process of a voice control function;

[0033] Figure 2B This is a schematic diagram illustrating a scenario example of a voice control method.

[0034] Figure 2C This is a schematic diagram illustrating a scenario example of a voice control method.

[0035] Figure 3 A flowchart illustrating a voice control method provided in an embodiment of this application;

[0036] Figure 4 A schematic diagram illustrating a scenario of the voice control method provided in this application embodiment;

[0037] Figure 5A This is a schematic diagram of the display interface of an electronic device;

[0038] Figure 5B A schematic diagram illustrating the activation process of a voice control function;

[0039] Figure 6 A flowchart illustrating a voice control method provided in an embodiment of this application;

[0040] Figure 7A A flowchart illustrating a voice control method provided in an embodiment of this application;

[0041] Figure 7B A schematic diagram illustrating a scenario of the voice control method provided in this application embodiment;

[0042] Figure 7C A schematic diagram illustrating a scenario of the voice control method provided in this application embodiment;

[0043] Figure 8A schematic diagram illustrating the determination of the simulated click area provided in this application embodiment;

[0044] Figure 9A A flowchart illustrating a voice control method provided in an embodiment of this application;

[0045] Figure 9B A schematic diagram illustrating a scenario example of a voice control method provided in this application embodiment;

[0046] Figure 10 A schematic diagram of the software architecture of an electronic device provided in an embodiment of this application;

[0047] Figure 11 A schematic diagram illustrating the interaction of various modules of an electronic device in order to implement the voice control method provided in the embodiments of this application;

[0048] Figure 12 A schematic diagram illustrating the interaction of various modules of an electronic device in order to implement the voice control method provided in the embodiments of this application;

[0049] Figure 13 A flowchart illustrating a voice control method provided in an embodiment of this application;

[0050] Figure 14 A schematic diagram of the software architecture of an electronic device provided in an embodiment of this application;

[0051] Figure 15 A schematic diagram illustrating the interaction of various modules of an electronic device in order to implement the voice control method provided in the embodiments of this application;

[0052] Figure 16 A schematic diagram illustrating the interaction of various modules of an electronic device in order to implement the voice control method provided in the embodiments of this application;

[0053] Figure 17 A flowchart illustrating a voice control method provided in an embodiment of this application;

[0054] Figure 18 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0055] Figure 19 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation

[0056] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:

[0057] The goal of automatic speech recognition (ASR) is to convert the lexical content of a user's speech into computer-readable input, such as keystrokes, binary codes, or character sequences.

[0058] Natural language understanding (NLU) is a general term for all methods, models, or tasks that support machines in understanding text content.

[0059] Dialogue management (DM) is used to control the process of human-computer dialogue and to determine the current response to the user based on dialogue history information.

[0060] An activity is a feature of the Android system. One of the four main components is the user interface; it provides users with a window to complete operation commands.

[0061] An intent is a thought or idea that aims to achieve a certain goal. In the field of voice control, intent recognition is a crucial technology. Accurately identifying and understanding a user's needs and intent allows for more precise responses to their voice commands, thus fulfilling their requirements. For example, if a user inputs "Check today's weather," the electronic device can perform intent recognition, extracting the entities "check" and "weather" to determine that the user's intent is to check the weather. Similarly, if a user inputs "Go back to the previous page," the electronic device can perform intent recognition, extracting the entities "back" and "previous page" to determine that the user's intent is to display the previous page. Furthermore, if a user inputs "Open calendar," the electronic device can extract the entities "open" and "calendar" to determine that the user's intent is to open the calendar application.

[0062] Voice control function:

[0063] Many electronic devices support voice control. These devices use microphones to capture user-inputted voice, analyze and recognize the voice, and execute the corresponding commands, allowing users to control the device via voice.

[0064] Generally, to conserve power and prevent accidental triggering of electronic devices, voice control functions need to be enabled before use. For example, a user enters a preset word (called a wake-up word) into the electronic device to wake it up. Once awakened, the electronic device can execute the corresponding voice command, thus activating the voice control function. For instance, a user can enable or disable the voice control function by toggling preset switches on or off through the device's user interface.

[0065] Voice control functions may have different names in different electronic devices, such as "voice control," "intelligent voice," "voice assistant," "see and speak," "voice command," "command at will," and "intelligent AI." The specific implementation of voice control functions with different names may also differ.

[0066] The following examples illustrate several different implementations of voice control functionality.

[0067] Voice assistant:

[0068] Before a user can control an electronic device using a voice assistant, the device needs to be woken up. In one example, the device is woken up when a wake word is detected by the user's voice input. In another example, the device is woken up when the user presses and holds the power button. In yet another example, the device is woken up when the user's breath is detected while inputting voice. Typically, before the device is woken up, its microphone operates in a power-saving mode (e.g., searching for signals at low power) to pick up ambient sounds. The microphone only detects the wake word at the kernel level; the corresponding recording channel for the voice assistant is not activated in the device's system or drivers.

[0069] When an electronic device is activated in response to a user action (such as receiving a wake word via voice input), the system and drivers initiate the corresponding recording channel for the voice assistant. After the electronic device is activated, the voice (audio stream) captured by the microphone is sent to the voice assistant application for processing through the corresponding recording channel. In this way, the electronic device can execute voice commands, enabling the user to control the electronic device via voice; it can also enable functions such as dialogue with the user.

[0070] Taking a mobile phone 100 as an example, for instance, Figure 1 This illustration shows a scenario where a user controls a mobile phone 100 using a voice assistant. Figure 1As shown, phone 100 displays a desktop interface, and the user inputs the voice command "Hello YOYO" into phone 100. In response to receiving the wake-up word "Hello YOYO," phone 100 is activated. For example, after phone 100 is activated, it plays the voice command "I'm here" to notify the user that phone 100 has been activated. After phone 100 is activated, the user can control phone 100 via voice. For example, as... Figure 1 As shown, the user inputs the voice command "Open video" into the phone. The phone 100 parses and recognizes the user's voice input and executes the command corresponding to "Open video". For example, in response to receiving the voice command "Open video", the phone 100 launches the video application.

[0071] In other examples, electronic devices may also wake up a voice assistant in response to a user pressing and holding the power button.

[0072] In some implementations, after the voice assistant on an electronic device is activated, the user can issue a command to the device by inputting voice, and the device will execute that command. Once the device has executed a command, or if the voice assistant does not receive a command from the user within a certain period (e.g., 8 seconds) after being activated, the device will no longer respond to voice commands. For example, the device may close the recording channel corresponding to the voice assistant. The user needs to input a wake word again to activate the device before they can issue commands by inputting voice again. In other words, after the voice assistant is activated, it enters a "short-reception" state, responding to user commands via voice for a short period (e.g., 8 seconds).

[0073] In some implementations, when the electronic device is connected to the network, it supports entering a continuous dialogue scenario after being woken up, allowing the user to engage in a sustained conversation with the device. After each broadcast, the electronic device will continue to pick up audio without needing to be woken up again, until the user exits the continuous dialogue using a command such as "exit."

[0074] What is visible can be said:

[0075] It can be seen that this is achieved locally by electronic devices and does not require a network connection.

[0076] In some examples, the "See and Say" feature is controlled by a preset switch. Users can enable or disable the "See and Say" feature by turning the preset switch on or off.

[0077] For example, such as Figure 2AAs shown, a user can access the settings function of mobile phone 100; for example, the user clicks the application icon of the "Settings" app on the desktop. In response to the user's click on the "Settings" app icon, mobile phone 100 displays the "Settings" interface 101. The "Settings" interface 101 includes a "Smart Voice" option 102, which is used to configure the smart voice function. For example, in response to the user's click on the "Smart Voice" option 102, mobile phone 100 displays the "Smart Voice" interface 103, which includes a "See and Speak" option 104. The user can click on the "See and Speak" option 104 to configure options related to the "See and Speak" function. For example, refer to... Figure 2A In response to a user's click on the "Speak When You See It" option 104, the mobile phone 100 displays the "Speak When You See It" interface 105. Optionally, the "Speak When You See It" interface 105 includes a prompt message 106 to instruct the user on how to use the "Speak When You See It" function. The "Speak When You See It" interface 105 also includes a "Speak When You See It" switch 107 (i.e., the aforementioned preset switch). The user can click the "Speak When You See It" switch 107 to turn the "Speak When You See It" switch on or off. In one example, in response to receiving a user's click on the "Speak When You See It" switch 107, the mobile phone 100's "Speak When You See It" switch turns on, enabling the "Speak When You See It" function. Optionally, the "Speak When You See It" interface 105 displays a prompt message 108 to instruct the user that the "Speak When You See It" function has been successfully enabled.

[0078] In one implementation, after the "See and Say" function is enabled, the mobile phone 100 displays a first recording icon, which indicates that the "See and Say" function is enabled. For example, as shown... Figure 2A As shown, after the "See and Talk" switch 107 is turned on, the status bar of the mobile phone 100 displays the recording icon 10a, indicating that the "See and Talk" function has been enabled.

[0079] In one scenario, the preset switch for "See and Say" on the electronic device is not turned on, and the device's microphone is not activated. When the preset switch for "See and Say" is turned on, the electronic device activates the microphone and starts the corresponding recording channel in the system and driver. In this way, the voice (audio stream) captured by the microphone can be sent to the "See and Say" application for processing through the corresponding recording channel, enabling the user to control the electronic device via voice.

[0080] In another scenario, the preset switch for "Visible & Talkable" on the electronic device is not turned on. The device's microphone operates in a power-saving mode (e.g., searching for signals at low power) to pick up ambient sound. When the preset switch for "Visible & Talkable" is turned on, the corresponding recording channel for "Visible & Talkable" is activated in the device's system and drivers. In this way, the voice (audio stream) captured by the microphone can be sent to the "Visible & Talkable" application for processing through the corresponding recording channel, enabling the user to control the electronic device via voice.

[0081] Once "Visible and Talkable" is enabled, the corresponding recording channel is activated in the electronic device's system and drivers. The electronic device enters a "long-continuous recording" state, continuously capturing ambient sound. Users can issue commands to the electronic device at any time via voice, without needing to input a wake-up word.

[0082] In one implementation, after any function on the electronic device activates the voice recording function (turns on the microphone and opens the recording channel), the electronic device will send a prompt message to the user to indicate that the electronic device is in voice recording mode. This can prevent the user's privacy from being leaked. For example, after the "see and speak" function is activated, the electronic device enters a continuous voice recording state, and a second recording icon is displayed on the electronic device's display interface. This second recording icon indicates that the recording channel is open, serving as a notification to the user that the microphone is recording voice. For example, as shown... Figure 2A As shown, the status bar of the mobile phone 100 displays a recording icon 10b, indicating that the recording channel is enabled.

[0083] Once the "See It, Speak It" function is enabled, the electronic device activates its recording channel and continuously captures voice input via the microphone. Users can input voice messages into the electronic device at any time. The device then parses and recognizes the user's voice input and executes the corresponding commands.

[0084] As you can see, the voice input function supports user commands that can be input via voice, including: system commands such as swipe left, swipe right, swipe up, return to desktop, return, increase volume, decrease volume; video application commands such as play, pause, stop, fast forward, rewind; and e-book playback application commands such as previous page, next page, table of contents, next chapter.

[0085] The "see-and-say" functionality of electronic devices allows users to input commands via voice, which can be categorized into several types, one of which is the operation category. For operation-type commands, when the electronic device responds to voice and executes the corresponding command, it simulates the user's action.

[0086] Hot words:

[0087] Electronic devices can be set with hot words. Once enabled, if the received voice message matches at least one of the set hot words, the electronic device will execute the corresponding voice command.

[0088] Hot keywords can be pre-configured on electronic devices, obtained from content displayed on the electronic device's interface, or input by the user. Depending on the scope and source of their use, hot keywords can be categorized into system-level hot keywords, scenario-level hot keywords, and interface-level hot keywords.

[0089] System-level hotwords are pre-configured and applicable to any application on the electronic device. When any application on the electronic device is running in the foreground, if the user's voice input matches at least one system-level hotword, the electronic device executes the command corresponding to the user's voice. System-level hotwords are global and do not depend on the application or application interface. For example, system-level hotwords may include: swipe left, swipe right, swipe up, return to desktop, back, etc.

[0090] Scene-level keywords apply to all applications within a given scene. In one implementation, multiple scenes are pre-defined in the electronic device, each corresponding to at least one application. For example, pre-defined scenes may include audio / video scenes, e-book scenes, home screen scenes, and permission pop-up scenes. Different scenes can have different pre-defined scene-level keywords. For example, scene-level keywords for the audio / video scene may include: play, pause, stop, fast forward, and rewind. Scene-level keywords for the e-book scene may include: previous page, next page, table of contents, and next chapter. Scene-level keywords for the home screen scene may include: application list information and settings menu items. Scene-level keywords for the permission pop-up scene may include: allow, confirm, deny, and understand.

[0091] Interface hot keywords are applicable to a specific interface. In some implementations, the electronic device obtains hot keywords from the display interface of the foreground application, i.e., it obtains interface hot keywords. For example, refer to... Figure 2BThe mobile phone 100 displays the "My" interface 110 of the video application. The text information in the "My" interface 110 includes "5G", "8:00", "Login / Register", "My Downloads", "Following Favorites", "My Purchases", "My Scenes", "History", "Coupon Package", "Settings", "Feedback", "Customer Service", "Homepage", "Member", "Short Videos", and "My". The mobile phone 100 sets the text in the currently displayed interface of the foreground application as interface hot keywords, namely, the interface hot keywords include "5G", "8:00", "Login / Register", "My Downloads", "Following Favorites", "My Purchases", "My Scenes", "History", "Coupon Package", "Settings", "Feedback", "Customer Service", "Homepage", "Member", "Short Videos", and "My". When the electronic device detects a change in the displayed interface, it can retrieve the text content of the currently displayed interface after the interface change and update the interface hot keywords.

[0092] In some scenarios, electronic devices can display multiple windows simultaneously. For example, in a scenario where an electronic device displays floating windows, two windows may be displayed at the same time. In this scenario, the interface hotspots acquired and saved by the electronic device are specifically the hotspots displayed in the focused window. The focused window represents the selected window among multiple windows, i.e., the window currently being operated on. In this scenario, global operations performed by the user apply to this focused window.

[0093] For example, the electronic device simultaneously displays a settings interface and a calculator interface in a floating window, with the settings interface being the focused window. At this time, the hot words retrieved and saved by the electronic device are those retrieved and saved from the settings interface.

[0094] Different hot words can correspond to different instructions. Since the interface hot words are set based on the displayed text of the current screen, in some embodiments, the instruction corresponding to the interface hot words is a control click instruction. Specifically, when the electronic device determines that the received voice matches at least one interface hot word, it can execute a control click instruction at the position of the text corresponding to that interface hot word on the screen. In some examples, when the electronic device executes a control click instruction, it can specifically execute a click event corresponding to the control click instruction, that is, perform a simulated click operation at the position of the corresponding text.

[0095] exist Figure 2BIn the example shown, mobile phone 100 receives the user's voice input "Open History". The voice "Open History" matches the interface's hotkey "History", and mobile phone 100 executes the command corresponding to the voice "Open History". Specifically, mobile phone 100 can perform a simulated click operation within the area corresponding to the "History" control, thereby fulfilling the user's intention to open the history. For example... Figure 2B As shown, after the mobile phone 100 performs a simulated click operation, the history interface 113 can be displayed.

[0096] When mobile phone 100 executes the voice command "Open History," it needs to determine the area corresponding to the "History" control. Then, based on the area of ​​the "History" control, it determines a simulated click location. In one example, the electronic device can determine the center position of the area corresponding to the "History" control as the simulated click location. Figure 2B As shown, the center position of the area corresponding to the "History" control is position 112, which can be identified as the simulated click position.

[0097] When the "My" screen of the video application is displayed on the phone 100, the user can also input other voice commands, such as "coupon package". The phone 100 receives the user's voice input "coupon package", determines that the voice command "coupon package" matches the interface's hot keyword "coupon package", and then executes the command corresponding to the voice command "coupon package". For example... Figure 2C As shown, mobile phone 100 performs a simulated click operation in the area corresponding to the "Coupon Package" control 114, which corresponds to the hot word "Coupon Package" on the interface.

[0098] However, in some embodiments, the "Coupon Package" control 114 is set to non-clickable. Therefore, after the mobile phone 100 responds to the voice command "Coupon Package" and performs a simulated click on the "Coupon Package" control 114, the mobile phone 100 cannot open the coupon package interface. Thus, the mobile phone 100 cannot fulfill the user's intent.

[0099] Based on this, this application proposes a voice control method that can be applied to electronic devices that support voice input. After the electronic device activates the voice control function in response to a user's operation, it can receive voice and execute the corresponding voice command. When the electronic device receives a first voice containing a control click intent on a second interface, it searches for a control (denoted as the first control) that matches the first voice on the second interface. If the first control is not clickable, the electronic device searches for a second control belonging to the same control group as the first control on the second interface. If the second control is clickable, a simulated click area is determined based on the second control, and a simulated click operation is performed within the simulated click area, thereby realizing the intent of the first voice. In this way, if the voice intent is accurately recognized, the possibility that the electronic device cannot realize the true intent of the user's voice input is reduced, and the probability of accurately executing a simulated click operation in response to the user's voice input is increased.

[0100] For example, the aforementioned electronic devices may be mobile phones, tablets, laptops, personal computers (PCs), ultra-mobile personal computers (UMPCs), handheld computers, netbooks, smart home devices (such as smart TVs, smart screens, large screens, smart speakers, smart air conditioners, etc.), personal digital assistants (PDAs), wearable devices (such as smartwatches, smart bracelets, etc.), in-vehicle devices, virtual reality devices, etc., and this application embodiment does not impose any limitations on them.

[0101] The specific implementation of the voice control method proposed in the embodiments of this application will be described in detail below with reference to the accompanying drawings. Figure 3 The flowcharts of voice control methods in some embodiments are shown.

[0102] S200. Display the first interface, which includes the first switch.

[0103] The first switch is used to turn the voice control function on or off. Users can turn the voice control function on or off via the first switch on the first interface. In some embodiments, the first interface may be... Figure 2A The "What you see is what you can say" interface 105 is shown; the first switch can be the "What you see is what you can say" switch 107.

[0104] S201. Receive the operation of turning on the first switch.

[0105] S202. In response to the operation of turning on the first switch, activate the voice control function.

[0106] In some embodiments, after the voice control function is enabled, the electronic device activates the recording channel corresponding to the voice control function. This recording channel is used to acquire the voice captured by the microphone. Furthermore, after the voice control function is enabled, the electronic device is in a continuous recording state, continuously capturing ambient sound. Users can issue commands to the electronic device at any time via voice, without needing to input a wake-up word.

[0107] In addition, activating voice control on an electronic device can further include displaying a first recording icon and a second recording icon. The first recording icon indicates that voice control is enabled, and the second recording icon indicates that the recording channel is open. These recording icons inform the user that voice control is enabled and that the electronic device is recording audio. This allows the user to quickly obtain information about the current status of the electronic device.

[0108] S203. Display the second interface.

[0109] The second interface can be any interface of the electronic device. After the voice control function is enabled, the electronic device can receive user-inputted voice and execute the corresponding command on any display interface. In some embodiments, the second interface can also be the first interface; that is, after the voice control function is enabled, the electronic device can receive user-inputted voice on the first interface and execute the corresponding command.

[0110] In some embodiments, after the electronic device enables voice control, it can parse the new display interface after each screen switch to obtain and save the control information of each control within that interface. This control information may include: control attributes (text control / image control), the area where the control is located (rectangular controls can be represented by the coordinates of their top-left and bottom-right corners), and whether the control is clickable. In this embodiment, after displaying a second interface, the electronic device can also parse that second interface to obtain the control information of each control within it. This facilitates accurate responses to received voice commands.

[0111] Since the electronic device has voice control enabled, the user can input voice commands into the device via voice control on the second interface. The electronic device then receives the user's voice input, as shown in S204.

[0112] S204. Receive the first voice message.

[0113] S205. Determine the user intent corresponding to the first voice.

[0114] In some embodiments, the first voice is matched with at least one hot word pre-set by the electronic device. This allows the electronic device to respond to the first voice and execute the corresponding instruction, which can be referred to as the first voice instruction.

[0115] In some embodiments, after S204, the electronic device can perform speech parsing on the first speech to obtain the parsed text corresponding to the first speech. Then, the electronic device can perform intent recognition on the obtained parsed text. In this embodiment, S205 may specifically include: parsing the first speech to obtain the parsed text corresponding to the first speech. Then, performing intent recognition on the parsed text corresponding to the first speech to obtain the user intent corresponding to the first speech.

[0116] As explained above, after receiving voice, the electronic device can match the voice with hot words. If the voice matches at least one hot word set by the electronic device, the electronic device can execute the command corresponding to the voice. Therefore, in some embodiments, the above-mentioned intention recognition of the parsed text corresponding to the first voice to obtain the user's intention corresponding to the first voice may specifically include: matching the parsed text corresponding to the first voice with hot words set by the electronic device. If the parsed text corresponding to the first voice matches at least one hot word of the electronic device, the user's intention can be determined according to the command corresponding to the hot word.

[0117] For example, if the parsed text corresponding to the first voice message matches a hot word on the interface, the electronic device can determine that the instruction corresponding to the first voice message is a control click instruction. Therefore, it can be determined that the user's intent for the first voice message is a control click intent. Similarly, if the hot word matching the parsed text of the first voice message belongs to the system-level hot words "swipe left," then the user's intent can be determined to be a swipe to the left. Or, if the hot word matching the parsed text of the first voice message belongs to the scene-level hot words "pause," then the user's intent can be determined to control the electronic device to pause playback.

[0118] As explained above, the interface hotwords set by the electronic device are obtained based on the currently focused window. In embodiments where the electronic device displays multiple windows simultaneously, the user's initial voice input may be content from a non-focused window. In this case, the electronic device will match the parsed text corresponding to the initial voice input with the interface hotwords, and will be unable to obtain an interface hotword that matches the parsed text corresponding to the initial voice input. In some embodiments, the electronic device will recognize the initial voice input as an invalid command in this situation.

[0119] In other embodiments, the electronic device may also pre-set multiple valid commands. When the user's input voice matches a valid command set by the electronic device, the electronic device can respond to the voice by executing the corresponding command. For example, valid commands may include: "back", "back to desktop", "play / pause", "previous page / next page", "swipe left / right", and "open [application name]", etc. Different valid commands may correspond to different intentions. In this embodiment, the above-mentioned S205 may specifically include: searching for a valid command that matches the first voice. Based on the valid command that matches the first voice, the user intention corresponding to the first voice is determined. Wherein, searching for a valid command that matches the first voice may first perform voice parsing on the first voice, and then compare the obtained parsed text with the valid commands stored in the electronic device one by one to determine whether the parsed text matches at least one valid command.

[0120] The above embodiments illustrate the scenario where the electronic device finds matching hot words or valid instructions based on the parsed text corresponding to the first voice input. In other embodiments, the user's first voice input may not find matching hot words or valid instructions. In this case, the electronic device cannot determine the user's intent and therefore cannot respond to the voice input by executing the corresponding instruction.

[0121] If the user intent corresponding to the first voice is determined to be a control click intent, the instruction corresponding to the first voice can be recorded as a voice click instruction.

[0122] In other embodiments, the electronic device may also determine the user intent corresponding to the first voice through other means.

[0123] S206. Determine if the user's intent is a control click intent.

[0124] After determining the user's intent, it can be determined whether the user's intent is a control click intent. In some embodiments, if the first voice matches an interface hotword set by at least one electronic device, the user's intent can be determined to be a control click intent.

[0125] If the judgment result of S206 is negative, it means that the user intent corresponding to the first voice is not a control click intent. In this case, the electronic device can directly execute the instruction corresponding to the first voice to realize the user intent. It should be noted that the case where the judgment result of S206 is negative is... Figure 3 Not shown in the image.

[0126] If the judgment result in S206 is yes, it means that the user intent corresponding to the first voice is a control click intent. Then, the electronic device can determine the control to be clicked corresponding to the first voice.

[0127] S207. On the second interface, find the first control that matches the first voice.

[0128] If the user's intent, determined from the first voice, is a control click intent, the electronic device needs to identify the control to be clicked. Then, it needs to obtain the area where the control to be clicked is located before executing the control click instruction. In some embodiments, the control matching the voice is a text control on the electronic device's current display screen.

[0129] In some embodiments, when determining the user intent corresponding to the first speech, the first speech is parsed to obtain parsed text, and then hot words matching the first speech are determined based on the parsed text. In this embodiment, S207 may specifically include: searching for a first control matching the first speech based on the parsed text corresponding to the first speech. Figure 2B In the example shown, the user inputs the voice command "Open History," and the electronic device determines that the voice command matches the interface keyword "History." The electronic device can then locate the corresponding "History" control based on this interface keyword. For example... Figure 2C In the example shown, the user inputs the voice message "coupon package," and the electronic device determines that the voice message matches the interface hotspot "coupon package." The electronic device can then locate the corresponding "coupon package" control based on this hotspot. In embodiments where the first voice message matches at least one interface hotspot, the first control matching the first voice message is a text control.

[0130] In other embodiments, the first control matching the first voice may also be an image control; such as Figure 2A The image control 109 corresponds to the return icon shown.

[0131] In an embodiment where an electronic device displays multiple windows simultaneously, the second interface is the interface displayed in the focus window.

[0132] When the electronic device executes the simulated click command corresponding to the first voice, in order to avoid the simulated click failing, it can first determine whether the area to be clicked is clickable, as in S208.

[0133] S208. Determine whether the first control is clickable.

[0134] It should be noted that S208 determines whether the first control itself has a clickable attribute.

[0135] As can be seen from the above embodiments, in some embodiments, after the electronic device displays the second interface, it can parse the displayed content of the second interface, obtain and save the control information on the second interface. In this embodiment, the above-mentioned S208 may specifically include: obtaining the control information of the first control, and determining whether the first control is clickable based on the control information of the target text control.

[0136] If the judgment result of S208 is yes, it means that the first control is clickable. At this time, a simulated click operation can be performed directly in the area corresponding to the first control. S213 and S214 can be executed.

[0137] If the result of S208 is negative, it means that the first control is not clickable, and S209 can be executed at this time.

[0138] S209. Determine if there exists a second control that belongs to the same control group as the first control.

[0139] In some embodiments, the control information obtained by the electronic device through interface parsing may further include: the control group to which the control belongs. In this embodiment, the electronic device can determine whether there is a sibling control belonging to the same control group as the first control by querying the control information; and the type and whether the sibling control is clickable.

[0140] The second control can be any type of control, such as a text control or an image control.

[0141] In this embodiment, S209 may specifically include: checking whether there is a sibling control belonging to the same control group as the first control. In one example, if the second interface includes a sibling control belonging to the same control group as the first control, then after S209, control information such as the type and whether the sibling control is clickable may also be obtained.

[0142] If the judgment result of S209 is negative, it means that the first control does not have any sibling controls belonging to the same control group. In this case, the electronic device may no longer execute the instruction corresponding to the first voice. Furthermore, in some examples, the electronic device may issue a prompt message (such as displaying a prompt message on the screen) to indicate to the user that it cannot be clicked.

[0143] If the judgment result of S209 is yes, it means that the first control has a sibling control belonging to the same control group. In this way, the instruction corresponding to the first voice can be executed based on the sibling control. Before this, the electronic device also needs to determine whether the sibling control is clickable, as in S210.

[0144] The interface displayed by an electronic device includes a control group with two or more controls, typically including an image control and a text control. In one scenario, the text control in the same control group is not clickable, while the image control is clickable. In some embodiments, S209 may specifically include: determining whether there is a sibling image control corresponding to the first control. In this embodiment, if there is a sibling image control corresponding to the first control, then S210 is executed.

[0145] S210. Determine whether the second control is clickable.

[0146] For details on how to determine whether a second control is clickable, please refer to the explanation on determining whether a first control is clickable.

[0147] If the determination result of S210 is negative, it indicates that the second control is not clickable. In this case, the electronic device may no longer execute the instruction corresponding to the first voice. Furthermore, the electronic device may issue a prompt message to inform the user that the control is not clickable, as in S215. For example, the electronic device may issue the prompt message by displaying text or image prompt messages on the screen; and / or by issuing voice prompt messages through a speaker; and / or by issuing prompt messages through vibration.

[0148] If the judgment result of S210 is yes, it means that the second control is clickable. In this way, the electronic device can execute the command corresponding to the first voice based on the second control.

[0149] S211. Get the area corresponding to the second control.

[0150] In some embodiments, the electronic device can obtain the area corresponding to the second control by acquiring the control information of the second control. Specifically, acquiring the area corresponding to the control can mean acquiring the coordinate position of the area corresponding to the control on the display interface.

[0151] Controls on the display interface of electronic devices can be presented in various shapes, the most common being rectangular controls. Taking a rectangular control as an example, the coordinate position of the area corresponding to the control on the display interface can be determined in the following way: obtain the coordinates of two corner points on one diagonal of the control (such as the coordinates of the top left corner and the bottom right corner, or the bottom left corner and the top right corner). If the second control is a rectangular control, then the above S211 can specifically obtain the coordinates of two corner points on one diagonal of the second control, such as the coordinates of the top left corner and the bottom right corner.

[0152] The coordinates of points on the display interface of an electronic device can be represented using coordinates in the screen coordinate system. In some examples, the top-left corner of the screen is the origin of the screen coordinate system (0,0). The positive X-axis extends to the right from the origin, and the positive Y-axis extends downwards from the origin.

[0153] If the control is of another shape, such as a circle, the coordinates of the area corresponding to the control on the display screen can be determined by obtaining the coordinates of the control's center point and radius. In other embodiments, if the control is an ellipse, polygon, or other shape, the coordinates of the area corresponding to the control on the display screen can also be determined in other ways.

[0154] S212. Perform a simulated click operation based on the area corresponding to the second control.

[0155] In some embodiments, S212 may specifically include: performing a simulated click operation within the area corresponding to the second control. Further, performing a simulated click operation within the area corresponding to the second control may specifically include: obtaining the center position of the area corresponding to the second control as the simulated click position, and performing a simulated click operation at the simulated click position.

[0156] In other embodiments, S212 may specifically include: splicing the areas corresponding to the first control and the second control to obtain a spliced ​​area. Then, a simulated click operation is performed within this spliced ​​area. In some examples, performing a simulated click operation within the spliced ​​area may specifically include: obtaining the center position of the spliced ​​area as the simulated click position; and then performing a simulated click operation at the simulated click position.

[0157] The specific implementation process of the electronic device performing a simulated click operation at the simulated click location can be referred to the description in related technologies, and will not be repeated in the embodiments of this application.

[0158] Please refer to Figure 4 In some examples, mobile phone 100 displays the "My" interface 301 of a video application, and the user inputs the voice command "Coupon Package" into the electronic device. In response to receiving the voice command "Coupon Package," mobile phone 100 searches for a text control, i.e., text control 302, that matches the voice command "Coupon Package" in the "My" interface 301. Then, mobile phone 100 determines whether text control 302 is clickable. If it determines that text control 302 is not clickable, mobile phone 100 can search for a second control belonging to the same control group as text control 302 in the "My" interface 301. In one example, mobile phone 100 finds that the second control belonging to the same control group as text control 302 is an image control 303. After mobile phone 100 determines that the image control 303 is clickable, it can execute the command corresponding to the voice command "Coupon Package" based on the image control 303, i.e., simulate a click command.

[0159] In some embodiments, the mobile phone 100 can obtain the area corresponding to the image control 303 and execute a simulated click command within the area corresponding to the image control 303.

[0160] In other embodiments, the mobile phone 100 can concatenate the area corresponding to the image control 303 with the area corresponding to the text control 302, and then execute a simulated click instruction based on the concatenated area; that is, perform a simulated click operation within the concatenated area.

[0161] For example, after the mobile phone 100 executes the simulated click command corresponding to the voice "coupon package" based on the image control 303, it can open the coupon package and display the coupon package interface 304.

[0162] In the technical solution proposed in this application, after the electronic device receives a first voice message containing a control click intention, if it determines that the first control corresponding to the control click intention is not clickable, it searches for a second control belonging to the same control group as the first control on the current display interface. If a clickable second control belonging to the same control group as the first control exists on the current display interface, the instruction corresponding to the first voice message can be executed based on the second control. In this way, if the received voice intention is accurately recognized, the possibility that the electronic device cannot realize the true intention of the user's voice input can be reduced, and the possibility of responding to the user's voice input and executing an accurate simulated click operation can be increased.

[0163] If the judgment result of S208 above is yes, the electronic device can directly perform a simulated click operation on the first control to realize the user's intention. For example... Figure 3 S213 and S214 are shown.

[0164] S213. Get the area corresponding to the first control.

[0165] S214. Perform a simulated click operation in the area corresponding to the first control.

[0166] like Figure 2B In the example shown, the first control corresponding to the user's voice input "Open History" is the "History" control 111. This "History" control 111 is clickable, allowing the mobile phone 100 to directly respond to the user's voice input and perform a simulated click operation on the "History" control 111. This also achieves the user's intent.

[0167] If the judgment result in S209 is negative or the judgment result in S210 is negative, the electronic device cannot perform a simulated click operation on the first control. At this time, the electronic device can issue a prompt message, such as S215.

[0168] S215. Issue a prompt message.

[0169] This message indicates to the user that the first control is not clickable. This serves as a reminder to the user of the electronic device's response to the first voice command.

[0170] Furthermore, electronic devices can display controls in various forms, and some control groups may include more than three controls. In other embodiments, if the determination result of S209 is yes, there may be more than two second controls belonging to the same control group as the first control. In this case, the electronic device cannot determine which sibling control to perform the simulated click operation on. If one sibling control is selected to perform the simulated click operation, it may result in a situation that does not conform to the user's intention. Therefore, in some embodiments, after S209 and before S210, the method may further include: determining whether the number of second controls is 1. In this embodiment, the electronic device executes S210 and the subsequent process only when it determines that the number of second controls is 1. That is, the determination of whether the second control is clickable is only made when the first control includes only one second control. And the electronic device will only perform the simulated click operation based on the second control when the second control is clickable.

[0171] In other embodiments, if the electronic device determines that there are two or more second controls belonging to the same control group as the first control, the electronic device may not respond to the first voice prompt. Furthermore, the electronic device may issue a prompt message, such as... Figure 3 S215 is shown.

[0172] In the technical solution proposed in this application, if the first control matching the first voice is unclickable, and only one second control belonging to the same control group as the first control is found, then a simulated click operation is performed based on the second control. This ensures that the electronic device performs the simulated click operation at the accurate location, and the execution result more closely matches the user's intent.

[0173] In some scenarios, electronic devices can display multiple windows simultaneously by stacking them. In scenarios where an electronic device responds to a user's voice to perform a simulated click, the last clicked location might be obscured by a window such as a floating window. If this location is obscured, the electronic device might click on another window during the simulated click operation, causing the simulated click to malfunction.

[0174] like Figure 5AAs shown, mobile phone 100 displays a settings interface 401 and a calculator floating window 402. The settings interface 401 includes a WLAN control 403, a Bluetooth control 404, and a mobile network control 405. Some controls on the settings interface 401 are obscured by the calculator floating window 402. Specifically, the WLAN control 403 in the settings interface 401 is completely obscured; the Bluetooth control 404 is partially obscured; and the mobile network control 405 is not obscured.

[0175] for Figure 5A In the settings interface 401 shown, the WLAN control 403 and Bluetooth control 404 may trigger a click on the calculator floating window 402 if the electronic device performs a simulated click operation on these two locations. Figure 5B As shown, when the user inputs the voice command "Bluetooth," the phone 100 performs a simulated click operation at position 404a when executing the command corresponding to that language. At this time, the phone 100 will click the control corresponding to the number "0" in the calculator floating window 402. Consequently, the calculator floating window 402 on the phone 100 updates to 402a, displaying the selected number "0." This will cause a problem where the simulated click on the electronic device fails.

[0176] or, Figure 4 The image control 303 shown may also be obscured by floating windows, etc., causing the electronic device to be unable to accurately perform simulated click operations on the image control 303.

[0177] Based on this, this application also proposes a voice control method, which can also be applied to electronic devices that support voice input. Figures 6-8 The specific implementation of this voice control method is shown.

[0178] Figure 6 Here is a flowchart of a voice control method. In this embodiment, the method includes:

[0179] S500. Displays the first interface, which includes the first switch.

[0180] S501. Receives the operation of turning on the first switch.

[0181] S502. In response to the operation of turning on the first switch, the voice control function is activated.

[0182] S503. Display the second interface.

[0183] S504. Receive the first voice message.

[0184] S505. Determine the user intent corresponding to the first voice.

[0185] S506. Determine if the user's intent is a control click intent.

[0186] S507. On the second interface, find the clickable control that matches the first voice.

[0187] In an embodiment where an electronic device displays multiple windows simultaneously, the second interface is the interface displayed by the focus window.

[0188] S508. Determine if the control to be clicked is obscured.

[0189] Controls displayed on the screen of an electronic device may be obscured by floating windows, pop-ups, or other windows, preventing the device from performing simulated click operations on these controls. Combined with... Figure 5A As the example shown illustrates, controls displayed on an electronic device may be obscured by a floating window. After the electronic device identifies the clickable control that matches the first voice, it can determine whether a floating window exists on the electronic device.

[0190] In other embodiments, the electronic device may also analyze the displayed content of the current display interface to determine whether the control to be clicked is obscured.

[0191] Combination Figure 5A As the examples show, there are two scenarios when a control is obscured: completely obscured and partially obscured. Understandably, if a control is completely obscured, the electronic device cannot perform a simulated click operation on that control. If a control is partially obscured, the electronic device can perform a simulated click operation on that control, but it might click on another window, causing the simulated click to fail. Therefore, the electronic device can also determine whether the control to be clicked is completely obscured, as in S509.

[0192] S509. Is the control to be clicked completely obscured?

[0193] A control being completely obscured means that the area corresponding to the control is completely covered by other windows. Figure 5A In the example shown, the WLAN control 403 is completely obscured; the Bluetooth control 404 is not completely obscured (i.e., it is partially obscured).

[0194] If the result of S509 is yes, it means that the control to be clicked is completely obscured. In this case, the electronic device cannot perform a simulated click operation on the control to be clicked. At this time, the electronic device can execute S510.

[0195] S510. Issue a prompt message.

[0196] This message indicates to the user that the control to be clicked is not clickable. This serves as a reminder of the response to the initial voice prompt.

[0197] If the result of S509 is negative, it means that the control to be clicked is not completely obscured. In this case, the electronic device can perform a simulated click operation within the area where the control is not obscured. To avoid simulation click errors, the electronic device can perform a simulated click operation within the unobscured area.

[0198] S511. Redefine the simulated click area.

[0199] Since the control to be clicked is not completely obscured, it means that there is still a portion of the control that can be used for simulated click operations. In some embodiments, S511 may specifically include: obtaining the unobscured portion of the control to be clicked as the simulated click area.

[0200] In one example, obtaining the unobstructed area of ​​the control to be clicked can specifically include: determining the overlapping area between the obstructing window and the control to be clicked; comparing the area corresponding to the control to be clicked with the overlapping area to determine the unobstructed area of ​​the control. This allows for quick determination of the simulated click area.

[0201] When determining the simulated click area, the border of the simulated click area can be determined. Taking a control to be clicked and an obscuring window as both being rectangles as an example, the determined simulated click area should also be a rectangle. In this embodiment, determining the simulated click area specifically involves determining the coordinates of two corner points on one diagonal of the simulated click area. For example, S511 above can specifically determine the coordinates of the upper left and lower right corner points of the simulated click area, or determine the coordinates of the upper right and lower left corner points of the simulated click area. In other embodiments, the electronic device can also determine the border of the simulated click area in other ways.

[0202] S512. Perform a simulated click operation based on the simulated click area.

[0203] In some embodiments, S512 may specifically include: obtaining the center position of the simulated click area and performing a simulated click operation at the center position of the simulated click area.

[0204] Taking the example of an electronic device performing a simulated click operation at the center of the area corresponding to the clickable control, in the above embodiment, if the judgment result of S509 is negative, S511 and S512 are executed directly. However, in reality, one situation where the judgment result of S509 is negative is when the clickable control is partially obscured, but the center of the area corresponding to the clickable control is not obscured. In this case, the electronic device does not need to redetermine the simulated click area, but can still directly perform the simulated click operation at the center of the area corresponding to the clickable control. Since the center of the area corresponding to the clickable control is not obscured, the electronic device will not encounter the problem of simulated click error when performing the simulated click operation.

[0205] Therefore, in some embodiments, if the determination result of S509 is negative, before S511 and S512, the above method may further include: determining whether the center position of the control to be clicked is obscured. In this embodiment, the electronic device can execute S511 and S512 only if the center position of the control to be clicked is obscured. In another example, if the determination result of S509 is negative and the center position of the control to be clicked is not obscured, the electronic device can directly perform a simulated click operation at the center position of the control to be clicked.

[0206] Furthermore, if the judgment result of S508 is negative, it indicates that the control to be clicked is not obscured. In this case, the electronic device can directly perform a simulated click on the control to be clicked, as in S513 and S514.

[0207] S513. Get the area corresponding to the control to be clicked.

[0208] S514. Perform a simulated click operation in the area corresponding to the control to be clicked.

[0209] It should be noted that the specific implementation process of some steps in S501-S514 above can be referred to the description of the corresponding steps of S201-S214 in the above embodiments; it will not be repeated here.

[0210] In the technical solution proposed in this application, if the electronic device detects that the clickable control matching the first voice is completely obscured, it will not perform the simulated click operation to avoid the simulated click operation clicking on other windows and causing the simulated click to fail. However, if the electronic device detects that the clickable control is not completely obscured, it will redetermine the simulated click area before performing the simulated click operation. This ensures that the clickable control is accurately clicked during the simulated click, avoiding the problem of simulated click errors in such scenarios. Therefore, the electronic device, responding to the user's voice input containing the intention to click the control, can perform the simulated click operation at the accurate location, thereby accurately realizing the user's intention.

[0211] Taking a floating window that might obscure a clickable control as an example, such as... Figure 7A As shown, the above S508 may include S601-S604:

[0212] S601. Check if a floating window exists.

[0213] In some embodiments, the application corresponding to the voice control function can register a floating window listener. When the electronic device displays the floating window, it can notify the application corresponding to the voice control function through this floating window listener. In this way, when the application corresponding to the voice control function of the electronic device needs to perform a simulated click operation in response to the user's voice input, it can check whether a floating window exists on the current display interface of the electronic device.

[0214] In this embodiment, if the determination result of S601 is negative, it means that the electronic device does not currently have a floating window. That is, the control to be clicked is not obscured by a floating window. In this case, the electronic device can execute S513 and S514.

[0215] If the result of S601 is yes, it indicates that a floating window exists. The electronic device can then determine whether the control to be clicked is obscured by the floating window based on the area corresponding to the floating window and the area corresponding to the control to be clicked, i.e., S602-S604.

[0216] S602. Determine if the floating window is the focused window.

[0217] In some embodiments, the electronic device may obtain the window identifier of the currently focused window, and then determine whether the floating window is the focused window based on the window identifier of the focused window.

[0218] In some examples, the window identifier of the focus window can also be the border of that focus window. In this embodiment, the electronic device can compare the border of the floating window with the border of the focus window to determine whether the floating window is the focus window. This makes it easier to confirm whether the control to be clicked is obscured by the floating window.

[0219] In other examples, the window identifier of the focused window can specifically be the application name of the application displayed in the focused window or the activity name of the activity. In this embodiment, the electronic device can compare the application name or activity name corresponding to the floating window with the focused window to determine whether the floating window is the focused window. This makes it easier to confirm whether the control to be clicked is obscured by the floating window.

[0220] If the floating window is the focused window, it means that the control to be clicked is a control on the floating window. In this case, the electronic device can directly perform a simulated click operation on the control to be clicked. That is, if the judgment result of S602 above is yes, then the electronic device can execute S513 and S514.

[0221] If the floating window is not the focused window, it means that the control to be clicked is not a control on the floating window, and the control to be clicked may be obscured by the floating window. In this case, the electronic device needs to further determine whether the control to be clicked is obscured by the floating window. For example, the electronic device can execute S603 and S604.

[0222] S603. Obtain the area 1 corresponding to the control to be clicked and the area 2 corresponding to the floating window.

[0223] For details on how to obtain the area corresponding to the control and the area corresponding to the floating window, please refer to the descriptions in the above embodiments and related technologies.

[0224] S604. Compare region 1 and region 2 to determine whether the control to be clicked is obscured by the floating window.

[0225] In this embodiment, the electronic device's display interface includes a floating window. The electronic device can determine whether the clickable control is obscured by the floating window based on whether the area corresponding to the floating window overlaps with the area corresponding to the clickable control. The floating window is typically displayed floating on the background interface. If the area corresponding to the floating window overlaps with the area corresponding to the clickable control, it indicates that the clickable control is obscured by the floating window.

[0226] If the result of S604 is negative, it means that although a floating window exists on the current display interface of the electronic device, the control to be clicked is not obscured by the floating window. In this case, the electronic device can also directly perform a simulated click operation on the target text, i.e., execute S513 and S514.

[0227] If the result of S604 is yes, the electronic device can continue to execute S509 to determine whether the control to be clicked is completely obscured.

[0228] Figure 7B This is a schematic diagram illustrating a scenario example of the voice control method according to an embodiment of this application. The mobile phone 100 displays a settings interface 401. In response to the user's voice input of "Bluetooth," the mobile phone 100 performs a simulated click operation at location 402b. Afterwards, the mobile phone 100 opens the Bluetooth submenu interface 402c. This accurately realizes the user's intent and avoids the problem of simulated click errors.

[0229] Figure 7CThis is a schematic diagram illustrating a scenario example of the voice control method according to an embodiment of this application. The mobile phone 100 displays a settings interface 401. In response to the user's voice input of "WLAN," the mobile phone 100 detects that the WLAN control 403 is completely obscured by a floating element. Therefore, the mobile phone 100 can display a prompt message 403a.

[0230] Figure 8 It shows Figure 5A In the example, the positional relationship between the partially obscured Bluetooth control 404 and the calculator floating window 402. In some embodiments, combined with Figure 8 Taking the control to be clicked and the floating window as both being rectangles as an example, the above S511 can be implemented in the following way: In this embodiment, the upper left corner of the screen is taken as the origin coordinate (0,0); the rightward extension of the origin is the positive direction of the X-axis, and the downward extension of the origin is the positive direction of the Y-axis. The established coordinate system is denoted as the screen coordinate system.

[0231] In the quadrant where x>0 and y>0, the area corresponding to the clickable control (i.e., Bluetooth control 404) is represented by coordinates as follows: top left corner point A(x1,y1), bottom right corner point B(x2,y2). The area corresponding to the floating window (i.e., calculator floating window 402) is represented by coordinates as follows: top left corner point C(x3,y3), bottom right corner point D(x4,y4).

[0232] The area corresponding to the control to be clicked overlaps with the area corresponding to the floating window. Figure 8 The simulated click area is defined as the shaded area 404-1, where the control to be clicked is not completely obscured. The simulated click area to be determined is the coordinates of the top-left and bottom-right corners of the largest unobscured rectangular area of ​​the control to be clicked. In some examples, the simulated click area can be represented by coordinates as follows:

[0233] The top-left corner point is M(max(x1,x3),max(y1,y3)), and the bottom-right corner point's coordinates are N(min(x2,x4),min(y2,y4)). In... Figure 8 In the example shown, the calculated simulated click area (i.e. Figure 8 The coordinates of region 404-2 in the image are M(x1,y4) and N(x2,y2).

[0234] It should be noted that, Figure 8 The method shown for determining the simulated click area is only one example; in other embodiments, the electronic device may determine the simulated click area in other ways.

[0235] In the technical solution proposed in this application embodiment, the electronic device determines whether the floating window is obscured by the floating window by judging whether the floating window is a focus window on the current display interface, and if the floating window is not a focus window, by comparing the area corresponding to the floating window with the area corresponding to the control to be clicked. This allows for quick and convenient determination of whether the control to be clicked is obscured, facilitating the electronic device to respond to the first voice command and perform a simulated click operation.

[0236] Figure 9A The flowchart of another voice control method proposed in the embodiments of this application is shown.

[0237] exist Figure 9A In the illustrated embodiment, the electronic device activates the voice control function in response to receiving a user's operation to turn on the first switch. If the electronic device determines that the received voice command corresponds to a control click intent, it can first determine whether the first control is clickable. The process by which the electronic device determines whether the first control is clickable can be referred to... Figure 3 The process flow in steps S207-S211 and S215 of the method shown is as follows: After determining that the first control is clickable, or that there is a clickable second control for the first control, it can be further determined whether the first control (or the second control) is obscured. Figure 6 The process flow in steps S506-S510 of the method shown is as follows. Finally, if the first control (or the second control) (i.e., the control to be clicked) is not obscured, or is not completely obscured, the electronic device can respond to the first voice command by performing a simulated click operation. Figure 6 The flowchart shown illustrates steps S511-S512 or S513-S514. In another example, if the first control is unclickable and there are no sibling controls, or if the first control (or the second control) is completely obscured, the electronic device can issue a prompt indicating that it is unclickable.

[0238] In the technical solution proposed in this application, when receiving voice indicating a control click intention, the electronic device first determines whether a first control matching the voice is clickable, and then determines whether it is obstructed. This reduces the possibility that the electronic device cannot realize the true intention of the user's voice input, while ensuring accurate voice intent recognition, and increases the likelihood of accurately executing simulated click operations in response to user voice input.

[0239] In other embodiments, Figure 9AIn the voice control method shown, if the electronic device determines that a first control matching the first voice is clickable, it will perform an occlusion check on that first control. Furthermore, if the electronic device determines that the first control is completely obscured, it means that the electronic device cannot execute a simulated click operation for the voice click command on that first control. In this case, the electronic device can re-determine the control to be clicked. Specifically, the electronic device can find a second control belonging to the same control group as the first control and determine that second control as the new control to be clicked. Then, the electronic device will determine whether the second control is obscured by a floating window.

[0240] If the second control is also completely obscured, it means that the electronic device cannot respond to the first voice command to perform a simulated click operation.

[0241] If the second control is not obscured, it means that the electronic device can respond to the first voice by performing a simulated click operation on the second control.

[0242] If the second control is partially obscured, the electronic device can redefine the simulated click area based on the areas corresponding to the floating window and the second control, respectively. Finally, the electronic device performs a simulated click operation within this newly determined simulated click area.

[0243] In this embodiment, the specific implementation process of the electronic device finding the second control that belongs to the same control group as the first control can be referred to the description in the above embodiment.

[0244] by Figure 9B For example, mobile phone 100 displays a settings interface 406, which also includes a calculator floating window 407. In this example, the text control 408 in the WLAN settings option is the first control mentioned above. When determining the click attribute of the text control 408, mobile phone 100 determines that the click attribute of the text control 408 is clickable. Then, mobile phone 100 performs an occlusion determination on the text control 408. It can be determined that the text control 408 is completely obscured by the calculator floating window 407. In this case, mobile phone 100 can find a second control belonging to the same control group as the text control 408, such as... Figure 9B The image control 409 in the WLAN settings options is shown. Then, the mobile phone 100 performs an occlusion check on the image control 409. In one example, such as... Figure 9B As shown, the image control 409 is not obscured by the floating window. At this time, the mobile phone 100 can perform a simulated click operation on the image control 409, which can also realize the user's intention to open the WLAN settings option and display the WLAN settings submenu.

[0245] Understandably, in another example, if the result obtained by the mobile phone 100 after judging the occlusion of the image control 409 is that the image control 409 is partially occluded, then the mobile phone 100 can redetermine the simulated click area based on the corresponding areas of the calculator floating window 407 and the image control 409, and then perform a simulated click operation on the redetermined simulated click area.

[0246] In another example, if the mobile phone 100 determines that the image control 409 is completely obscured after performing an occlusion check, the mobile phone 100 can issue a prompt message. This prompt message indicates that the image control 409 is not clickable.

[0247] In the technical solution proposed in this application, when the electronic device performs a simulated click operation in response to voice, it simultaneously considers whether the control to be clicked is clickable and whether it is obscured. This further ensures that the control to be clicked can execute a click event. Therefore, while ensuring accurate recognition of voice intent, this reduces the possibility that the electronic device cannot realize the true intent of the user's voice input, and increases the likelihood of accurately performing a simulated click operation in response to user voice input.

[0248] Figures 10-13 Implementation shown Figures 3-4 The software architecture of the electronic device using the voice control method is shown, as well as the interaction flow of each module of the electronic device when implementing the voice control method.

[0249] Figure 10 This document describes the software architecture of an electronic device in some embodiments. In some embodiments, the software system of the electronic device may adopt a layered architecture, an event-driven architecture, a microkernel architecture, or a cloud architecture. This application embodiment uses a layered architecture. Taking the system as an example, the software structure of the electronic device is illustrated.

[0250] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, [the following is omitted as the text is incomplete and likely refers to a specific implementation or feature]. The system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0251] The application layer may include a series of application packages (APKs). Examples include applications for camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS. In this embodiment, the application layer includes a voice control APK for providing voice control functionality for the electronic device. The voice control APK includes an interface monitoring module, an interface content acquisition module (content sensor), an interface parsing module (view fetcher), and an application interaction module. The interface monitoring module monitors the application interface's startup, exit, and switching processes. The interface content acquisition module acquires the top-level activity. The interface parsing module parses the top-level activity to obtain the content of the currently displayed page (including images and / or text). The application interaction module manages the process of the application executing commands based on voice.

[0252] The application layer also includes a speech processing engine for parsing, recognizing, and processing speech. Specifically, the speech parsing module converts speech into text; this module can be part of the ASR (Automatic Speech Recognition) engine. The speech recognition module understands and recognizes the semantics of the text and determines whether it matches the interface text; this module can be part of the NLU (Natural Language Understanding) engine. The instruction mapping module converts the recognized semantics into machine-executable instructions; this module can be part of the DM (Demand Mapping) engine.

[0253] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0254] like Figure 11 As shown, the application framework layer may include a window manager, content provider, view system, resource manager, activity manager service (AMS), application manager service (PMS), and multi-modal control module, etc.

[0255] AMS is primarily responsible for starting, switching, and scheduling the four major components of the system, as well as managing and scheduling application processes. Its responsibilities are similar to those of the process management and scheduling module in an operating system. When a process or component starts, the request is passed to AMS through the inter-process communication (binder) mechanism, and AMS then processes it uniformly.

[0256] The multi-mode control module is used to manage the execution of voice commands. For example, it sends commands to applications, causing the applications to execute those commands.

[0257] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0258] The Android runtime is responsible for scheduling and managing the Android system. The Android runtime includes core libraries and a virtual machine.

[0259] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0260] The kernel layer is the layer between hardware and software. The kernel layer can contain display drivers, sensor drivers, microphone drivers, Wi-Fi drivers, etc.

[0261] Combination Figure 10 , Figure 11 This diagram illustrates the interaction of various modules of an electronic device when implementing the voice control method provided in the embodiments of this application.

[0262] AMS can monitor the lifecycle of the activity corresponding to each interface of each application on an electronic device.

[0263] In some examples, the UI content acquisition module can send a request to the ActivityManagerService (AMS). In response to this request, the AMS returns the top-level activity to the UI content acquisition module. In some embodiments, sending a request to the AMS by the UI content acquisition module may specifically include: the UI content acquisition module sending a request to the AMS via ActivityManagerEx.RequestContentNode.

[0264] In other examples, the UI content fetching module can register a listener with AMS. When AMS detects changes in the application's lifecycle, it can notify the top-level activity via a callback to the UI content fetching module.

[0265] After receiving the top-level activity from AMS, the UI content acquisition module sends the top-level activity to the UI parsing module. The UI parsing module parses the top-level activity to obtain the display content of the electronic device, including the text controls and / or image controls included therein.

[0266] Furthermore, the interface parsing module can save the text controls and / or image controls obtained from parsing the top-level activity to a cache. In some embodiments, the top-level activity obtained by the interface parsing module may include: control information of each control group. Specifically, the control information of a control group may include: control information of each control contained in the control group. Taking a control group that includes an image control and a text control as an example, the control information of the control group may include: control group [coordinate 1, coordinate 2]: image control [coordinate 3, coordinate 4], text control [coordinate 5, coordinate 6].

[0267] In one example, the control information for the image control in the control group may include at least the contents of Table 1.

[0268] Table 1

[0269]

[0270] The control information for the text controls in the control group can include at least the contents of Table 2.

[0271] Table 2

[0272]

[0273] The label indicates the control's sequence number within the control group.

[0274] The interface parsing module parses the top-level activity to obtain control information for each control or control group on the currently displayed screen. In some embodiments, the interface parsing module can save the control information of text controls and / or image controls to a cache. This control information may include: the area where the control is located (rectangular controls can be represented by the coordinates of the top-left and bottom-right corners), whether the control is clickable, and the control group to which the control belongs.

[0275] After a user inputs voice into an electronic device, the device's microphone picks up the voice. The microphone then inputs the picked-up voice into the application interaction module of the voice control APK. The application interaction module is responsible for distributing the received voice; in some examples, the application interaction module can distribute the voice to the voice parsing module in the voice processing engine, where the voice parsing module parses the user's input.

[0276] The speech parsing module parses the speech to obtain the corresponding text. Then, the speech parsing module sends the obtained text to the speech recognition module.

[0277] The speech recognition module performs semantic understanding and recognition on the text corresponding to the speech to obtain the intent behind the speech; that is, the user's intent, the operation the user wants to control the electronic device to perform through the voice. Then, the speech recognition module can send the intent corresponding to the speech to the command mapping module.

[0278] The instruction mapping module converts the intent mapping corresponding to the speech into machine-executable instructions (denoted as instruction a).

[0279] In some embodiments, after the instruction mapping module converts the user intent into instruction 'a', it can determine whether instruction 'a' is a control click instruction. If the obtained instruction 'a' is not a control click instruction, the instruction mapping module can directly send instruction 'a' to the corresponding application through the multi-mode control module, notifying the application to execute instruction 'a'.

[0280] If the received instruction 'a' is a control click instruction, the instruction mapping module can send the control click instruction to the interface parsing module. This control click instruction carries text that matches the parsed text corresponding to the first speech (i.e., the matched text).

[0281] In other embodiments, the command mapping module can, after receiving the user intent sent by the speech recognition module, also determine whether the user intent is a control click intent. Furthermore, if the command mapping module determines that the user intent is a control click intent, it can convert the user intent into instruction 'a' and send instruction 'a' to the interface parsing module. If it determines that the user intent is not a control click intent, the command mapping module, after converting the user intent into instruction 'a', can directly send it to the corresponding application through the multi-mode control module, notifying the application to execute instruction 'a'.

[0282] The following example illustrates the concept of user-inputted voice indicating a control click intention, where instruction 'a' is a control click instruction. In this embodiment, after the interface parsing module receives the control click instruction sent by the instruction mapping module, it can search the cache for the control information of the corresponding control (i.e., the first control, hereinafter referred to as the target text control) based on the matched text carried in the control click instruction.

[0283] In this embodiment, the interface parsing module retrieves the control information of the target text control from the cache and determines whether the target text control is clickable. In some examples, if the target text control is not clickable, the interface parsing module can search the cache for a second control belonging to the same control group as the target text control (hereinafter referred to as the target sibling control).

[0284] If the interface parsing module finds a target sibling control corresponding to the target text control, it can also query the number of target sibling controls, the area corresponding to the target sibling control, and whether the target sibling control is clickable.

[0285] In some embodiments, if the interface parsing module determines that there exists a target sibling control belonging to the same control group as the target text control and that the target text control is clickable, it can execute instruction 'a' based on the area corresponding to the found target sibling control. In another example, if the interface parsing module determines that there is no target sibling control belonging to the same control group as the target text control, or that all sibling controls belonging to the same control group as the target text control are not clickable, then it determines that the control click instruction cannot be executed.

[0286] In other embodiments, if the interface parsing module determines that there exists a target sibling control belonging to the same control group as the target text control, and there is only one target sibling control that is clickable, it can execute instruction a based on the area corresponding to the target sibling control. In this embodiment, if the interface parsing module determines that there are multiple sibling controls belonging to the same control group as the target text control, or that there is only one sibling control belonging to the same control group as the target text control that is not clickable, then it determines that instruction a cannot be executed.

[0287] Furthermore, after determining that instruction a can be executed, the interface parsing module can notify the corresponding application to execute instruction a through the multi-mode control module.

[0288] Once the interface parsing module determines that the click command for a control cannot be executed, it can send a notification message to the application interaction module. Upon receiving this notification message, the application interaction module can issue a prompt indicating that the control is not clickable.

[0289] The following is combined Figure 12 The module interaction describes the specific process by which the interface parsing module determines whether a target text control is clickable. In this embodiment, the area corresponding to the control is the control's border.

[0290] In this embodiment, the interface parsing module includes a clickability determination submodule and a sibling control query submodule. The clickability determination submodule, upon receiving a control click instruction from the instruction mapping module, queries the cache to determine whether the target text control is clickable. The sibling control query submodule can be used to query control information corresponding to the target text control.

[0291] If the target text control is clickable, the interface parsing module sends a control click instruction to the multi-module control module through the clickability determination submodule. This control click instruction carries the border of the target text control; that is... Figure 12 The first case shown.

[0292] If the target text control is not clickable, the clickability determination module sends a notification message to the sibling control query submodule. In response to this notification message, the sibling control query submodule retrieves sibling controls belonging to the same control group as the target text control from the cache, along with the clickability properties (i.e., whether they are clickable) and borders of these sibling controls.

[0293] The sibling control query submodule can determine whether a target text control has sibling controls based on the query results.

[0294] If the target text control has a sibling control that is clickable, the interface parsing module sends a control click instruction to the multi-module control module through the clickability determination submodule. This control click instruction includes the border of the sibling control; that is... Figure 12 The second scenario shown.

[0295] If the target text control has no sibling controls, the interface parsing module returns a notification message indicating that it is not clickable to the application interaction module through the clickability determination submodule; that is... Figure 12 The third scenario shown.

[0296] If the target text control has no sibling controls, the interface parsing module returns a notification message indicating that it is not clickable to the application interaction module through the clickability determination submodule; that is... Figure 12 The fourth scenario shown.

[0297] If the target text control has multiple sibling controls, the interface parsing module returns a notification message indicating that it is not clickable to the application interaction module through the clickability determination submodule; that is... Figure 12 The fifth case shown.

[0298] When the application interaction module receives a notification message from the clickability judgment submodule, it can issue a prompt message.

[0299] Figure 13 This diagram illustrates the timing interaction between modules in the voice control method proposed in this application. In this embodiment, the user's first voice input containing a control click intent is used as an example for explanation. In other embodiments, the commands corresponding to the user's voice input are specific implementations of other commands, which can be referred to... Figure 13 The process shown is not repeated in the embodiments of this application.

[0300] The user turns on the first switch. In response to receiving the user's action of turning on the first switch, the electronic device activates its voice control function.

[0301] The UI content acquisition module registers a listener with AMS. It's worth noting that this registration can occur when voice control is enabled on the electronic device. When AMS detects changes in the application's lifecycle, it can notify the top-level activity via a callback to the UI content acquisition module.

[0302] Then, the UI content acquisition module can send the top-level activity to the UI parsing module, notifying the UI parsing module to perform parsing. The UI parsing module parses the top-level activity to obtain the controls and control information of the currently displayed UI. It then saves the control information to a cache. Furthermore, the UI parsing module sends the text of the currently displayed UI, i.e., the UI hot words, to the speech recognition module.

[0303] The user inputs a first voice message into the electronic device. Specifically, the microphone picks up the user's voice input. Since the voice control function is currently enabled, the recording channel corresponding to the first voice control function is activated. Therefore, the microphone sends the first voice message to the application interaction module of the voice control APK.

[0304] The application interaction module distributes the received first voice message to the voice processing engine for processing. Specifically, the application interaction module sends the message to the voice parsing module within the voice processing engine. Upon receiving the first voice message, the voice parsing module parses it to obtain parsed text. Then, the voice parsing module sends the parsed text to the voice recognition module for intent recognition. The voice recognition module identifies the intent corresponding to the received parsed text, obtaining the user intent. The voice recognition module identifies the user intent as a control click intent. Afterward, the voice recognition module sends the control click intent to the command mapping module for command mapping. Upon receiving the control click intent, the command mapping module maps the control click intent to an executable command of the electronic device, i.e., a control click command.

[0305] Then, the instruction mapping module sends the control click instruction to the interface parsing module, which carries the text of the hit.

[0306] After receiving a control click instruction, the interface parsing module sends a query request to the cache. The cache returns the click attributes and border of the target text control corresponding to the matched text to the interface parsing module. Then, the interface parsing module determines whether the target text control is clickable.

[0307] If the target text control is clickable, the interface parsing module can send a control click instruction to the multimodal control module. This control click instruction carries the border of the target text control. The multimodal control module can then instruct the corresponding application to execute the control click instruction, that is, to perform a simulated click operation within the border of the target text control.

[0308] If the target text control is not clickable, the interface parsing module can cache the query request. The cache returns the sibling control, click attribute, and border to the interface parsing module. Then, the interface parsing module determines if there is a sibling control. If a sibling control exists (i.e., the target sibling control mentioned above), the interface parsing module continues to determine if the number of sibling controls is 1. If the number of sibling controls is 1, the interface parsing module continues to determine if the sibling control is clickable. Finally, if the unique sibling control is clickable, the interface can send a control click instruction to the multimodal control module; this control click instruction carries the border of the sibling control. The multimodal control module can then notify the corresponding application to execute the control click instruction.

[0309] If the target text control has no sibling controls, or if the number of target sibling controls is greater than one, or if the number of target sibling controls is one and that sibling control is not clickable, the interface parsing module notifies the application interaction module that it is not clickable. The application interaction module can then issue a prompt message to indicate that it is not clickable.

[0310] Figures 14-17 Implementation shown Figures 6-8 The software architecture of the electronic device using the voice control method is shown, as well as the interaction flow of each module of the electronic device when implementing the voice control method.

[0311] Figure 14 The following describes the software architecture of an electronic device in some embodiments. In this embodiment, the voice control APK of the electronic device further includes a floating window monitoring module and an occlusion detection module. The floating window monitoring module is used to monitor whether the electronic device displays a floating window and to obtain the border information of the floating window. The occlusion detection module is used to determine whether the control to be clicked is occluded.

[0312] Combination Figure 14 , Figure 15 This diagram illustrates the interaction of various modules of an electronic device when implementing the voice control method provided in the embodiments of this application.

[0313] The floating window monitoring module registers its monitoring with AMS. After AMS detects that a floating window is displayed on the electronic device, it notifies the floating window monitoring module of the window's border via a callback. In some examples, the floating window monitoring module registers its monitoring with AMS when the electronic device enables voice control.

[0314] When the user inputs voice, the voice control APK's application interaction module distributes the voice to the voice processing engine, obtains control click commands, and returns the control click commands to the interface parsing module. For the specific implementation process, please refer to [link / reference needed]. Figure 11 The description process of the corresponding steps.

[0315] In this embodiment, after the interface parsing module queries the border of the clickable control that matches the hit text, it sends the border of the clickable control to the occlusion judgment module.

[0316] The occlusion detection module can query the floating window monitoring module to determine if a floating window currently exists on the electronic device. Simultaneously, the occlusion detection module can also query the currently focused window via AMS. Then, the occlusion detection module can determine whether a floating window exists, and if so, whether the clickable control in the currently focused window is completely obscured. If a floating window exists and the clickable control is completely obscured, the occlusion detection module notifies the application interaction module that it is not clickable. The application interaction module can then issue a prompt indicating that it is not clickable. In other cases, the occlusion detection module sends a control click command to the multi-mode control module, which carries a simulated click area. The multi-mode control module can then instruct the corresponding application to execute the control click command, i.e., perform a simulated click operation within the simulated click area.

[0317] Figure 16 The diagram illustrates the specific process by which the occlusion detection module determines the presence of a floating window, and when a floating window exists, whether the clickable control in the currently focused window is completely obscured. In this embodiment, the occlusion detection module includes a border comparison submodule and a border calculation submodule. The border comparison submodule is used to determine whether a floating window exists, and if so, to compare the clickable control with the floating window to determine whether the clickable control is completely obscured by the floating window. The border calculation submodule is used to redetermine the simulated click area when the clickable control is not completely obscured.

[0318] After receiving the border of the clickable control from the interface parsing module, the border comparison submodule retrieves the border of the floating window from the floating window listening module. In one example, if there is no floating window currently, the floating window listening module returns null to the border comparison submodule.

[0319] If the border comparison submodule receives an empty value from the floating window listener module, it falls into case ①, where there is no floating window. The border submodule can send control click commands to the multi-mode control module, and these commands carry the border of the control to be clicked.

[0320] After the border comparison submodule obtains the border of the floating window, it retrieves the focus window from AMS and determines whether the floating window is the focus window. If a floating window exists and is the focus window, it means that the control to be clicked is a control within the floating window, and there is no possibility that the control to be clicked is obscured by the floating window.

[0321] In other examples, if a floating window exists and is not the focused window, there is a possibility that the clickable control may be obscured by the floating window. In this case, the border comparison submodule can compare the borders of the floating window and the clickable control to determine whether the clickable control is obscured by the floating window.

[0322] If the border comparison submodule determines that the control to be clicked is not obscured by the floating window, then it is case ②. The border comparison sub-control can also send a control click instruction to the multi-mode control module. This control click instruction carries the border of the control to be clicked.

[0323] If the border comparison submodule determines that the control to be clicked is not completely obscured by the floating window, then it falls into case ③. In this case, the border comparison submodule can notify the border calculation submodule to re-determine the simulated click area. Specifically, the border comparison submodule sends the borders of the floating window and the control to be clicked to the border calculation submodule. The border calculation submodule calculates and determines the border of the simulated click area based on the borders of the floating window and the control to be clicked. Afterward, the border calculation submodule can send a control click command to the multi-mode control module, which carries the border of the simulated click area.

[0324] If the border comparison submodule determines that the control to be clicked is completely obscured by the floating window, then it falls into case ④. In this case, the border comparison submodule notifies the application interaction module that the control is not clickable. The application interaction module can then issue a prompt message.

[0325] In cases ①, ②, and ③ above, after receiving a control click instruction, the multi-mode control module can notify the corresponding application to execute the control click instruction.

[0326] Figure 17 This diagram illustrates the timing interaction between modules in the voice control method proposed in this application. In this embodiment, the user's first voice input containing a control click intent is used as an example for explanation. In other embodiments, the commands corresponding to the user's voice input are specific implementations of other commands, which can be referred to... Figure 13 The process shown is not repeated in the embodiments of this application.

[0327] The user turns on the first switch. In response to receiving the user's action of turning on the first switch, the electronic device activates its voice control function.

[0328] The UI content acquisition module registers a listener with AMS. It's worth noting that this registration can occur when voice control is enabled on the electronic device. When AMS detects changes in the application's lifecycle, it can notify the top-level activity via a callback to the UI content acquisition module.

[0329] Then, the UI content acquisition module can send the top-level activity to the UI parsing module, notifying the UI parsing module to perform parsing. The UI parsing module parses the top-level activity to obtain the controls of the currently displayed UI. It then saves the control information to a cache. Finally, the UI parsing module sends the text of the currently displayed UI to the speech recognition module.

[0330] In addition, the floating window listener module registers a floating window listener with AMS, and AMS can notify the floating window listener module of the border of the floating window through callbacks.

[0331] The user inputs a first voice message into the electronic device. Specifically, the microphone picks up the user's voice input. Since the voice control function is currently enabled, the recording channel corresponding to the first voice control function is activated. Therefore, the microphone sends the first voice message to the application interaction module of the voice control APK.

[0332] The application interaction module distributes the received first voice message to the voice processing engine for processing. Specifically, the application interaction module sends the message to the voice parsing module within the voice processing engine. Upon receiving the first voice message, the voice parsing module parses it to obtain parsed text. Then, the voice parsing module sends the parsed text to the voice recognition module for intent recognition. The voice recognition module identifies the intent corresponding to the received parsed text, obtaining the user intent. The voice recognition module identifies the user intent as a control click intent. Afterward, the voice recognition module sends the control click intent to the command mapping module for command mapping. Upon receiving the control click intent, the command mapping module maps the control click intent to an executable command of the electronic device, i.e., a control click command.

[0333] Then, the instruction mapping module sends the control click instruction to the interface parsing module, which carries the text of the hit.

[0334] After receiving a control click instruction, the interface parsing module sends a query request to the cache. The cache returns the border of the clicked control to the interface parsing module, matching the text it found. Then, the interface parsing module sends the border of the clicked control to the occlusion detection module.

[0335] The occlusion detection module can send a query request to the floating window listening module. The floating window listening module returns the border / empty state of the floating window to the occlusion detection module. Then, the occlusion detection module determines whether a floating window exists. If no floating window exists, the occlusion detection module can directly send a control click command to the multi-mode control module, which carries the border of the control to be clicked.

[0336] If a floating window exists, the occlusion detection module sends a query request to the AMS. In response, the AMS returns the current focus window of the electronic device to the occlusion detection module. The occlusion detection module can then determine whether the floating window is the focus window. If it is, the occlusion detection module directly sends a control click instruction to the multi-mode control module, which includes the border of the control to be clicked.

[0337] If the floating window is not the focused window, the occlusion detection module can continue to determine whether the control to be clicked is obscured by the floating window. If the control to be clicked is not obscured by the floating window, the occlusion detection module directly sends a control click command to the multi-mode control module, which carries the border of the control to be clicked.

[0338] If the control to be clicked is obscured by the floating window, the occlusion detection module can further determine whether the control is completely obscured. If the control is not completely obscured, the occlusion detection module combines the border of the floating window and the border of the control to be clicked to calculate a new simulated click area. Then, the occlusion detection module sends a control click command to the multi-mode control module, which carries the border of the new simulated click area.

[0339] If the control to be clicked is completely obscured, the obscuration detection module notifies the application interaction module that it is not clickable. The application interaction module can then issue a notification message indicating that the control is not clickable.

[0340] In some embodiments, having Figure 14 and Figure 15 The electronic device with the software architecture shown can also be used to implement the voice control method shown in Figure 9. In this embodiment, Figure 15 In the module interaction shown, the control to be clicked is either the first control that matches the first voice as determined by the interface parsing module, or the second control that belongs to the same control group as the first control.

[0341] It should be noted that, in this embodiment, during the process of the electronic device recognizing the received voice and executing the corresponding command, the following correspondence exists: The electronic device parses the display interface to obtain: text controls and image controls; wherein, the text contained in the text controls is saved as interface hot words; that is, the text of the text controls in the display interface corresponds one-to-one with the interface hot words of the display interface. The electronic device recognizes the user's input voice to obtain: parsing text - user intent - user intent converted into command - hitting interface hot words - the text control corresponding to the hit interface hot words (denoted as the first control). That is, the voice corresponds one-to-one with the parsed text, user intent, hit interface hot words, the first control, and the command. Some text controls belong to the same control group as other controls. The text controls under the same control group have a one-to-one correspondence with other controls.

[0342] Figure 18 A schematic diagram of the structure of an electronic device 700 provided in an embodiment of this application is shown.

[0343] Electronic device 700 may include a processor 710, an external memory interface 720, an internal memory 721, a universal serial bus (USB) interface 730, a charging management module 740, a power management module 741, a battery 742, antenna 1, antenna 2, a mobile communication module 750, a wireless communication module 760, an audio module 770, a speaker 770A, a receiver 770B, a microphone 770C, a headphone jack 770D, a sensor module 780, buttons 790, a motor 791, a camera 792, a display screen 793, and a subscriber identification module (SIM) card interface 794, etc. The sensor module 780 may include a pressure sensor 780A, a touch sensor 780B, etc.

[0344] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 700. In other embodiments of this application, the electronic device 700 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0345] Processor 710 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. For example, processor 710 is used to execute the voice control method in the embodiments of this application.

[0346] The controller can serve as the nerve center and command center of the electronic device 700. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.

[0347] The processor 710 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 710 is a cache memory. This memory can store instructions or data that the processor 710 has just used or that are used repeatedly. If the processor 710 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 710, and thus improves the efficiency of the system.

[0348] The USB interface 730 is a USB standard compliant interface, which can be a Mini USB interface, Micro USB interface, USB Type-C interface, etc. The USB interface 730 can be used to connect a charger to charge the electronic device 700, and can also be used for data transfer between the electronic device 700 and peripheral devices.

[0349] The external memory interface 720 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 700. The external memory card communicates with the processor 710 through the external memory interface 720 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0350] Internal memory 721 can be used to store executable program code, which includes instructions. Processor 710 executes various functional applications and data processing of electronic device 700 by running the instructions stored in internal memory 721. Internal memory 721 may include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function (such as sound playback, image playback, etc.).

[0351] In addition, the internal memory 721 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0352] The charging management module 740 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 740 can receive charging input from the wired charger via a USB interface 730.

[0353] The power management module 741 is used to connect the battery 742, the charging management module 740, and the processor 710. The power management module 741 receives input from the battery 742 and / or the charging management module 740 to power the processor 710, internal memory 721, external memory, display 793, camera 792, and wireless communication module 760, etc.

[0354] In some other embodiments, the power management module 741 may also be located within the processor 710. In still other embodiments, the power management module 741 and the charging management module 740 may also be located in the same device.

[0355] The wireless communication function of electronic device 700 can be implemented through antenna 1, antenna 2, mobile communication module 750, wireless communication module 760, modem processor and baseband processor, etc.

[0356] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 700 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0357] The mobile communication module 750 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on the electronic device 700. The mobile communication module 750 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 750 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 750 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1.

[0358] The wireless communication module 760 can provide solutions for wireless communication applications on the electronic device 700, including wireless local area networks (WLAN) (such as Wi-Fi), Bluetooth, Global Navigation Satellite System (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR). The wireless communication module 760 can be one or more devices integrating at least one communication processing module. The wireless communication module 760 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signal, and sends the processed signal to processor 710. The wireless communication module 760 can also receive signals to be transmitted from processor 710, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0359] In some embodiments, antenna 1 of electronic device 700 is coupled to mobile communication module 750, and antenna 2 is coupled to wireless communication module 760, enabling electronic device 700 to communicate with networks and other devices via wireless communication technology.

[0360] Electronic device 100 can implement audio functions such as music playback and recording through audio module 770, speaker 770A, receiver 770B, microphone 770C, headphone jack 770D, and application processor.

[0361] Audio module 770 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal. Audio module 770 can also be used for encoding and decoding audio signals. In some embodiments, audio module 770 can be located in processor 710, or some functional modules of audio module 770 can be located in processor 710. Speaker 770A, also called a "loudspeaker," is used to convert audio electrical signals into sound signals. Receiver 770B, also called a "handpiece," is used to convert audio electrical signals into sound signals. Microphone 770C, also called a "microphone" or "microphone," is used to convert sound signals into electrical signals. Headphone jack 770D is used to connect wired headphones. Headphone jack 770D can be a USB interface 730, or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, or a Cellular Telecommunications Industry Association of the USA (CTIA) standard interface. In this embodiment, after the voice control function is activated, electronic device 100 continuously collects ambient sound through microphone 770C.

[0362] Pressure sensor 780A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 780A may be disposed on display screen 793. There are many types of pressure sensors 780A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When a force is applied to pressure sensor 780A, the capacitance between the electrodes changes. Electronic device 700 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 793, electronic device 700 detects the touch operation intensity based on pressure sensor 780A. Electronic device 700 can also calculate the touch position based on the detection signal from pressure sensor 780A.

[0363] Touch sensor 780B, also known as a "touch panel," can be located on display screen 793. The touch sensor 780B and display screen 793 together form a touchscreen, also known as a "touch screen." Touch sensor 780B is used to detect touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 793. In other embodiments, touch sensor 780B may also be located on the surface of electronic device 700, in a different position than display screen 793.

[0364] Buttons 790 include a power button, volume buttons, etc. Buttons 790 can be mechanical buttons or touch-sensitive buttons. Electronic device 700 can receive button input and generate key signal inputs related to user settings and function control of electronic device 700.

[0365] Motor 791 can generate vibration alerts. Motor 791 can be used for incoming call vibration alerts or for touch vibration feedback.

[0366] The camera 792 is used to capture still images or videos. In some embodiments, the electronic device 700 may include one or N cameras 792, where N is a positive integer greater than 1.

[0367] Electronic device 700 implements display functions through a GPU, a display screen 793, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 793 and the application processor. The GPU performs mathematical and geometric calculations and is used for graphics rendering. Processor 710 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0368] The display screen 793 is used to display images, videos, etc. In some embodiments, the electronic device 700 may include one or N display screens 793, where N is a positive integer greater than 1.

[0369] The SIM card interface 794 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 794 to make contact with or separate from the electronic device 700. The electronic device 700 can support one or N SIM card interfaces, where N is a positive integer greater than 1.

[0370] The voice control methods described in the above embodiments can all be executed in electronic devices with the above hardware structure.

[0371] Other embodiments of this application provide an electronic device (such as a mobile phone 100). The electronic device may include a microphone, a display screen, a processor, and a memory. The microphone, display screen, and memory are respectively coupled to the processor. The microphone is used to collect voice; the display screen is used to display the interface of the electronic device; the memory is coupled to the processor. The memory is also used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device can perform various functions or steps performed by the mobile phone 100 in the above method embodiments. The structure of the electronic device can be referred to... Figure 18 The structure of the electronic device 700 shown is illustrated.

[0372] This application also provides a chip system, such as... Figure 19 As shown, the chip system 1000 includes at least one processor 1001 and at least one interface circuit 1002. The processor 1001 and the interface circuit 1002 are interconnected via lines. For example, the interface circuit 1002 can be used to receive signals from other devices (e.g., a computer's memory). As another example, the interface circuit 1002 can be used to send signals to other devices (e.g., the processor 1001). Exemplarily, the interface circuit 1002 can read instructions stored in memory and send those instructions to the processor 1001. When the instructions are executed by the processor 1001, the computer can perform the steps in the above embodiments. Of course, the chip system may also include other discrete devices, which are not specifically limited in this application embodiment.

[0373] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device (such as mobile phone 100), cause the electronic device to perform various functions or steps performed by mobile phone 100 in the above method embodiments.

[0374] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps described in the method embodiments above. The computer may be an electronic device, such as a mobile phone 100.

[0375] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0376] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0377] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0378] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0379] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0380] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A voice control method, characterized in that, Applied to an electronic device, the electronic device including a microphone, the method includes: Display a first interface, the first interface including a first switch; In response to the user's operation of turning on the first switch, the voice control function is activated; activating the voice control function includes: activating the recording channel corresponding to the voice control function, the recording channel being used to acquire the voice collected by the microphone; In response to a voice click command collected by the microphone, a clickable control matching the voice click command is located on the current display interface of the electronic device; When the control to be clicked is partially obscured by the floating window, a first simulated click area is determined based on the area corresponding to the floating window and the area corresponding to the control to be clicked. Within the first simulated click area, execute the click event corresponding to the voice click command; Wherein, the clickable control is partially obscured by the floating window, including: The current display interface of the electronic device includes a floating window, which is not the focus window, and the area corresponding to the floating window and the area corresponding to the control to be clicked overlap; the overlapping area is smaller than the area corresponding to the control to be clicked.

2. The method according to claim 1, characterized in that, Determining the first simulated click area based on the area corresponding to the floating window and the area corresponding to the control to be clicked includes: The first simulated click area is obtained by subtracting the overlapping area from the area corresponding to the control to be clicked.

3. The method according to claim 1, characterized in that, The method further includes: If the current display interface of the electronic device includes a floating window, obtain the area corresponding to the current focus window of the electronic device; Compare the area corresponding to the focus window with the area corresponding to the floating window; If the area corresponding to the focus window is exactly the same as the area corresponding to the floating window, then the floating window is determined to be the focus window.

4. The method according to any one of claims 1-3, characterized in that, The method further includes: When the control to be clicked is not obscured, the click event corresponding to the voice click command is executed within the second simulated click area; the second simulated click area includes the area corresponding to the control to be clicked. Wherein, the fact that the clickable control is not obscured includes: the current display interface of the electronic device does not include a floating window; Alternatively, the current display interface of the electronic device includes a floating window, and the floating window is the focus window; Alternatively, the current display interface of the electronic device includes a floating window, which is not the focus window, and the area corresponding to the clickable control does not overlap with the area corresponding to the floating window.

5. The method according to any one of claims 1-3, characterized in that, The method further includes: If the control to be clicked is completely obscured by the floating window, a prompt message is issued to indicate that it is not clickable.

6. The method according to any one of claims 1-3, characterized in that, The step of searching for a clickable control that matches the voice click command on the current display interface of the electronic device includes: Locate the first control that matches the voice click command on the current display screen of the electronic device; Get the click attribute of the first control; If the clickable property of the first control is clickable, then the first control is determined as the control to be clicked.

7. The method according to claim 6, characterized in that, The voice click command is matched with at least one interface hotword set by the electronic device, and the first control is a text control.

8. The method according to claim 6, characterized in that, The step of searching for a clickable control that matches the voice click command on the current display interface of the electronic device further includes: If the click property of the first control is not clickable, then a second control belonging to the same control group as the first control will be searched. Get the click property of the second control; If the clickable property of the second control is clickable, then the second control is determined as the control to be clicked.

9. The method according to claim 6, characterized in that, The method further includes: If the first control has a clickable property and is completely obscured by the floating window, then find the second control that belongs to the same control group as the first control. Get the click property of the second control; If the clickable property of the second control is clickable, then the second control is identified as the new clickable control, and an occlusion judgment is performed on the new clickable control.

10. The method according to any one of claims 2-3, characterized in that, The electronic device includes a voice control application package (APK), a voice processing engine, and a window activity manager (AMS); the method further includes: The voice control APK receives the current display interface of the electronic device returned by the AMS, parses the current display interface, and obtains and saves the text control and image control of the current display interface; The voice control APK sends the text contained in the text control of the currently displayed interface as interface hot words to the voice processing engine; In response to receiving the voice click instruction, the voice control APK distributes the voice click instruction to the voice processing engine; The voice processing engine searches for target interface hot words that match the voice click command; The voice processing engine returns the target interface hot words to the voice control APK; The voice control APK queries whether the electronic device includes a floating window, and if the electronic device includes a floating window, it determines whether the floating window is the focus window, and whether the area corresponding to the floating window and the area corresponding to the control to be clicked overlap.

11. An electronic device, characterized in that, The electronic device includes: a microphone, a display screen, a memory, and a processor; the microphone, the display screen, and the memory are respectively coupled to the processor; The microphone is used to capture speech, the display screen is used to display the interface of the electronic device; the memory is used to store computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1-10.

12. A computer-readable storage medium, characterized in that, It includes computer instructions, which, when executed by a processor of an electronic device, cause the electronic device to perform the method as described in any one of claims 1-10.