Voice control method and electronic equipment
By introducing preset prompt information and status judgment in voice control electronic devices, the problem of users requiring multiple operations is solved, more efficient voice response is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202311872224.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
AI Technical Summary
In voice control scenarios, users need to operate multiple times in a row to achieve the response of electronic devices, resulting in untimely response and poor user experience.
By introducing preset prompt information into the electronic device, combining the display interface changes and the execution object status of the voice command, we judge whether the voice command is successful or not, and automatically repeat the command when the recognition is accurate, ensuring that the user can realize the intention by one voice input.
It improves the response success rate of voice control, reduces the number of times users need to input continuously, and improves the user experience.
Smart Images

Figure CN120279902A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of speech processing technologies, and in particular, to a speech control method and an electronic device. Background Art
[0002] With the development of technologies, speech control has gradually become a commonly used human-computer interaction method. In particular, when the user's hands are occupied in scenarios such as driving, cooking, or reading, it is very convenient and fast to control an electronic device through speech.
[0003] After the speech control function is started, the user can control the electronic device to execute corresponding instructions by means of speech. The instructions supported by the electronic device for the user to input through speech may include return, go back to the desktop, previous page / next page, volume increase / decrease, and play / pause, etc. Among them, when the electronic device executes some instructions input by the user through speech, it is executed by simulating the user's operations on the electronic device, and these instructions can be called operation vertical instructions.
[0004] However, in some scenarios, some instructions require the user to perform continuous multiple operations before the electronic device can make a response. In this case, it is possible that when the user utters an instruction, the electronic device only simulates one operation of the user, and thus does not respond according to the user's intention. Summary of the Invention
[0005] Embodiments of the present application provide a speech control method and an electronic device, which are used to enable the electronic device to respond according to the user's intention when the user utters an instruction once in different scenarios.
[0006] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, there is provided a speech control method, which is applied to an electronic device. The electronic device includes a microphone, and the method includes:
[0008] The electronic device displays a first interface, and the first interface includes a first switch. In response to the user's operation of turning on the first switch, the voice control function is enabled. The enabling of this voice control function is controlled by the first switch and does not require waking up the electronic device. For example, this voice control function is a visible-and-speak function. After the voice control function is enabled, upon receiving the first voice command collected by the microphone, the electronic device can execute the processing event corresponding to the first voice command. After that, if it is detected that the electronic device emits a preset prompt message, it means that the processing event corresponding to the first voice command has not been successfully executed, or the user needs to operate again. At this time, the electronic device executes the processing event corresponding to the first voice command again. In this way, when the user inputs voice once, the electronic device can respond to the voice to achieve the user's true intention. Thus, when the voice intention recognition is accurate, the possibility that the electronic device cannot achieve the true intention of the voice input by the user is reduced, and the possibility of successful response to the voice input by the user is increased.
[0009] Among them, the preset prompt message includes a hot word matching the first voice command, or the preset prompt message is used to indicate that the execution of the processing event corresponding to the first voice command fails. Enabling the voice control function includes: enabling the recording channel corresponding to the voice control function, and the recording channel is used to obtain the voice collected by the microphone. After the recording channel corresponding to the voice control function is enabled, the voice control function can obtain the voice (audio stream) collected by the microphone through this recording channel.
[0010] In a possible implementation manner of the first aspect, when the processing event corresponding to the first voice command is successfully executed, the execution result of the electronic device includes: the state of the execution object of the processing event corresponding to the first voice command is updated. Further, after executing the processing event corresponding to the first voice command, the above method may further include: determining whether the state of the execution object of the processing event corresponding to the first voice command is updated. In this implementation manner, when the electronic device emits a preset prompt message, executing the processing event corresponding to the first voice command again includes: if the state of the execution object of the processing event corresponding to the first voice command is updated, then when the electronic device emits a preset prompt message, executing the processing event corresponding to the first voice command again.
[0011] In this solution, after the electronic device executes the processing event corresponding to the first voice command, by querying whether the state of the execution object is updated, it is determined whether the processing event corresponding to the first voice command is successfully executed. In this way, by combining the state change of the execution object and whether the electronic device emits a preset prompt message, it is possible to more accurately determine whether the processing event corresponding to the first voice command is successfully executed.
[0012] In a possible implementation of the first aspect, the electronic device includes a display screen. The execution object of the processing event corresponding to the first voice command is the display screen of the electronic device. In this embodiment, the execution object of the processing event corresponding to the first voice command undergoes a status update, which may specifically include: the current display interface of the electronic device is switched. In this solution, by combining whether the display interface is switched and whether a prompt message is sent, it is possible to determine whether the processing event corresponding to the first voice command of the electronic device is successfully executed, thereby improving the judgment accuracy.
[0013] In another possible implementation of the first aspect, the electronic device can specifically determine whether the current display interface is switched in the following manner: obtain the name of the first active component of the electronic device; the name of the first active component includes the application name and the activity name; compare the name of the first active component with the name of the second active component obtained before the electronic device executes the first command; if the name of the first active component is the same as the name of the second active component, then the display page of the electronic device has not been switched.
[0014] In a possible implementation of the first aspect, the execution object of the processing event corresponding to the first voice command is the audio output device of the electronic device. The status update of the execution object of the processing event corresponding to the first voice command includes: the volume of the audio output device changes. In this solution, by combining whether the volume of the audio output device changes and whether the electronic device sends a prompt message, it is possible to determine whether the processing event corresponding to the first voice command of the electronic device is successfully executed, thereby improving the judgment accuracy.
[0015] In a possible implementation of the first aspect, the audio output device of the electronic device can be a speaker, headphones, or other sound playback device (such as a stereo) connected to the electronic device.
[0016] In a possible implementation of the first aspect, the above method further includes: when the first voice command is input, the electronic device displays a second interface. The processing event corresponding to the first voice command includes returning to the previous level from the second interface. After executing the processing event corresponding to the first voice command, the electronic device continues to display the second interface, and the second interface includes preset prompt information, and the preset prompt information includes hot words. In response to the second interface including hot words, the electronic device performs the operation of exiting the second interface again.
[0017] In this solution, the voice commands input by the user are used to control the electronic device to exit the current display interface. For example, the first voice command can be "return" or "exit". In this way, by combining whether the display interface has switched and whether a prompt message is sent, it is possible to determine whether the processing event corresponding to the first voice command of the electronic device is successfully executed, which can improve the judgment accuracy. For some applications, in scenarios where the user needs to continuously execute the return operation multiple times to exit, when the user inputs a voice command once, the electronic device can execute the processing events corresponding to the voice command multiple times, so as to ensure that the electronic device can respond according to the user's intention, that is, exit the application or return to the previous level.
[0018] In a possible implementation manner of the first aspect, the above method may further include: when the first voice command is input, the volume of the audio output device of the electronic device is the first volume. The processing event corresponding to the first voice command includes: increasing the volume of the audio output device. After executing the processing event corresponding to the first voice command, the volume of the audio output device of the electronic device is still the first volume, and the current display interface of the electronic device includes a preset prompt message. In response to the current display interface of the electronic device including the preset prompt message, the electronic device increases the volume of the audio output device again.
[0019] In this solution, the voice commands input by the user are used to control the electronic device to increase the volume. For example, the first voice command can be "increase volume". In this way, by combining whether the volume of the audio output device has changed and whether the electronic device has sent a prompt message, it is possible to determine whether the processing event corresponding to the first voice command of the electronic device is successfully executed, which can improve the judgment accuracy. In some scenarios where the user needs to continuously execute the increase volume operation multiple times to increase the volume, when the user inputs a voice command once, the electronic device can execute the processing events corresponding to the voice command multiple times, so as to ensure that the electronic device can respond according to the user's intention, that is, increase the volume of the audio output device.
[0020] In a possible implementation manner of the first aspect, the electronic device includes a voice control application package APK and a window activity manager AMS; the preset prompt message is a preset text prompt message. After executing the processing event corresponding to the first voice command, the method further includes: the voice control APK obtains the display content of the current display interface of the electronic device from the AMS. When it is determined that the display content includes the preset prompt message, the voice control APK notifies the application corresponding to the first voice command to execute the processing event corresponding to the first voice command again.
[0021] In a possible implementation of the first aspect, the voice control APK includes an interface content acquisition module and an interface parsing module. In this implementation, the above voice control APK obtains the display content of the current display interface of the electronic device from the AMS. Specifically, it may include: the interface content acquisition module sends an acquisition request to the AMS. In response to the acquisition request, the AMS sends the top-level activity of the electronic device to the interface content acquisition module. The interface content acquisition module sends the top-level activity to the interface parsing module. The interface parsing module parses the top-level activity to obtain the display content of the current display interface of the electronic device.
[0022] In a possible implementation of the first aspect, the above method may further include: in response to the user's operation of closing the first switch, the electronic device turns off the voice control function; turning off the voice control function includes: turning off the recording channel corresponding to the voice control function.
[0023] In another possible implementation of the first aspect, turning on the voice control function further includes: displaying a first recording icon and a second recording icon. The first recording icon indicates that the voice control function is turned on, and the second recording icon indicates that the recording channel is turned on. In this way, the user can conveniently obtain the on or off state of the voice control function through the recording icon. The second recording icon indicates that the recording channel is turned on, which is used to prompt the user that the terminal device is in the voice collection state, and can prevent the user's privacy from being leaked.
[0024] In another possible implementation of the first aspect, the first voice matches at least one hot word preset in the electronic device.
[0025] In another possible implementation of the first aspect, the first instruction belongs to an operation instruction. In this embodiment, executing the first instruction corresponding to the first voice may specifically include: simulating the first operation corresponding to the first voice performed by the user on the electronic device.
[0026] In another possible implementation of the first aspect, the first voice is used to control the electronic device to return to the previous level. In this embodiment, simulating the first operation corresponding to the first voice performed by the user on the electronic device may specifically include: querying the navigation mode of the electronic device. When the navigation mode is the three-button navigation, obtaining the position of the return key in the three-button navigation. At the position of the return key, simulating the user's click operation.
[0027] In another possible implementation of the first aspect, the first voice is used to control the electronic device to return to the previous level. In this embodiment, simulating the first operation corresponding to the first voice performed by the user on the electronic device may specifically include: querying the navigation mode of the electronic device. When the navigation mode is gesture navigation, obtaining the preset return gesture in the gesture navigation. Simulating the user to perform the preset return gesture on the display interface of the electronic device.
[0028] In a second aspect, the present application further provides an electronic device. The electronic device may include: a microphone, a display screen, a processor, and a memory. The microphone, the display screen, and the memory are respectively coupled to the processor. The microphone is used to collect language, the display screen is used to display the interface of the electronic device. The memory is used to store computer execution instructions. When the electronic device runs, the processor executes the computer execution instructions stored in the memory to enable the electronic device to execute the voice control method according to any one of the above first aspects.
[0029] In a third aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium includes computer instructions. When the computer instructions are executed by the processor of the electronic device, the electronic device can execute the voice control method according to any one of the above first aspects.
[0030] In a fourth aspect, a computer program product containing instructions is provided. When it runs on an electronic device, the electronic device can execute the voice control method according to any one of the above first aspects.
[0031] In a fifth aspect, a device (for example, the device may be a chip system) is provided. The device includes a processor for supporting the electronic device to implement the functions involved in the above first aspect. In a possible design, the device further includes a memory for storing the necessary program instructions and data of the electronic device. When the device is a chip system, it may be composed of chips or may include chips and other discrete devices.
[0032] Among them, the technical effects brought by any one of the design methods in the second aspect to the fifth aspect can refer to the technical effects brought by different design methods in the first aspect, which will not be elaborated here. Description of the Drawings
[0033] Figure 1 It is a schematic diagram of a scenario example of a voice control method;
[0034] Figure 2A It is a schematic diagram of an example of the opening process of another voice control function;
[0035] Figure 2B It is a schematic diagram of a display page of a navigation mode;
[0036] Figure 2C Schematic illustration of a scenario example of another voice control method Figure 1 ;
[0037] Figure 2D Schematic diagram two of a scenario example of another voice control method;
[0038] Figure 2E Schematic diagram three of a scenario example of another voice control method;
[0039] Figure 2F Schematic illustration of a scenario example of another voice control method Figure 4 ;
[0040] Figure 3A Flow schematic illustration of a voice control method provided by an embodiment of the present application Figure 1 ;
[0041] Figure 3B Flow schematic diagram two of a voice control method provided by an embodiment of the present application;
[0042] Figure 4 Software architecture schematic diagram of an electronic device provided by an embodiment of the present application;
[0043] Figure 5 Interaction schematic illustration of each module of an electronic device when implementing the voice control method provided by an embodiment of the present application Figure 1 ;
[0044] Figure 6 Interaction schematic diagram two of each module of an electronic device when implementing the voice control method provided by an embodiment of the present application;
[0045] Figure 7A Schematic illustration of a scenario example of the voice control method provided by an embodiment of the present application Figure 1 ;
[0046] Figure 7B Schematic diagram two of a scenario example of the voice control method provided by an embodiment of the present application;
[0047] Figure 8 Flow schematic diagram of a voice control method provided by an embodiment of the present application;
[0048] Figure 9 Hardware structure schematic diagram of an electronic device provided by an embodiment of the present application;
[0049] Figure 10 Structure schematic diagram of a chip system provided by an embodiment of the present application. Detailed implementation manners
[0050] To facilitate a clear description of the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:
[0051] The goal of automatic speech recognition (ASR) is to convert the lexical content in the user's speech into computer-readable input, such as keystrokes, binary codes, or character sequences.
[0052] Natural language understanding (NLU) is a general term for all method models or tasks that support machines to understand the content of text.
[0053] Dialogue management (DM) is used to control the process of human-machine dialogue and determine the response to the user at this moment based on the dialogue history information.
[0054] An activity is one of the four major components of the Android system and is a visual interface for user operations; it provides a window for the user to complete operation instructions.
[0055] An intention is an idea of hoping to achieve a certain goal. In the field of voice control, intention recognition is an important technology. By accurately identifying and understanding the user's needs and intentions, more accurate instructions can be executed in response to the user's speech, thereby meeting the user's needs. Taking the user's input speech "query today's weather" as an example, the electronic device can perform intention recognition on this speech, extract the entity content in this speech: "query", "weather", and thus determine that the user's intention is to query the weather. Another example is that the user's input speech is "return to the previous level", and the electronic device can perform intention recognition on this speech, extract the entity content in this speech: "return", "previous level", and thus determine that the user's intention is to display the previous page.
[0056] Voice control function:
[0057] Many electronic devices support the voice control function. The electronic device collects the user's input speech through the microphone, parses and recognizes the speech, and executes the instruction corresponding to the speech, realizing the control of the electronic device by the user through speech.
[0058] Generally speaking, in order to save power consumption of electronic devices and avoid false triggering, the voice control function needs to be turned on before it can be used. For example, the user inputs a preset word (called a wake-up word) to the electronic device through voice to wake up the electronic device. After the electronic device is awakened, it can execute the command corresponding to the voice, that is, the voice control function is turned on. For example, the user can turn on or off the voice control function by turning on or off the preset switch in the human-computer interaction interface of the electronic device.
[0059] In different electronic devices, the voice control function may have different names, such as "voice control", "intelligent voice", "voice assistant", "see and speak", "voice command", "free command", "intelligent AI", etc. The specific implementation of voice control functions with different names may also be different.
[0060] The following is an illustrative introduction to several different implementations of the voice control function.
[0061] Voice Assistant:
[0062] Before the user uses a voice assistant to control an electronic device, the electronic device needs to be woken up first. In one example, the wake-up word input by the user into the electronic device is detected, and the electronic device is awakened. In another example, the electronic device is awakened by detecting the user's long press of the power button. In another example, the breath generated when the user inputs voice to the electronic device is detected, and the electronic device is awakened. Generally speaking, before the electronic device is awakened, the microphone of the electronic device works in a power-saving mode (such as searching for signals at a lower power) to pick up sounds from the surrounding environment. The voice collected by the microphone is only detected for wake-up words at the kernel layer, and the corresponding recording channel of the voice assistant is not started in the system and driver of the electronic device.
[0063] The electronic device is awakened in response to a user operation (for example, receiving a wake-up word input by voice), and the corresponding recording channel of the voice assistant is started in the system and driver. After the electronic device is awakened, the voice (audio stream) collected by the microphone is sent to the voice assistant application for processing through the recording channel corresponding to the voice assistant. In this way, the electronic device can execute the instructions corresponding to the voice, enabling the user to control the electronic device through voice; it can also realize functions such as dialogue with the user.
[0064] Taking the electronic device as a mobile phone 100 as an example, illustratively, Figure 1 FIG. 1 is a schematic diagram showing a scenario example in which a user uses a voice assistant to control a mobile phone 100. Figure 1As shown, the mobile phone 100 displays the desktop interface, and the user inputs the voice "Hello YOYO" to the mobile phone 100. In response to receiving the wake-up word "Hello YOYO", the mobile phone 100 is awakened. Exemplarily, after the mobile phone 100 is awakened, it plays the voice "I'm here" to prompt the user that the mobile phone 100 has been awakened. After the mobile phone 100 is awakened, the user can control the mobile phone 100 by voice. Exemplarily, as Figure 1 shown, the user inputs the voice "Open the video" to the mobile phone, and the mobile phone 100 parses and recognizes the voice input by the user, and executes the instruction corresponding to the voice "Open the video". Exemplarily, in response to receiving the voice "Open the video", the mobile phone 100 starts the video application.
[0065] In some other examples, the electronic device can also be awakened by the voice assistant in response to receiving the user's operation of long pressing the power button.
[0066] In some implementation manners, after the voice assistant of the electronic device is awakened, the user can issue an instruction to the electronic device by inputting voice to the electronic device, and the electronic device executes the instruction corresponding to the voice. After the electronic device executes an instruction, or, within a period of time (such as within 8 seconds) after the voice assistant of the electronic device is awakened, if no instruction is received from the user through voice, the electronic device no longer responds to the voice instruction. For example, the electronic device will close the recording channel corresponding to the voice assistant. The user needs to input the wake-up word to the electronic device again to wake up the electronic device before being able to issue an instruction to the electronic device by inputting voice again. That is to say, after the voice assistant of the electronic device is awakened, it enters a "short voice reception" state and can respond to the instructions issued by the user through voice within a short period of time (such as within 8 seconds).
[0067] In some implementation manners, when the electronic device is connected to the network, it supports entering a continuous conversation scenario after being awakened, and the user can have a continuous conversation with the electronic device. After each broadcast by the electronic device, it will continue to pick up the voice and does not need to be awakened repeatedly. Until the user exits the continuous conversation through an instruction such as "Exit".
[0068] See and speak:
[0069] See and speak is implemented locally by the electronic device and does not require a network connection.
[0070] In some examples, see and speak is controlled by a preset switch. The user can turn on the preset switch to enable the see and speak function, or turn off the preset switch to disable the see and speak function.
[0071] Exemplarily, as Figure 2AAs shown, the user can open the settings function of the mobile phone 100; for example, the user clicks on the application icon of the "Settings" application on the desktop. In response to the user's click operation on the application icon of the "Settings" application, the mobile phone 100 displays the "Settings" interface 101. The "Settings" interface 101 includes a "Smart Voice" option 102, and the "Smart Voice" option 102 is used to set the smart voice function. Exemplarily, in response to the user's click operation on the "Smart Voice" option 102, the mobile phone 100 displays the "Smart Voice" interface 103, and the "Smart Voice" interface 103 includes a "See and Speak" option 104. The user can click on the "See and Speak" option 104 to set the options related to the see and speak function. Exemplarily, referring to Figure 2A , in response to the user's click operation on the "See and Speak" option 104, the mobile phone 100 displays the "See and Speak" interface 105. Optionally, the "See and Speak" interface 105 includes a prompt message 106 for prompting the user about the usage method of the see and speak function. The "See and Speak" interface 105 also includes a "See and Speak" switch 107 (i.e., the above-mentioned preset switch). The user can click on the "See and Speak" switch 107 to turn on or off the "See and Speak" switch. In one example, in response to receiving the user's click operation on the "See and Speak" switch 107, the "See and Speak" switch of the mobile phone 100 is turned on, enabling the see and speak function. Optionally, the "See and Speak" interface 105 displays a prompt message 108 for prompting the user that the see and speak function has been successfully enabled.
[0072] In one implementation, after the see and speak function is enabled, the mobile phone 100 displays a first recording icon, and this first recording icon indicates that the see and speak function has been enabled. Exemplarily, as Figure 2A shown, after the "See and Speak" switch 107 is turned on, the status bar of the page displayed by the mobile phone 100 shows a recording icon 10a, indicating that the see and speak function has been enabled.
[0073] In one scenario, the preset switch corresponding to see and speak on the electronic device is not turned on, and the microphone of the electronic device is not enabled. When the preset switch corresponding to see and speak is turned on, the electronic device starts the microphone and starts the corresponding recording channel for see and speak in the system and the driver. In this way, the voice (audio stream) collected by the microphone can be sent to the see and speak application for processing through the corresponding recording channel for see and speak, and the user can control the electronic device through voice.
[0074] In another scenario, when the corresponding preset switch of the "visible-to-speak" function is not turned on, the microphone of the electronic device operates in a power-saving mode (for example, searching for signals at a lower power) to pick up the surrounding environment sounds. When the corresponding preset switch of the "visible-to-speak" function is turned on, the corresponding recording channel of the "visible-to-speak" function is activated in the system and driver of the electronic device. In this way, the voice (audio stream) collected by the microphone can be sent to the "visible-to-speak" application for processing through the corresponding recording channel of the "visible-to-speak" function, enabling the user to control the electronic device by voice.
[0075] After the "visible-to-speak" function is turned on, the corresponding recording channel of the "visible-to-speak" function is activated in the system and driver of the electronic device, and the electronic device enters a "long recording" state, continuously collecting surrounding sounds. The user can issue commands to the electronic device by voice at any time without the need to enter a wake-up word to wake up the electronic device.
[0076] In one implementation, after any function on the electronic device activates the voice input function of the electronic device (turning on the microphone and the recording channel), the electronic device will send a prompt message to the user to indicate that the electronic device is in a voice collection state. This can prevent the leakage of the user's privacy. For example, after the "visible-to-speak" function is turned on, the electronic device enters a continuous voice collection state, and a second recording icon is displayed on the display page of the electronic device. This second recording icon indicates that the recording channel is open, used to prompt the user that the microphone is collecting voice. Exemplarily, as Figure 2A shown, the status bar on the display page of the mobile phone 100 shows a recording icon 10b, indicating that the recording channel is open.
[0077] After the "visible-to-speak" function is turned on, the electronic device activates the recording channel and continuously collects voice through the microphone. The user can input voice to the electronic device at any time. The electronic device analyzes and recognizes the voice input by the user and executes the command corresponding to the voice.
[0078] The commands that the "visible-to-speak" function supports the user to input by voice can include: system commands, such as swiping left, swiping right, swiping up, returning to the desktop, going back, turning up the volume, turning down the volume; video application commands, such as playing, pausing, stopping, fast-forwarding, rewinding; and e-book playback application commands, such as going to the previous page, going to the next page, going to the table of contents, going to the next chapter.
[0079] The "visible-to-speak" function of the electronic device supports classifying the commands input by the user by voice into multiple vertical categories, and one of the vertical categories is the operation vertical category. For the operation vertical category commands, when the electronic device responds to the voice and executes the command corresponding to the voice, it will simulate the user's operation.
[0080] Hot words:
[0081] An electronic device can set some hotwords. After the "visible and speakable" function is enabled, if the received voice matches at least one of the set hotwords, the electronic device executes the instruction corresponding to the voice.
[0082] The hotwords can be pre-configured on the electronic device, can be obtained according to the content in the display page of the electronic device, can also be input by the user, etc.
[0083] Common mobile phone system navigation modes include two types: gesture navigation and three-button navigation.
[0084] Taking the mobile phone system navigation mode as gesture navigation as an example, users can achieve different functions through different gestures. For example, the mobile phone can return to the previous level in response to the gesture of the user swiping from the left or right side of the screen to the middle of the screen. Another example is that the mobile phone can return to the desktop in response to the gesture of the user swiping up from the bottom of the screen. Another example is that the mobile phone can display the recent tasks in response to the gesture of the user swiping up from the bottom of the screen and pausing.
[0085] Taking the mobile phone system navigation mode as three-button navigation as an example, Figure 2B The schematic diagram of the interface of the three-button navigation of the mobile phone 100 is shown. The mobile phone 100 displays three buttons: the function key 109, the home key 110, and the back key 111. The mobile phone 100 can display the recent tasks in response to the user's click operation on the function key 109. The mobile phone 100 can return to the desktop in response to the user's click operation on the home key 110. The mobile phone 100 can return to the previous level in response to the user's click operation on the back key 111.
[0086] In the process of the mobile phone executing the "return" instruction in the operation vertical instruction in response to the voice, when the mobile phone detects that the user inputs the voice corresponding to the "return" instruction, the mobile phone will simulate the user's return operation on the interface. For example, in the scenario where the mobile phone system navigation mode is gesture navigation, the mobile phone will simulate the gesture of the user swiping from the left or right side of the screen to the middle of the screen in response to the voice "return", so as to execute the operation of returning to the previous level. Another example is that in the scenario where the mobile phone system navigation mode is three-button navigation, the mobile phone will simulate the user's click operation on the back key (such as Figure 2B the back key 111) in response to the voice "return", so as to execute the operation of returning to the previous level.
[0087] After the mobile phone executes the return instruction in response to the voice "return", there is usually an interface switch. As Figure 2C shown, the mobile phone 100 displays the "Smart Voice" interface 112. The user inputs the voice "return" to the mobile phone 100. In response to receiving the voice "return" input by the user, the mobile phone 100 simulates the user's operation, returns to the previous level, and displays the settings interface 113.
[0088] Further, in the scenario of the main interface or the main menu interface of the mobile phone display application, when the mobile phone responds to the voice "return" input by the user and executes the return instruction, it will exit the current application. Please continue to refer to Figure 2C , the mobile phone 100 displays the settings interface 113 (i.e., the main menu interface of the settings application). The user inputs the voice "return" to the mobile phone 100. In response to receiving the voice "return" input by the user, the mobile phone 100 simulates the user operation, returns to the previous level, that is, exits the settings application, returns to the desktop, and displays the desktop interface 114.
[0089] However, some applications have set up an anti-misoperation mechanism. To prevent misoperation, after the user performs a return operation once, these applications will not immediately exit the application, but will display prompt messages such as "Press again to exit the application". As Figure 2D shown, the mobile phone 100 displays the message interface 120 of the short video application (the message interface 120 belongs to the main interface of the short video application). The user inputs the voice "return" to the mobile phone 100. The mobile phone 100 responds to receiving the voice "return" input by the user and simulates the user to perform a return operation. At this time, the mobile phone 100 does not exit the short video application, but still displays the message interface 120. And, the message interface 120 also displays the prompt message 121 of "Press again to exit the application". Usually, the user needs to perform a return operation again within a short time after the prompt message 121 appears to exit the short video application. That is, in this scenario, the mobile phone 100 detects two consecutive return operations before it can exit the application.
[0090] In some other examples, as Figure 2E shown, the mobile phone 100 displays the short video playing interface 122 of the short video application (the short video playing interface 122 belongs to the main interface of the short video application). The user inputs the voice "return" to the mobile phone 100. The mobile phone 100 responds to receiving the voice "return" input by the user and simulates the user to perform a return operation. At this time, the mobile phone 100 also does not exit the short video application, but displays the short video playing interface 123. And, the short video playing interface 123 also displays the prompt message 124 of "Press again to exit the application". At the same time, compared with the short video playing interface 122, the short video playing interface 123 refreshes the currently playing short video, that is, the displayed content is updated.
[0091] It should be noted that the above Figure 2D and Figure 2E examples are based on the gesture navigation mode. The same problem also exists in the three-button navigation mode.
[0092] For Figure 2EIn the scenario shown, after the user inputs the voice "Return", the electronic device fails to respond to the voice to obtain a response result that meets the user's true purpose, that is, to exit the short video application. This easily leads to the problem of the electronic device failing to respond to the voice, thus bringing a poor voice control experience to the user.
[0093] In the scenario where the electronic device outputs audio through the earphone, in order to protect the user, the electronic device usually sets a safe volume threshold. Moreover, when the electronic device responds to the user's operation and turns up the volume above the safe volume threshold, the electronic device will display a prompt message to prompt the user that the volume is relatively high and ask whether the user still wants to continue turning up the volume. In the voice control scenario, when the user inputs the voice "Turn up the volume", the electronic device responds to the voice and executes the "Turn up the volume" instruction, thereby turning up the volume of the electronic device.
[0094] As Figure 2F shown, the mobile phone 100 displays the short video playing interface 125, and the mobile phone 100 can display the prompt message 126 to indicate that the current mobile phone 100 outputs audio through the mobile phone. After the mobile phone 100 responds to the voice "Turn up the volume" input by the user and detects that the current volume has exceeded the safe volume threshold, the mobile phone 100 will not execute the "Turn up the volume" instruction, but will first give a prompt to the user and ask whether the user continues to execute the "Turn up the volume" instruction. In one example, the mobile phone 100 can display the prompt message 127 in the short video playing interface 125 as Figure 2F shown. The prompt message 127 is used to prompt that listening to a high volume for a long time may damage the ears and ask whether to continue turning up the volume.
[0095] Based on this, an embodiment of the present application proposes a voice control method. This method is applied to an electronic device that supports voice input. After the electronic device responds to the voice input by the user and executes the instruction corresponding to the voice, the electronic device detects that the true intention of the user is not realized. Exemplarily, the electronic device can, after executing the instruction corresponding to the voice, detect whether the displayed content includes a preset text to determine whether the true intention of the user is realized. In the case where it is detected that the displayed content includes the preset text, the electronic device can execute the instruction again. Among them, the preset text can be used to instruct the user to perform the operation again, or the preset text is used to indicate that the instruction is not executed successfully. Thus, when the voice intention recognition is accurate, the possibility that the electronic device cannot realize the true intention of the voice input by the user is reduced, and the possibility of successful response to the voice input by the user is improved.
[0096] Exemplarily, the above-mentioned electronic device may be a mobile phone, a tablet computer, a laptop computer, a personal computer (PC), an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a smart home device (such as a smart TV, a smart screen, a large screen, a smart speaker, a smart air conditioner, etc.), a personal digital assistant (PDA), a wearable device (such as a smart watch, a smart bracelet, etc.), a vehicle-mounted device, a virtual reality device, etc. The embodiments of the present application do not make any restrictions on this.
[0097] The voice control method proposed by the embodiments of the present application can be applied to a scenario where a user controls an electronic device to perform corresponding operations through voice input. After the voice control function of the electronic device is turned on, in response to receiving a first voice command collected by the microphone, after executing the processing event corresponding to the first voice command, the electronic device can query whether the electronic device emits a prompt message to determine the execution result of the processing event corresponding to the first voice command. If it is detected that the electronic device emits a preset prompt message, the electronic device can execute the processing event corresponding to the first voice command again. Thus, it is ensured that when the user inputs voice once, and the electronic device accurately recognizes the intention of the voice, it can respond according to the user's intention. Furthermore, the problem that the user needs to say twice continuously to control the electronic device to execute an operation is avoided, and the user experience is improved.
[0098] In some embodiments, to query whether the electronic device emits a prompt message, specifically, it can be by querying whether the display content includes a preset text prompt message. Or, to query whether the electronic device emits a prompt message, it can also be to query whether a preset prompt message in the form of voice is emitted.
[0099] Generally, after an electronic device executes a corresponding processing event in response to a user's voice command, some changes will occur to the electronic device, such as display interface switching, volume increase / decrease, changing from an unlocked state to a locked screen state, and exiting the voice control function, etc. Therefore, in addition to detecting whether a prompt message is issued to determine the execution result of the processing event corresponding to the first voice command, the electronic device can also monitor the state change of the execution object of the first voice command and jointly determine the execution result of the processing event corresponding to the first voice command with whether the electronic device issues a prompt message. In some other embodiments, when the processing event corresponding to the first voice command is successfully executed and the execution result includes a state update of the execution object of the processing event corresponding to the first voice command. In this embodiment, after the electronic device executes the processing event corresponding to the first voice command, the electronic device can also determine whether the execution object of the processing event corresponding to the first voice command has a state update. In one example, when the execution object of the processing event corresponding to the first voice command does not have a state update and the electronic device issues a preset prompt message, the electronic device can execute the processing event corresponding to the first voice command again.
[0100] In another example, if after the electronic device executes the processing event corresponding to the first voice command, the execution object of the processing event corresponding to the first voice command has a state update, it may no longer be necessary to determine whether the electronic device issues a preset prompt message.
[0101] Among them, taking the case where the electronic device switches the display interface when the processing event corresponding to the first voice command is successfully executed as an example, the above-mentioned monitoring of the state change of the execution object of the first voice command may specifically include: monitoring whether the display interface of the display screen is switched.
[0102] In some other embodiments, taking the case where the electronic device increases the volume when the processing event corresponding to the first voice command is successfully executed as an example, the above-mentioned monitoring of the state change of the execution object of the first voice command may specifically include: monitoring whether the volume of the audio output device changes.
[0103] In other embodiments, for the embodiment where the processing event corresponding to the first voice command is other events, reference may be made to the description of the implementation process of the voice control method in the examples listed in this application embodiment.
[0104] Hereinafter, the specific implementation manners of the voice control method proposed in the embodiments of the present application will be described in detail with reference to the accompanying drawings. Figure 3A The flowchart of the voice control method in some embodiments of the present application is shown.
[0105] S200. Display a first interface, and the first interface includes a first switch.
[0106] The first switch is used to turn on or off the voice control function. The user can turn on or off the voice control function through the first switch on the first interface. In some embodiments, the first interface may be Figure 2A the "speak as you see" interface 105 shown; the first switch may be the "speak as you see" switch 107.
[0107] S201. Receive an operation to turn on the first switch.
[0108] S202. In response to the operation of turning on the first switch, enable the voice control function.
[0109] In some embodiments, after enabling the voice control function, the electronic device activates the recording channel corresponding to the voice control function. And, after the voice control function is enabled, the electronic device is in a long recording state, continuously collecting surrounding sounds. The user can issue commands to the electronic device through voice at any time without having to input a wake word to wake up the electronic device.
[0110] In some embodiments, for the electronic device to enable the voice control function, it may specifically include: displaying a first recording icon and a second recording icon, where the first recording icon indicates that the voice control function is enabled, and the second recording icon indicates that the recording channel is enabled. Informing the user that the voice control function has been enabled and the electronic device is recording in the form of recording icons. In this way, the user can quickly obtain the status information of the current electronic device.
[0111] S203. Display a second interface.
[0112] The second interface can be any interface of the electronic device. After the voice control function is enabled, on any display page, the electronic device can receive the voice input by the user and execute the instruction corresponding to the voice.
[0113] The user can input voice to the electronic device through the voice control function on the second interface. Correspondingly, the electronic device receives the voice input by the user, as in S204.
[0114] S204. Receive a first voice.
[0115] S205. In response to the first voice, execute the first instruction corresponding to the first voice.
[0116] In some embodiments, the first voice matches at least one hot word preset in the electronic device. In some embodiments, the first instruction corresponding to the first voice may be recorded as a first voice instruction.
[0117] In some embodiments, the electronic device executes a first instruction, which may specifically include: the electronic device executes a processing event corresponding to the first voice or a processing event corresponding to the first instruction. For example, if the first instruction is a return instruction, the electronic device's execution of the return instruction may specifically include: the electronic device simulating a user's execution of a return operation.
[0118] For instructions corresponding to some voices, when the electronic device executes an instruction corresponding to a voice, it is achieved by simulating a user's operation on the electronic device. In some embodiments, the first instruction belongs to an operation instruction. The above S205 may specifically include: simulating a first operation corresponding to the first voice that the user executes on the electronic device.
[0119] Taking the user inputting the voice "return" as an example, when the electronic device responds to this voice and executes the return instruction corresponding to the voice, the electronic device can achieve it by simulating a user's execution of a return operation. Exemplarily, in the gesture navigation mode, when the electronic device responds to receiving the voice "return", it will simulate a preset return gesture of the user on the screen. Since the gestures set for return on different electronic devices may be different, before the electronic device simulates the preset return gesture of the user on the screen, it may further include: obtaining the preset return gesture. In some examples, the preset return gesture may be a gesture of swiping from the left or right side of the screen towards the middle of the screen.
[0120] In some embodiments, the above S205 may specifically include: querying the navigation mode of the electronic device. When the navigation mode is the gesture navigation mode, obtaining the preset return gesture in the gesture navigation. Simulating the user's execution of the preset return gesture on the display page of the electronic device.
[0121] In some other embodiments, the above S205 may specifically further include: querying the navigation mode of the electronic device. When the navigation mode is the three - key navigation mode, obtaining the position of the return key in the three - key navigation. Then, at the position of the return key, simulating the user's click operation.
[0122] The specific implementation processes of the above simulating the user's execution of the preset return gesture and simulating the user's click operation may refer to the descriptions in related technologies and will not be elaborated in the embodiments of the present application.
[0123] The first instruction may also be other instructions. For example, when the first instruction is to increase the volume, the above S205 may specifically include: calling a preset volume adjustment interface, which is used to increase the volume of the electronic device. In other embodiments, when the first instruction is other instructions, the specific implementation process of the electronic device executing the first instruction may refer to the descriptions in related technologies.
[0124] S206. Obtain the display content.
[0125] The displayed content specifically represents the content included in the display page of the electronic device. The displayed content may include pictures and / or texts on the display page. In some embodiments, the electronic device obtains the displayed content. Specifically, it can obtain the top-level activity from the AMS, and then parse the top-level activity to obtain the displayed content. The top-level activity specifically refers to the topmost activity currently displayed on the electronic device, that is, the activity corresponding to the interface that the user sees on the electronic device. Parsing the top-level activity can obtain the pictures and / or texts on the display page.
[0126] S207. Determine whether the displayed content includes preset text.
[0127] In some embodiments, the preset text is used to instruct the user to perform an operation again, such as Figure 2D the prompt message 121 shown or Figure 2E the prompt message 124 shown. In some embodiments, when the first instruction is a return instruction, the preset text is used to instruct the user to perform an operation again.
[0128] In some embodiments, the first voice matches at least one hotword. In this embodiment, the above-mentioned preset text can also be a hotword that matches the first voice.
[0129] In other embodiments, the preset text is used to indicate that the first instruction has not been executed successfully; such as Figure 2F the prompt message 127 shown. In some embodiments, when the first instruction is to increase the volume, the preset text can specifically be that listening to a high volume for a long time may damage the ears. Do you want to continue increasing the volume?
[0130] The preset text can also be set as keywords, such as any one of "return", "exit", "unsuccessful" or "failed", etc.; in this embodiment, when the displayed content includes any one of the above keywords, it can be determined that the displayed content includes the preset text. In other embodiments, the preset text can also be set as multiple keywords. For example, the preset text is set to include two keywords, "again" and "return"; in this embodiment, when the displayed content includes both the two keywords, "again" and "return", it can be determined that the displayed content includes the preset text.
[0131] After obtaining the displayed content, the electronic device can obtain the text in the displayed content. Then, the electronic device can match the text of the displayed content with the preset text. When the electronic device detects text that matches the preset text, it means that the displayed content includes the preset text.
[0132] If the judgment result of S207 is No, it means that after the electronic device executes the first instruction, the display content does not include the preset text. In this case, the electronic device cannot determine whether the first instruction is executed successfully, so it can stop repeating the execution of the first instruction, as shown in S208.
[0133] S208. End.
[0134] It should be noted that the "end" here means that the electronic device stops executing other instructions or operations in response to the first voice input by the user. This can avoid the problem of incorrect response to the voice input by the user. After S207, if the user continues to input other voices (such as the second voice), the electronic device can still continue to recognize, analyze the intention of the second voice, and execute the instruction corresponding to the second voice.
[0135] If the judgment result of S207 is Yes, it means that after the electronic device executes the first instruction, the display content includes the preset text. In this way, the electronic device can determine that the first instruction is not executed successfully, or the user needs to perform the operation again, that is, the electronic device does not achieve the user's intention when executing the first instruction. In the embodiments of the present application, the electronic device can execute S209.
[0136] S209. Execute the first instruction again.
[0137] When the electronic device executes the first instruction in S205 above, it is in response to receiving the first voice. When executing the first instruction again in S209, it does not require the user to repeat the voice input, but is executed based on the fact that the display content includes the preset text after executing the first instruction. That is to say, in this scenario, the user inputs the first voice once, and the electronic device can execute the first instruction twice continuously. Thus, when the voice intention recognition is accurate, the possibility of the electronic device successfully executing the first instruction in response to the first voice is increased.
[0138] For applications with an anti-misoperation mechanism set, usually after the user performs a return operation, the user needs to perform the return operation again within a short time to successfully return to the previous level. Therefore, the electronic device in S209 above needs to execute the first instruction again within a short time. In some embodiments, in S209 above, the electronic device executes the first instruction again within a preset time after S205. The preset time can be set according to the actual situation. For example, the preset time can be set to 0.5 seconds or 1 second, etc. In this way, it is ensured that the electronic device continuously executes the first instruction twice in response to the voice input by the user, and the user's intention can be achieved.
[0139] In the technical solution provided by the embodiment of the present application, after the electronic device responds to the voice input by the user and executes the instruction corresponding to the voice, if it is detected that the display content of the electronic device includes a preset text, it is determined that the electronic device has not successfully executed the instruction corresponding to the voice, or the user needs to perform an operation again. At this time, the electronic device will execute the instruction corresponding to the voice again. In this way, when the user inputs the voice once, the electronic device can respond to the voice and realize the true intention of the user. Therefore, when the voice intention recognition is accurate, the possibility that the electronic device cannot realize the true intention of the voice input by the user is reduced, and the possibility of successful response to the voice input by the user is increased.
[0140] The execution results obtained after the electronic device successfully executes the first instruction can be divided into two cases. One case is that successfully executing the first instruction will cause the electronic device to switch the display page. For example, the first instruction corresponds to returning, going back to the desktop, opening an "application" (such as opening a video), swiping left / right / up / down, swiping to the bottom, going to the previous page / next page, and opening a directory, etc. After the electronic device executes this instruction, the display page usually switches. And in the other case, after the electronic device successfully executes the first instruction, the electronic device does not switch the display page; for example, the first instruction corresponds to turning up the volume, turning down the volume, etc. After the electronic device executes this first instruction, the display page usually does not switch.
[0141] Taking the case where successfully executing the first instruction by the electronic device will cause the electronic device to switch the display page as an example, after the electronic device executes the first instruction, it can be combined with whether the electronic device switches the display page to judge the execution result of the first instruction. As Figure 3B shown, in some embodiments, after the above S205, the above method further includes S301.
[0142] S301. Determine whether the display page switches.
[0143] In some embodiments, the above S301 may specifically include: obtaining the name of the first active component of the electronic device. Comparing the name of the first active component with the name of the second active component obtained before the electronic device executes the first instruction to determine whether the display page switches.
[0144] In some examples, if the name of the first active component is the same as the name of the second active component, the electronic device has not switched the display page. In other examples, if the name of the first active component is different from the name of the second active component, the electronic device has switched the display page.
[0145] Among them, the activity component name includes the application name and the activity name. If the application name and the activity name are the same before and after the electronic device executes the first instruction, it means that the electronic device has not switched the display page. If the application name and / or the activity name are different before and after the electronic device executes the first instruction, it means that the electronic device has switched the display page.
[0146] It should be noted that after S205, the electronic device may first execute S301, and then execute S206 and S207; or it may first execute S206 and S207, and then execute S301; alternatively, the electronic device may also execute S301, as well as S206 and S207 simultaneously. In the embodiments of the present application, the order of S301, as well as S206 and S207, is not limited.
[0147] The above S301 corresponds to determining whether the state of the execution object (of the display screen) corresponding to the first instruction has been updated (whether the display interface has been switched).
[0148] In some embodiments, when the judgment result of S301 is yes, and / or when the judgment result of S207 is no, the electronic device will no longer make other responses to the first voice, that is, end, such as S208.
[0149] In this embodiment, the above S209 may specifically be executed when the judgment result of S301 is no and the judgment result of S207 is yes. In this way, by combining whether the display page has been switched after the electronic device executes the first instruction and whether the display content includes the preset text, the execution result of the first instruction can be determined, making the judgment of the execution result of the first instruction more accurate. Thus, the problem of inaccurate response to the voice caused by incorrect recognition of the execution result of the first instruction can be avoided.
[0150] In the above embodiments, it is described by taking the electronic device as an example that responds to the first voice and continuously executes the first instruction twice under certain conditions. In actual application scenarios, some operations may require the user to operate continuously multiple times to be successfully executed, such as three times or more. Therefore, in other embodiments, after S209, the electronic device may continue to determine whether the first instruction has been successfully executed, that is, return to make a judgment again through S301 and S207, and when it is determined that the electronic device executes the first instruction again and still fails to be executed successfully or the user needs to perform the operation for the third time, the electronic device may continue to execute the first instruction until a preset threshold is reached, or until the time since the first execution of the first instruction reaches a preset time.
[0151] In some embodiments, after the user turns on the voice control function by the first switch, when the voice control function is not needed, the user can also turn off the voice control function by the first switch. In some embodiments, the electronic device turns off the voice control function in response to the user's operation of turning off the first switch. Wherein, turning off the voice control function includes: turning off the recording channel corresponding to the voice control function.
[0152] In other embodiments, after the user turns on the voice control function by the first switch, when the voice control function is not needed, the user can also turn off the voice control function in a voice control manner. In some embodiments, the electronic device turns off the voice control function in response to the target voice input by the user. Wherein, the target voice is used to instruct the electronic device to turn off the voice control function. Exemplarily, the target voice may specifically be: Exit voice control (function), Turn off voice control (function), or Stop voice control (function), etc. In this way, the voice control function can be quickly and conveniently exited in a voice control manner.
[0153] Next, when implementing the voice control method proposed in the embodiments of the present application, the interaction between the modules of the electronic device will be described in detail. Figure 4 Shows the software architecture of the electronic device in some embodiments.
[0154] In some embodiments, the software system of the electronic device may adopt a layered architecture, an event-driven architecture, a microkernel architecture, or a cloud architecture. The embodiments of the present application take the layered architecture system as an example to exemplarily illustrate the software structure of the electronic device.
[0155] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the system is divided into four layers, from top to bottom are the application layer, the application framework layer, the Android runtime and the system library, and the kernel layer.
[0156] The application layer may include a series of application packages (APKs). For example, applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc. In the embodiments of the present application, the application layer includes a voice control APK for providing the voice control function of the electronic device. The voice control APK includes an interface monitoring module, an interface content acquisition module, an interface parsing module, and an application interaction module, etc. Among them, the interface monitoring module is used to monitor processes such as the startup, exit, and switching of the application interface. The interface content acquisition module is used to acquire the top-level activity. The interface parsing module is used to parse the top-level activity to obtain the content of the current display page (including pictures and / or texts). The application interaction module is used to manage the process of the application executing instructions according to voice.
[0157] The application layer further includes a voice processing engine for parsing, recognizing, and processing voices. Among them, the voice parsing module is used to convert voice into text; the voice parsing module may belong to the ASR engine. The voice recognition module is used to understand and recognize the semantics of the text and determine whether it matches the interface text; the voice recognition module may belong to the NLU engine. The instruction mapping module is used to convert the recognized semantics into machine-executable instructions; the instruction mapping module may belong to the DM engine.
[0158] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0159] As Figure 4 shown, the application framework layer may include a window manager, a content provider, a view system, a resource manager, an activity manager service (AMS), a package manager service (PMS), and a multimode control module, etc.
[0160] AMS is mainly responsible for the startup, switching, scheduling of the four major components in the system and the management and scheduling of application processes, etc. Its responsibilities are similar to the process management and scheduling module in the operating system. When initiating a process startup or component startup, the request will be passed to AMS through the inter-process communication (binder) mechanism, and AMS will then make unified processing.
[0161] The multimode control module is used to manage the execution of voice instructions. For example, sending the instruction to the application to make the application execute the instruction.
[0162] The system library can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing library (e.g., openGL ES), 2D graphics engine (e.g., SGL), etc.
[0163] Android runtime is responsible for the scheduling and management of the Android system. Android runtime includes core libraries and a virtual machine.
[0164] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0165] The kernel layer is the layer between hardware and software. The kernel layer can include display drivers, sensor drivers, microphone drivers, Wi-Fi drivers, etc.
[0166] Combined Figure 4 , Figure 5 shows an interaction schematic diagram of each module of the electronic device when implementing the voice control method provided by the embodiments of the present application.
[0167] AMS can monitor the lifecycle of the activity corresponding to each interface of each application on the electronic device.
[0168] Please refer to Figure 6 , the interface monitoring module registers a perception fence with AMS. In this way, when AMS monitors the lifecycle changes of the activities in each application, it can callback and notify the interface monitoring module. Among them, the notification sent by AMS to the interface monitoring module can include the component name corresponding to the activity, and the component name can include the application package name and the activity name. The interface monitoring module can determine whether a display page switch has occurred according to the component names corresponding to the activities received successively. Specifically, the display page switch can include a display page switch caused by an application switch, such as switching from the display page of application 1 to the display page of application 2. Or, the display page switch can also include a display page switch within an application, such as switching from display page a of application 1 to display page b of application 1; such as switching from the short video playing interface of a short video application to the message interface.
[0169] In some embodiments, the interface monitoring module registering a perception fence with AMS can specifically include: the interface monitoring module registers a perception fence with AMS through the awarenessrequest.builder.addfence method.
[0170] In some examples, the interface monitoring module compares the received component name with the component name received from the AMS last time. If the component names are the same in the two consecutive times, it indicates that the display page has not changed. In some cases, the display content of the same display page may be updated. As Figure 2E shown, compared with the short video playing interface 122, the display content of the short video playing interface 123 has been updated; however, the application names and activity names corresponding to the short video playing interface 123 and the short video playing interface 122 are the same. Therefore, the short video playing interface 123 and the short video playing interface 122 belong to the same display page.
[0171] After the user inputs voice to the electronic device, the microphone of the electronic device collects the voice. The microphone inputs the collected voice to the application interaction module of the voice control APK. The application interaction module is responsible for distributing the received voice; in some examples, the application interaction module can distribute the voice to the voice parsing module in the voice processing engine, and the voice parsing module parses the voice input by the user.
[0172] The voice parsing module parses the voice and can obtain the parsed text corresponding to the voice. Then, the voice parsing module sends the parsed text corresponding to the voice obtained by parsing to the voice recognition module.
[0173] The voice recognition module performs semantic understanding and recognition on the parsed text corresponding to the voice, and obtains the intention corresponding to the voice; that is, the intention of the user, the operation that the user hopes to control the electronic device to perform through the voice. After that, the voice recognition module can send the intention corresponding to the voice to the instruction mapping module.
[0174] The instruction mapping module maps and converts the intention corresponding to the voice into a machine-executable instruction (denoted as instruction a). Then, the instruction mapping module can send instruction a to the corresponding application through the multimode control module to notify the application to execute instruction a.
[0175] In the embodiment of the present application, after the multimode control module sends instruction a to the corresponding application, it can also notify the voice control APK that instruction a has been executed. In this way, the voice control APK can determine the execution result of instruction a by obtaining relevant information. And, in the case where it is determined that instruction a may not have been executed successfully, or the electronic device does not achieve the user intention after executing instruction a, the voice control APK can notify the corresponding application to execute instruction a again through the multimode control module.
[0176] In some embodiments, the multimode control module notifying the corresponding application to execute instruction a may specifically include: the multimode control module calls the following interfaces to notify the corresponding application to execute instruction a:
[0177] InputManagerEx.injectInputEvent(InputManager, MotionEvent, InputManagerEx.getInjectInputEventModeWaitForFinish()). Here, parameter 1 is the system InputManager object; parameter 2 is the click event object; parameter 3 is the execution mode, generally injecting the event and waiting for execution.
[0178] Figure 5 Taking the example that the successful execution of instruction a by the electronic device causes the display page of the electronic device to switch. Since the execution of instruction a by the electronic device causes the display page of the electronic device to switch, therefore, after the electronic device executes instruction a, it is possible to determine the execution result of instruction a by the electronic device in combination with whether the display page has switched. In some embodiments, after the multi-mode control module sends instruction a to the corresponding application, it can also notify the interface monitoring module in the voice control APK that instruction a has been executed.
[0179] After the interface monitoring module receives the notification message from the multi-mode control module (which can be recorded as the first notification message), if it determines that the display page has not switched, it means that the electronic device has not achieved the user's intention. In the embodiments of the present application, at this time, instruction a can be executed again through the multi-mode control module.
[0180] In some embodiments, after determining that the electronic device has executed instruction a and the display page of the electronic device has not switched, the interface monitoring module can notify the corresponding application to execute instruction a again through the multi-mode control module.
[0181] In addition, when the execution of instruction a by the electronic device fails, the electronic device will also issue a prompt message on the interface to prompt the user that the execution has failed or to instruct the user to perform the operation again, such as Figure 2E the prompt message 124 shown. Therefore, in some other embodiments, when it is determined that the display page of the electronic device has not switched after the electronic device executes instruction a, it is also possible to determine the execution result of instruction a by the electronic device in combination with the display content of the electronic device. In some examples, when the electronic device does not switch the display page after executing instruction a and the display content of the electronic device includes the preset text, it means that the execution of instruction a by the electronic device has failed.
[0182] Exemplarily, when it is determined that the display page is not switched after the electronic device executes instruction a, the interface monitoring module may send a notification message (which may be denoted as the second notification message) to the interface content acquisition module. In response to receiving the second notification message from the interface monitoring module, the interface content acquisition module acquires the display content of the electronic device. The interface content acquisition module sends the display content to the interface parsing module, and the interface parsing module parses the text displayed on the electronic device. After obtaining the text displayed on the electronic device, the interface parsing module can determine whether the text displayed on the electronic device contains a preset text, so as to determine the execution result of instruction a by the electronic device.
[0183] Among them, the preset text can be set and stored in the electronic device in advance. After obtaining the text displayed on the electronic device, the electronic device can match each text with the preset text one by one to determine whether it contains the preset text. In some embodiments, the preset text (which may be denoted as the first preset text) can be used to instruct the user to perform an operation again; such as Figure 2E the prompt message 124 shown. In other embodiments, the first preset text may also only include keywords, such as: "return", "exit", etc.
[0184] Specifically, when it is detected that after executing instruction a, the display page of the electronic device does not switch and the displayed text contains the first preset text, it can be determined that instruction a has not been executed successfully, or the electronic device does not achieve the user's intention after executing instruction a.
[0185] In some examples, for the interface content acquisition module to acquire the display content of the electronic device, it may specifically include: the interface content acquisition module sends a request to the AMS. In response to this request, the AMS returns the top-level activity of the electronic device to the interface content acquisition module.
[0186] In some embodiments, for the interface content acquisition module to send a request to the AMS, it may specifically include: the interface content acquisition module sends a request to the AMS through activitymanagerEx.requestContentNode.
[0187] Taking the user's input voice as "return" as an example, the speech recognition module performs semantic understanding and recognition on the text corresponding to the voice, and can obtain the intention corresponding to the voice as the return intention. At this time, the speech recognition module can send the return intention to the instruction mapping module. After receiving the return intention, the instruction mapping module can map the return intention to a return instruction (denoted as return instruction 1). Then, the instruction mapping module sends return instruction 1 to the multimode control module, so that the multimode control module notifies the corresponding application to execute the return instruction. Further, the multimode control module sends a first notification message to the interface monitoring module, for notifying the interface monitoring module that return instruction 1 has been executed. The interface monitoring module can monitor whether the display page of the electronic device has switched from the AMS according to the sensing fence.
[0188] After the interface monitoring module determines that it has received the first notification message and the display page of the electronic device has not switched, the interface monitoring module can send a second notification message to the interface content acquisition module. The second notification message is used to notify the interface content acquisition module to acquire the display content of the electronic device. In response to the second notification message, the interface content acquisition module can send a request to the AMS. In response to the request, the AMS returns the current top-level activity of the electronic device to the interface content acquisition module. The interface content acquisition module sends the top-level activity to the interface parsing module. The interface parsing module parses the top-level activity to obtain the display content of the electronic device and the text included therein.
[0189] Then, the interface parsing module analyzes whether the text displayed on the electronic device includes a first preset text. Since the user's input voice is return, when it is determined that the text displayed on the electronic device includes the first preset text, the interface parsing module can notify the corresponding application to execute the return instruction again through the multimode control module (which can be denoted as return instruction 2). And the multimode control module can execute return instruction 2 for the corresponding application. In this way, for an application with an anti-misoperation mechanism set, when the electronic device responds to a single "return" voice input by the user, it can execute two return instructions, so as to ensure that the electronic device can achieve the user's intention, that is, return to the previous level.
[0190] Such as Figure 7AAs shown, the mobile phone 100 displays the message interface 601 of the short video application, and the user inputs the voice "Return" to the mobile phone 100. In the embodiment of the present application, in response to the voice "Return", after the mobile phone 100 executes a return instruction, it displays the message interface 602, and the message interface 602 includes a prompt message 603. The voice control APK determines that the text of the display content of the mobile phone 100 includes a first preset text through the interface content acquisition module and the interface parsing module, and the interface parsing module notifies the corresponding application to execute the return instruction again through the multimode control module. After the corresponding application executes the return instruction for the second time, the mobile phone 100 can exit the short video application. In an example, after the mobile phone 100 exits the short video application, it can display the desktop interface 604. Thus, when the user says the voice "Return" once, the mobile phone 100 can realize the user's true intention, exit the short video application, and return to the previous level, that is, return to display the desktop.
[0191] In another example, as Figure 7B shown, the mobile phone 100 displays the short video playback interface 605 of the short video application, and the user inputs the voice "Return" to the mobile phone 100. In the embodiment of the present application, in response to the voice "Return", after the mobile phone 100 executes a return instruction, it displays the short video playback interface 606, and the short video playback interface 606 includes a prompt message 607. Moreover, the short video playback interface 606 and the short video playback interface 605 have the same application name and activity name, so they belong to the same activity. Thus, the interface monitoring module can determine that there is no display page switching. Then, the voice control APK determines that the text of the display content of the mobile phone 100 includes a first preset text through the interface content acquisition module and the interface parsing module, and the interface parsing module notifies the corresponding application to execute the return instruction again through the multimode control module. After the corresponding application executes the return instruction for the second time, the mobile phone 100 can exit the short video application. In an example, after the mobile phone 100 exits the short video application, it can display the desktop interface 608. Thus, when the user says the voice "Return" once, the mobile phone 100 can realize the user's true intention, exit the short video application, and return to the previous level, that is, return to display the desktop.
[0192] In addition, for applications with an anti-misoperation mechanism set, usually after the user executes a return operation, the user needs to execute the return operation again within a short period of time to successfully return to the previous level. Therefore, in some embodiments, when the interface parsing module determines that the display content of the electronic device includes a preset text and notifies the corresponding application to execute the second return instruction through the multimode control module, the application needs to execute the second return instruction within a preset time after the first execution of the return instruction.
[0193] In some other embodiments, the electronic device does not switch the display page after executing instruction a. When the electronic device fails to execute successfully, a prompt message is usually issued. This prompt message can be used to ask the user whether to execute instruction a, such as Figure 2F prompt message 127 shown. Or in some other examples, this prompt message can be used to prompt the user that instruction a has failed. In this embodiment, after the electronic device executes instruction a, it can determine the execution result of instruction a by judging whether the displayed content contains a preset text (which can be denoted as the third preset text). Among them, the third preset text can be set according to the actual situation. In the example, the third preset text can be set to, for example, Figure 2F prompt message 127 shown, or the third preset text can also be set to keywords such as "failure", "unsuccessful", etc.
[0194] In this embodiment, since the electronic device does not switch the display page when instruction a is executed successfully. Therefore, after the electronic device responds to the voice input by the user and for the first time notifies the corresponding application to execute instruction a through the multimode control module, the multimode control module can directly notify the interface content acquisition module that instruction a has been executed. In response to receiving the notification message from the multimode control module, the interface content acquisition module can send a request to the AMS. In response to this request, the AMS returns the top-level activity to the interface content acquisition module. The interface content acquisition module sends the top-level activity to the interface parsing module, and the interface parsing module parses the top-level activity to obtain the display content of the electronic device and the text included therein. Then, the interface parsing module determines whether the text displayed on the electronic device includes the third preset text. And, when it is determined that the text displayed on the electronic device includes the third preset text, the corresponding application is notified through the multimode control module to execute instruction a again.
[0195] Taking the scenario where the electronic device outputs audio through the earphone as an example, the user inputs the voice "turn up the volume". The electronic device responds to receiving the voice "turn up the volume" and executes the instruction to turn up the volume. At the same time, if it is detected that if the instruction to turn up the volume is executed, the volume of the electronic device will exceed the safe volume threshold, the electronic device will issue a prompt message to ask the user whether to continue to execute the instruction to turn up the volume. If the electronic device detects that the displayed content includes the third preset text, the electronic device will again notify the corresponding application to execute the instruction to turn up the volume through the multimode control module.
[0196] Generally, when a user performs the same operation twice in a relatively short period of time, it can indicate confirmation of the execution of that operation. In some embodiments, when the interface parsing module determines that the display content of the electronic device includes a preset text and notifies the corresponding application to execute the instruction to increase the volume again through the multimode control module, the application needs to execute the second instruction to increase the volume within a preset time after the first execution of the instruction to increase the volume.
[0197] In addition, if the execution of instruction a fails, the execution object corresponding to instruction a will not have its status updated either. Taking instruction a as increasing the volume as an example, the execution object corresponding to instruction a can be the audio output device of the electronic device. If the execution of instruction a fails, the volume of the corresponding audio output device does not change.
[0198] In the technical solution provided by the embodiments of the present application, after the electronic device responds to the voice input by the user and executes the instruction corresponding to the voice, if it is detected that the display content of the electronic device includes a preset text, it is determined that the electronic device has not successfully executed the instruction corresponding to the voice, or the user needs to perform the operation again. At this time, the electronic device will execute the instruction corresponding to the voice again. In this way, when the voice intention is accurately recognized, the possibility that the electronic device cannot realize the true intention of the voice input by the user can be reduced, and the possibility of successful response to the voice input by the user can be improved.
[0199] Figure 8 Shows the timing interaction diagram among the modules during the process of the voice control method proposed by the embodiments of the present application. In this embodiment, the voice input by the user is "return" as an example for illustration. The specific implementation manners of other instructions corresponding to the voice input by the user in other embodiments can be referred to Figure 8 the process shown and will not be elaborated in the embodiments of the present application.
[0200] The user turns on the first switch. The electronic device responds to the operation of the user turning on the first switch and enables the voice control function. Enabling the voice control function may specifically include: enabling the recording channel corresponding to the voice control function, and this recording channel is used to obtain the voice collected by the microphone. In some other embodiments, enabling the voice control function may further include: displaying a first recording icon and a second recording icon, where the first recording icon indicates that the voice control function is enabled, and the second recording icon indicates that the recording channel is enabled.
[0201] The interface monitoring module registers the perception fence with the AMS. It should be noted that the interface monitoring module can register the perception fence with the AMS when the voice control function of the electronic device is turned on. When the AMS perceives changes in the application's life cycle, it can make a callback through the perception fence to notify the interface monitoring module. Exemplarily, when the AMS makes a callback to notify the interface monitoring module, it notifies the interface monitoring module of the component name corresponding to the activity, and the component name can include the application package name and the activity name.
[0202] The user inputs the voice "Return" to the electronic device. Specifically, the microphone collects the voice input by the user. Since the current voice control function is turned on and the recording channel corresponding to the voice control function is turned on, the microphone sends the voice to the application interaction module of the voice control APK.
[0203] The application interaction module distributes the received voice to the voice processing engine for processing. Specifically, the application interaction module sends the voice to the voice parsing module in the voice processing engine. After receiving the voice, the voice parsing module can parse the voice to obtain the parsed text. Then, the voice parsing module sends the parsed text to the voice recognition module for intent recognition. The voice recognition module recognizes the intent of the received parsed text and obtains the intent corresponding to the parsed text. The voice input by the user is "Return", and the voice recognition module recognizes the intent as the return intent. After that, the voice recognition module can send the return intent to the instruction mapping module for instruction mapping. After receiving the return intent, the instruction mapping module maps the return intent to an instruction that the electronic device can execute, that is, the return instruction.
[0204] The instruction mapping module sends the mapped return instruction (denoted as return instruction 1) to the multi-mode control module. After receiving return instruction 1, the multi-mode control module notifies the corresponding application to execute return instruction 1. Thus, the response to the user's input voice "Return" is completed. At the same time, the multi-mode control module can also notify the interface parsing module that return instruction 1 has been executed.
[0205] The interface monitoring module can determine whether the display page has switched according to the activity name received from the AMS successively. The switching of the display page can specifically be the switching of the display page between applications or the switching of the display page within an application. When the activity name received by the interface monitoring module from the AMS is the same as the activity name received last time, it can be determined that the display page has not switched.
[0206] When the electronic device successfully executes the return instruction, it usually causes the display page of the electronic device to switch. Moreover, if the electronic device fails to execute the return instruction successfully, the electronic device will display a prompt message containing preset text on the display page to instruct the user to perform the operation again or indicate that the return instruction has not been executed successfully. Therefore, in the embodiments of the present application, after the electronic device executes return instruction 1, the interface monitoring module determines whether the user intention is achieved for return instruction 1 of the electronic device by combining whether the display page switches.
[0207] When the interface monitoring module determines that the display page has switched, it may no longer make other responses to the first voice, that is, end.
[0208] When the interface monitoring module determines that the display page has not switched, it sends a notification message to the interface content acquisition module. After receiving the notification message from the interface monitoring module, the interface content acquisition module sends a request to the AMS. In response to this request, the AMS returns the top-level activity to the interface content acquisition module.
[0209] After receiving the top-level activity from the AMS, the interface content acquisition module sends the top-level activity to the interface parsing module. The interface parsing module parses the top-level activity to obtain the text of the current display page. Then the interface parsing module determines whether the text includes "return" or "exit".
[0210] When the interface parsing module determines that the text includes "return" or "exit", it sends a return instruction (denoted as return instruction 2) to the multimode control module. The multimode control module notifies the corresponding application to execute return instruction 2.
[0211] Figure 8 The specific implementation manners of the steps in the shown process can be referred to the previous description and will not be elaborated here.
[0212] In the technical solution provided by the embodiments of the present application, after the electronic device responds to the voice input by the user and executes the instruction corresponding to the voice, if it is detected that the display page of the electronic device has not switched and the display content includes the preset text, it is determined that the electronic device has not successfully executed the instruction corresponding to the voice or the user needs to perform the operation again. At this time, the electronic device will execute the instruction corresponding to the voice again. In this way, when the user inputs the voice once, the electronic device can respond to the voice to achieve the real intention of the user. Therefore, when the voice intention recognition is accurate, the possibility that the electronic device cannot achieve the real intention of the voice input by the user is reduced, and the possibility of successful response to the voice input by the user is increased.
[0213] Next, the hardware structure of the electronic device to which the method proposed in the embodiments of the present application is applied will be introduced.
[0214] As Figure 9 shown is a schematic structural diagram of an electronic device 500 provided by an embodiment of the present application. The electronic device 500 may include a processor 510, an external memory interface 520, an internal memory 521, a universal serial bus (USB) interface 530, a charging management module 540, a power management module 541, a battery 542, an antenna 1, an antenna 2, a mobile communication module 550, a wireless communication module 560, an audio module 570, a speaker 570A, a receiver 570B, a microphone 570C, a headphone jack 570D, a sensor module 580, a button 590, a motor 591, a camera 592, a display screen 593, and a subscriber identification module (SIM) card interface 594, etc. Among them, the sensor module 580 may include a pressure sensor 580A, a touch sensor 580B, etc.
[0215] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 500. In other embodiments of the present application, the electronic device 500 may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0216] The processor 510 may include one or more processing units. For example, the processor 510 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. For example, the processor 510 is used to execute the voice control method in the embodiments of the present application.
[0217] Among them, the controller may be the nerve center and command center of the electronic device 500. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.
[0218] A memory can also be set in the processor 510 for storing instructions and data. In some embodiments, the memory in the processor 510 is a cache memory. This memory can save the instructions or data that the processor 510 has just used or recycled. If the processor 510 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 510, and thus improves the efficiency of the system.
[0219] The USB interface 530 is an interface that conforms to the USB standard specification. Specifically, it can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 530 can be used to connect a charger to charge the electronic device 500, and can also be used to transfer data between the electronic device 500 and peripheral devices.
[0220] The external memory interface 520 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 500. The external memory card communicates with the processor 510 through the external memory interface 520 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0221] The internal memory 521 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 510 executes various functional applications and data processing of the electronic device 500 by running the instructions stored in the internal memory 521. The internal memory 521 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function (such as a sound playback function, an image playback function, etc.).
[0222] In addition, the internal memory 521 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0223] The charge management module 540 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charge management module 540 can receive the charging input of the wired charger through the USB interface 530.
[0224] The power management module 541 is used to connect the battery 542, the charge management module 540, and the processor 510. The power management module 541 receives the inputs of the battery 542 and / or the charge management module 540 to supply power to the processor 510, the internal memory 521, the external memory, the display screen 593, the camera 592, and the wireless communication module 560, etc.
[0225] In some other embodiments, the power management module 541 may also be disposed in the processor 510. In some other embodiments, the power management module 541 and the charging management module 540 may also be disposed in the same device.
[0226] The wireless communication function of the electronic device 500 may be implemented by the antenna 1, the antenna 2, the mobile communication module 550, the wireless communication module 560, the modulation and demodulation processor, and the baseband processor, etc.
[0227] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 500 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: The antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0228] The mobile communication module 550 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 500. The mobile communication module 550 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 550 can receive electromagnetic waves through the antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 550 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves through the antenna 1 and radiate it out.
[0229] The wireless communication module 560 can provide solutions for wireless communications such as wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied to the electronic device 500. The wireless communication module 560 may be one or more devices integrating at least one communication processing module. The wireless communication module 560 receives electromagnetic waves through the antenna 2, frequency-modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 510. The wireless communication module 560 can also receive the signals to be transmitted from the processor 510, frequency-modulate and amplify them, and convert them into electromagnetic waves through the antenna 2 and radiate them out.
[0230] In some embodiments, antenna 1 of electronic device 500 is coupled to mobile communication module 550, and antenna 2 is coupled to wireless communication module 560, such that electronic device 500 can communicate with a network and other devices via wireless communication technologies.
[0231] Electronic device 100 can implement audio functions via audio module 570, speaker 570A, receiver 570B, microphone 570C, headphone jack 570D, and an application processor, etc. For example, music playback, recording, etc.
[0232] Audio module 570 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. Audio module 570 can also be used for encoding and decoding audio signals. In some embodiments, audio module 570 can be disposed in processor 510, or some functional modules of audio module 570 can be disposed in processor 510. Speaker 570A, also referred to as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. Receiver 570B, also referred to as an "earpiece", is used to convert an audio electrical signal into a sound signal. Microphone 570C, also referred to as a "microphone", "transmitter", is used to convert a sound signal into an electrical signal. Headphone jack 570D is used to connect a wired headphone. Headphone jack 570D can be a USB interface 530, or can be a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface. In the embodiments of the present application, after the voice control function is turned on, electronic device 100 continuously collects surrounding sounds via microphone 570C.
[0233] Pressure sensor 580A is used to sense a pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, pressure sensor 580A can be disposed in display screen 593. There are many types of pressure sensors 580A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. A capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on pressure sensor 580A, the capacitance between the electrodes changes. Electronic device 500 determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on display screen 593, electronic device 500 detects the intensity of the touch operation based on pressure sensor 580A. Electronic device 500 can also calculate the position of the touch based on the detection signal of pressure sensor 580A.
[0234] The touch sensor 580B, also known as the "touch panel". The touch sensor 580B can be disposed on the display screen 593, and the touch sensor 580B and the display screen 593 form a touch screen, also known as the "touch display screen". The touch sensor 580B is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 593. In some other embodiments, the touch sensor 580B can also be disposed on the surface of the electronic device 500, at a different position from where the display screen 593 is located.
[0235] The keys 590 include a power-on key, volume keys, etc. The keys 590 can be mechanical keys or touch keys. The electronic device 500 can receive key inputs and generate key signal inputs related to the user settings and function control of the electronic device 500.
[0236] The motor 591 can generate vibration prompts. The motor 591 can be used for incoming call vibration prompts or touch vibration feedback.
[0237] The camera 592 is used to capture still images or videos. In some embodiments, the electronic device 500 can include one or N cameras 592, where N is a positive integer greater than 1.
[0238] The electronic device 500 realizes the display function through the GPU, the display screen 593, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 593 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 510 can include one or more GPUs, which execute program instructions to generate or change the display information.
[0239] The display screen 593 is used to display images, videos, etc. In some embodiments, the electronic device 500 can include one or N display screens 593, where N is a positive integer greater than 1.
[0240] The SIM card interface 594 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 594 to achieve contact and separation with the electronic device 500. The electronic device 500 can support one or N SIM card interfaces, where N is a positive integer greater than 1.
[0241] The voice control methods introduced in the above embodiments can all be executed in the electronic device with the above hardware structure.
[0242] Some other embodiments of the present application provide an electronic device (such as mobile phone 100). The electronic device may include: a microphone, a display screen, a processor, and a memory. The microphone, the display screen, and the memory are respectively coupled to the processor. The microphone is used to collect voices; the display screen is used to display the interface of the electronic device; the memory is coupled to the processor. The memory is further used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device can perform each function or step that the mobile phone 100 performs in the above method embodiments. The structure of the electronic device may refer to Figure 9 the structure of the electronic device 500 shown.
[0243] Embodiments of the present application further provide a chip system, as Figure 10 shown, the chip system 1000 includes at least one processor 1001 and at least one interface circuit 1002. The processor 1001 and the interface circuit 1002 can be interconnected through a line. For example, the interface circuit 1002 can be used to receive signals from other devices (such as the memory of a computer). For another example, the interface circuit 1002 can be used to send signals to other devices (such as the processor 1001). Exemplarily, the interface circuit 1002 can read the instructions stored in the memory and send the instructions to the processor 1001. When the instructions are executed by the processor 1001, the computer can perform each step in the above embodiments. Of course, the chip system can also include other discrete devices, and the embodiments of the present application do not make specific limitations thereto.
[0244] Embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium includes computer instructions. When the computer instructions run on the above electronic device (such as mobile phone 100), the electronic device is enabled to perform each function or step that the mobile phone 100 performs in the above method embodiments.
[0245] Embodiments of the present application further provide a computer program product. When the computer program product runs on a computer, the computer is enabled to perform each function or step that the mobile phone 100 performs in the above method embodiments. Among them, the computer can be an electronic device, such as mobile phone 100.
[0246] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0247] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0248] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0249] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0250] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes.
[0251] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A voice control method, characterized in that, Applied to an electronic device, the electronic device including a microphone, the method comprising: Displaying a first interface, the first interface including a first switch; In response to a user's operation of turning on the first switch, enabling a voice control function; the enabling of the voice control function includes: enabling a recording channel corresponding to the voice control function, the recording channel being used to acquire voices collected by the microphone; In response to a first voice command collected by the microphone, executing a processing event corresponding to the first voice command; In the case where the electronic device issues a preset prompt message, executing again the processing event corresponding to the first voice command; the preset prompt message includes a hot word matching the first voice command, or the preset prompt message is used to indicate that the execution of the processing event corresponding to the first voice command fails.
2. The method according to claim 1, wherein In the case where the execution of the processing event corresponding to the first voice command is successful, the execution result of the electronic device includes: the state of the execution object of the processing event corresponding to the first voice command is updated; After executing the processing event corresponding to the first voice command, the method further includes: Judging whether the state of the execution object of the processing event corresponding to the first voice command is updated; The executing again the processing event corresponding to the first voice command in the case where the electronic device issues a preset prompt message includes: If the state of the execution object of the processing event corresponding to the first voice command is not updated, judging whether the electronic device issues a preset prompt message; In the case where the electronic device issues a preset prompt message, executing again the processing event corresponding to the first voice command.
3. The method according to claim 2, wherein The electronic device includes a display screen; the execution object of the processing event corresponding to the first voice command is the display screen of the electronic device; The state of the execution object of the processing event corresponding to the first voice command is updated, including: the current display interface of the electronic device is switched.
4. The method according to claim 2, wherein The execution object of the processing event corresponding to the first voice command is the audio output device of the electronic device; the state of the execution object of the processing event corresponding to the first voice command is updated, including: the volume of the audio output device changes.
5. The method according to any one of claims 1-4, characterized in that The method further includes: When the first voice command is input, the electronic device displays a second interface; the processing event corresponding to the first voice command includes returning to the previous level from the second interface; After executing the processing event corresponding to the first voice command, the electronic device continues to display the second interface, the second interface including the preset prompt message, the preset prompt message including the hot word; In response to the second interface including the hot word, the electronic device executes again the operation of exiting the second interface.
6. The method according to any one of claims 1-4, characterized in that, The method further includes: When the first voice command is input, the volume of the audio output device of the electronic device is a first volume; the processing event corresponding to the first voice command includes: increasing the volume of the audio output device; After executing the processing event corresponding to the first voice command, the volume of the audio output device of the electronic device remains at the first volume, and the current display interface of the electronic device includes the preset prompt information; In response to the current display interface of the electronic device including the preset prompt information, the electronic device executes again to increase the volume of the audio output device.
7. The method according to any one of claims 1-6, characterized in that, The electronic device includes a voice control application package APK and a window activity manager AMS; The preset prompt information is preset text prompt information; After executing the processing event corresponding to the first voice command, the method further includes: The voice control APK obtains the display content of the current display interface of the electronic device from the AMS; When it is determined that the display content includes the preset prompt information, the voice control APK notifies the application corresponding to the first voice command to execute again the processing event corresponding to the first voice command.
8. The method according to claim 7, wherein The voice control APK includes an interface content acquisition module and an interface parsing module; The voice control APK obtains the display content of the current display interface of the electronic device from the AMS, including: The interface content acquisition module sends an acquisition request to the AMS; In response to the acquisition request, the AMS sends the top-level activity of the electronic device to the interface content acquisition module; The interface content acquisition module sends the top-level activity to the interface parsing module; The interface parsing module parses the top-level activity to obtain the display content of the current display interface of the electronic device.
9. The method according to any one of claims 1-8, characterized in that The method further includes: In response to the user's operation of closing the first switch, the electronic device turns off the voice control function; turning off the voice control function includes: turning off the recording channel corresponding to the voice control function.
10. An electronic device, characterized in that, The electronic device includes: a processor, a memory, a microphone, and a display screen; the memory and the display screen are respectively coupled to the processor; The microphone is used to collect voices; the display screen is used to display the interface of the electronic device; the memory stores computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the method according to any one of claims 1-9.
11. A computer-readable storage medium, characterized in that, Including computer instructions, when the computer instructions run on an electronic device, the electronic device executes the method according to any one of claims 1-9.
Citation Information
Patent Citations
Speech control method, device and terminal equipment
CN105957530A
Voice control method and electronic equipment
CN109584879A
Voice control method and device of terminal, storage medium and terminal
CN110865755A
Voice control method and electronic equipment
CN113794800A
Display device, server and wakeup-free voice control method
CN115396709A