A method of controlling a terminal device by voice and a terminal device
By pre-setting the on/off control voice function and setting multiple levels of hot words in the terminal device, the problems of cumbersome operation and limited scope of voice control on the terminal device are solved, realizing convenient voice control without frequent wake-up and more efficient voice response.
Patent Information
- Application Number
- CN202310809133.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2043-06-30
AI Technical Summary
The voice control function of existing terminal devices requires frequent wake-up and de-wake-up, which is cumbersome. Furthermore, the scope of voice control is limited by the text displayed on the interface, resulting in a poor user experience.
The voice control function can be turned on and off by a preset switch. Hot words are pre-set in the terminal device, including system-level, scene-level and interface-level hot words, which support a wide range of voice commands. The cached hot words are updated when the interface is switched to improve response speed and save power.
It enables voice control anytime without frequently waking up the terminal device, supports more comprehensive and richer voice commands, improves user experience and response speed, and reduces power consumption.
Smart Images

Figure CN119229861B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of terminals, and in particular to a method for controlling a terminal device by voice and a terminal device. BACKGROUND
[0002] With the development of terminal technology, terminal devices are increasingly used in people's daily life. In order to improve the operation efficiency of users using terminal devices, terminal devices generally provide multiple human-computer interaction modes such as human-computer interaction interface, voice control, gesture control, etc. Among them, voice control is a main and fast human-computer interaction mode. Especially when the user is driving, cooking, reading, etc. in a scenario where both hands are occupied, it is very convenient and fast to control the terminal device by voice.
[0003] Due to the needs of user information security, terminal device power consumption, etc., the voice control function needs to be turned on before it can be used. In some designs, the voice control function is turned on by a wake-up word. The user inputs the wake-up word to the terminal device by voice, and after the terminal device is woken up, it can execute corresponding operations according to the user voice. After the user stops voice input (such as pausing for a few seconds), the voice control function is automatically turned off. The user needs to wake up the terminal device by voice input of the wake-up word before using the voice control function each time. If the terminal device needs to be controlled by voice for a long time, the terminal device needs to be frequently woken up, which is relatively cumbersome to operate.
[0004] In some other designs, the voice control function is turned on or off by setting a switch. The user can turn on the voice control function by opening a preset switch on the terminal device, or turn off the voice control function by closing the preset switch on the terminal device. During the opening of the switch, the terminal device can execute corresponding instructions according to the user input voice, and the user can conveniently control the terminal device by voice. At present, the voice range supported by the terminal device in this scenario is relatively small, such as only supporting voice matching the text on the display interface; the user's experience of using the voice control function is poor. SUMMARY
[0005] The embodiments of the present application provide a method for controlling a terminal device by voice and a terminal device, which do not require the user to input a wake-up word. After the preset switch is opened, the user can interact with the terminal device by voice at any time, the terminal device supports more comprehensive voice, and the user's experience of controlling the terminal device by voice is improved.
[0006] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, a method for controlling a terminal device by voice is provided. The method is applied to the terminal device, and the terminal device comprises a microphone. The method comprises: displaying, by the terminal device, a first interface, the first interface comprising a first switch; in response to an operation of the user opening the first switch, starting, by the terminal device, a voice control function. In response to an operation of the user closing the first switch, stopping, by the terminal device, the voice control function. The starting or stopping of the voice control function is controlled by the first switch, without the need to wake up the terminal device. For example, the voice control function is a see-to-say function. After the voice control function is started, a first voice input by the user when the terminal device displays a second interface is received, and the terminal device executes a first instruction corresponding to the first voice; wherein the second interface does not comprise a first text corresponding to the first instruction. That is, the voice instructions supported by the terminal device are not limited to the texts in the interface. The terminal device can support more, more comprehensive and more rich voice to control the terminal device.
[0008] In the method, starting the voice control function comprises starting a recording channel corresponding to the voice control function, and the recording channel is used to acquire the voice collected by the microphone. Stopping the voice control function comprises stopping the recording channel corresponding to the voice control function. After the recording channel corresponding to the voice control function is started, the voice control function can acquire the voice (audio stream) collected by the microphone through the recording channel.
[0009] In combination with the first aspect, in a possible implementation, starting, by the terminal device, the voice control function comprises: displaying, by the terminal device, a first recording icon and a second recording icon. The first recording icon indicates that the voice control function is started, so that the user can conveniently acquire the state of starting or stopping the voice control function through the recording icon. The second recording icon indicates that the recording channel is started, and is used to prompt the user that the terminal device is in a state of collecting voice, so as to avoid the privacy of the user being leaked.
[0010] In combination with the first aspect, in a possible implementation, the first voice matches at least one hot word set in advance.
[0011] In the method, some hot words are set in the terminal device in advance, and the hot words are not acquired from the display interface. If the voice input by the user matches any hot word, the instruction corresponding to the voice can be executed. Since the hot words are not acquired from the display interface, the voice instructions supported by the terminal device are not limited to the texts in the interface. In this way, the terminal device can support more comprehensive and more rich voice instructions.
[0012] In some implementations, the hot words set in advance comprise at least one of the following: left swipe, right swipe, up swipe, back to desktop, and back. The hot words are system-level hot words, and are applicable to any application on the terminal device.
[0013] In some embodiments, the application to which the second interface belongs belongs to an audio category, a short video category, or a long video category, and the pre-set hot words include at least one of the following: play, pause, stop, fast forward, and fast backward.
[0014] In some embodiments, the application to which the second interface belongs belongs to an e-book playing category, and the pre-set hot words include at least one of the following: previous page, next page, table of contents, and next chapter.
[0015] In combination with the first aspect, in a possible implementation, before receiving the first voice input by the user while the terminal device displays the second interface, it is detected that the second interface is switched to foreground display, and the terminal device updates the hot words corresponding to the second interface in the cache; after receiving the first voice input by the user while the terminal device displays the second interface, the terminal device obtains the hot words corresponding to the second interface from the cache.
[0016] In the method, when an interface is switched to the foreground, the corresponding hot words in the cache are updated. After receiving the user voice, the hot words corresponding to the interface are directly obtained from the cache, without the need for real-time analysis from the currently displayed interface. The speech recognition rate is faster, the terminal device responds to the user voice faster, and the user's use experience is improved.
[0017] In combination with the first aspect, in a possible implementation, the terminal device updates the hot words corresponding to the second interface in the cache, including: the terminal device obtains the scene-level hot words corresponding to the second interface, and marks the scene-level hot words corresponding to the second interface in the cache. The scene is pre-configured according to the type of the application, and the correspondence between the scene and the scene-level hot words is pre-configured.
[0018] In combination with the first aspect, in a possible implementation, the cache of the terminal device includes at least one list item, and each list item is used to save the interface hot words corresponding to an interface (the interface hot words are applicable to one interface). The terminal device updates the hot words corresponding to the second interface in the cache, including: if the list item corresponding to the second interface does not exist in the cache of the terminal device, the terminal device creates the list item corresponding to the second interface in the cache; the list item corresponding to the second interface is used to save the interface hot words corresponding to the second interface. Then, the terminal device obtains the text included in the second interface; and the terminal device saves the interface hot words corresponding to the second interface in the list item corresponding to the second interface according to the text included in the second interface.
[0019] In the method, the interface hot words corresponding to each interface are stored in the cache. The cache can include a plurality of list items, and each list item stores the interface hot words corresponding to an interface. In this way, when the display interface is frequently switched back and forth, since the interface hot words corresponding to a plurality of display interfaces are stored in the cache, it is not necessary to frequently parse the text from the display interface again when the display interface is switched, thereby reducing the power consumption generated by parsing the text and improving the response speed of the terminal device.
[0020] In combination with the first aspect, in a possible implementation, before the terminal device creates the list item corresponding to the second interface in the cache, if the number of list items in the cache is greater than or equal to a preset threshold, the terminal device removes one or more list items in the cache.
[0021] In the method, the number of list items stored in the cache is less than or equal to the preset threshold. In this way, the storage space can be saved, and the cache can be prevented from being too large due to a list exception.
[0022] In combination with the first aspect, in a possible implementation, removing one or more list items in the cache includes removing the list items corresponding to one or more interfaces with the shortest total accumulated calling time from the cache.
[0023] In combination with the first aspect, in a possible implementation, the first voice is the first user voice received after the second interface is switched to the foreground display. That is, the hot words corresponding to the interface are obtained from the cache only when the first user voice is received after the interface is switched to the foreground display. If the first user voice is not received after the interface is switched to the foreground display, the voice processing engine can directly parse and recognize the user voice without obtaining the hot word set again.
[0024] In the method, since the hot words correspond to the interface, if the interface is not switched, the hot words do not need to be obtained again. The processing flow is more concise.
[0025] In combination with the first aspect, in a possible implementation, after the second interface is switched to the foreground, the terminal device obtains the system-level hot words, the scene-level hot words corresponding to the second interface, and the interface hot words from the cache after the terminal device receives the first voice of the user for the first time.
[0026] The terminal device obtains the parsed text of the first voice, and performs voice recognition on the parsed text of the first voice according to the hot words corresponding to the second interface. In this way, the terminal device first matches the parsed text of the first voice with the system-level hot words, and if the matching fails, matches the parsed text of the first voice with the scene-level hot words corresponding to the second interface, and if the matching fails, matches the parsed text of the first voice with the interface hot words corresponding to the second interface.
[0027] The priority of the system-level hotword is greater than the priority of the scene-level hotword, and the priority of the scene-level hotword is greater than the priority of the interface hotword. In this way, the voice analysis and recognition can be closer to the real intention of the user, and misoperation can be avoided.
[0028] In a second aspect, a terminal device is provided, which has the function of implementing the method of the first aspect. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0029] In a third aspect, a terminal device is provided, which includes a processor, a memory, and a microphone. The memory is used to store computer execution instructions. When the terminal device is running, the processor executes the computer execution instructions stored in the memory, so that the terminal device executes the method of any one of the first aspect.
[0030] In a fourth aspect, a terminal device is provided, which includes a processor. The processor is used to be coupled with a memory, and after reading the instructions in the memory, the processor executes the method of any one of the first aspect according to the instructions.
[0031] In a fifth aspect, a computer readable storage medium is provided, which stores instructions. When the instructions are run on a computer, the computer can execute the method of any one of the first aspect.
[0032] In a sixth aspect, a computer program product is provided, which contains instructions. When the instructions are run on a computer, the computer can execute the method of any one of the first aspect.
[0033] In a seventh aspect, an apparatus (for example, the apparatus can be a chip system) is provided, which includes a processor for supporting a terminal device to implement the functions involved in the first aspect. In a possible design, the apparatus further includes a memory for storing necessary program instructions and data of the terminal device. When the apparatus is a chip system, it can be composed of a chip, or can include a chip and other discrete devices.
[0034] The technical effects brought by any one of the designs of the second aspect to the seventh aspect can be referred to the technical effects brought by the different designs of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0035] FIG. 1A A scene example schematic diagram of a method for controlling a terminal device by voice;
[0036] FIG. 1B Another scene example schematic diagram of a method for controlling a terminal device by voice;
[0037] FIG. 2A A schematic diagram illustrating a scenario of another method for controlling terminal devices via voice.
[0038] FIG. 2B Schematic diagram 2 illustrates another scenario for controlling terminal devices via voice.
[0039] FIG. 3 This is a flowchart illustrating a method for controlling a terminal device via voice.
[0040] FIG. 4A A schematic diagram illustrating a scenario example of a method for controlling a terminal device via voice, as provided in this application embodiment;
[0041] FIG. 4B Schematic diagram 2 illustrating a scenario example of a method for controlling a terminal device via voice, as provided in an embodiment of this application;
[0042] FIG. 5 This application provides a scenario example illustrating the method for controlling a terminal device via voice according to an embodiment of the present application. FIG. 3 ;
[0043] FIG. 6 Schematic diagram four illustrating a scenario example of a method for controlling a terminal device via voice, as provided in this application embodiment;
[0044] FIG. 7 This application provides a scenario example illustrating the method for controlling a terminal device via voice according to an embodiment of the present application. FIG. 5 ;
[0045] FIG. 8 This application provides a scenario example illustrating the method for controlling a terminal device via voice according to an embodiment of the present application. FIG. 6 ;
[0046] FIG. 9 This is a schematic diagram of the hardware structure of a terminal device provided in an embodiment of this application;
[0047] FIG. 10 This application provides a schematic diagram of the software architecture of a terminal device.
[0048] FIG. 11 A schematic diagram illustrating the interaction of various modules of the terminal device when implementing the method of controlling the terminal device via voice provided in this application embodiment;
[0049] FIG. 12 The second diagram illustrates the interaction between various modules of the terminal device when implementing the method of controlling the terminal device via voice provided in this application embodiment;
[0050] FIG. 13A flowchart of a method for controlling a terminal device by voice provided in an embodiment of the present application is shown in FIG. 1.
[0051] FIG. 14 A schematic diagram of a terminal device structure provided in an embodiment of the present application is shown in FIG. 2.
[0052] FIG. 15 A schematic diagram of a chip system structure provided in an embodiment of the present application is shown in FIG. 3. DETAILED DESCRIPTION
[0053] In order to clearly describe the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application:
[0054] In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing the specific embodiments of the present application, and are not intended to be limiting to the present application. As used in the specification and the appended claims of the present application, the singular forms "a," "an," and "the" are intended to include the plural forms, such as "one or more," unless the context clearly indicates otherwise. It will be further understood that the terms "at least one," "one or more," refer to one or two or more (including two) in the following embodiments of the present application. The term "and / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships; for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects.
[0055] In the present specification, the reference to "one embodiment" or "some embodiments" means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in further some embodiments" and the like appearing in different places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "including but not limited to", unless otherwise specifically emphasized. The term "connected" includes direct connection and indirect connection, unless otherwise specified. "First", "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features.
[0056] In the embodiments of the present application, the word "exemplary" or "for example" is used to mean "an example of" or "an example, only. Any embodiment or design solution described in the embodiments of the present application as "exemplary" or "for example" should not be construed as being more preferred or advantageous than other embodiments or design solutions. In fact, the use of the word "exemplary" or "for example" is intended to present relevant concepts in a concrete manner.
[0057] Voice control function:
[0058] Many terminal devices support voice control function. The terminal device collects user voice through the microphone, analyzes and identifies the user voice, and executes the instruction corresponding to the user voice, so as to realize the control of the terminal device by the user through the voice.
[0059] Generally speaking, in order to save the power consumption of the terminal device and avoid false triggering, the voice control function needs to be turned on before it can be used. For example, the user inputs a preset word (called wake-up word) to the terminal device through voice to wake up the terminal device. After the terminal device is woken up, it can execute the instruction corresponding to the user voice, that is, the voice control function is turned on. For example, the user can turn on or off the voice control function by opening or closing a preset switch on the human-computer interaction interface of the terminal device.
[0060] In different terminal devices, the voice control function can have different names, such as "voice control", "intelligent voice", "voice assistant", "visible and speakable", "voice instruction", "heart instruction", "intelligent AI", etc. The specific implementation of the voice control function with different names can also have some differences.
[0061] The following exemplary introduces several different implementation forms of the voice control function.
[0062] Voice assistant:
[0063] Before the user uses the voice assistant to control the terminal device, the terminal device needs to be woken up. In one example, the terminal device is woken up when the wake-up word input by the user to the terminal device is detected. In another example, the terminal device is woken up when the operation of the user pressing the power key for a long time is detected. In yet another example, the terminal device is woken up when the breath generated when the user inputs voice to the terminal device is detected. Generally speaking, before the terminal device is woken up, the microphone of the terminal device works in a power-saving mode (such as searching for signals at a lower power) to pick up the sound of the surrounding environment. The voice collected by the microphone is only detected for the wake-up word in the kernel layer, and the corresponding recording channel of the voice assistant in the system and driver of the terminal device is not started.
[0064] In response to a user operation (e.g., receiving a wake-up word of a user voice input), the terminal device is woken up, and a corresponding recording channel of the voice assistant in the system and the driver is started. After the terminal device is woken up, the voice (audio stream) collected by the microphone is sent to the voice assistant application for processing through the corresponding recording channel of the voice assistant. In this way, the terminal device can execute the instruction corresponding to the user voice, and realize the control of the terminal device by the user through the voice; and the function of chatting with the user can also be realized.
[0065] An exemplary, FIG. 1A A scenario example schematic diagram in which a user controls a mobile phone by using a voice assistant is shown. As shown in FIG. 1A The mobile phone displays a desktop interface, and the user inputs a voice "Hello YOYO" to the mobile phone. In response to receiving the wake-up word "Hello YOYO", the mobile phone is woken up. Exemplarily, after the mobile phone is woken up, a voice "I am here" is played to prompt the user that the mobile phone has been woken up. After the mobile phone is woken up, the user can control the mobile phone through the voice. Exemplarily, as shown in FIG. 1A The user inputs a voice "Open video" to the mobile phone, and the mobile phone analyzes and identifies the user voice and executes the instruction corresponding to the voice "Open video". Exemplarily, in response to receiving the user voice "Open video", the mobile phone starts a video application.
[0066] An exemplary, FIG. 1B Another scenario example schematic diagram in which a user controls a mobile phone by using a voice assistant is shown. As shown in FIG. 1B The mobile phone displays a desktop interface, and the user long-presses a power key. In response to receiving the operation of long-pressing the power key, the mobile phone is woken up. After the mobile phone is woken up, the user can control the mobile phone through the voice. Exemplarily, as shown in FIG. 1B The user inputs a voice "Open video" to the mobile phone, and the mobile phone analyzes and identifies the user voice and executes the instruction corresponding to the voice "Open video". Exemplarily, in response to receiving the user voice "Open video", the mobile phone starts a video application.
[0067] In some implementations, after the terminal device is woken up, the user can issue an instruction to the terminal device by inputting a voice to the terminal device, and the terminal device executes the instruction corresponding to the voice. After the terminal device executes an instruction, or within a period of time (such as 8 seconds) after the terminal device is woken up, the terminal device does not respond to the instruction issued by the user through the voice. For example, the terminal device closes the recording channel corresponding to the voice assistant. The user needs to input a wake-up word to the terminal device to wake up the terminal device again, and then can issue an instruction to the terminal device by inputting a voice to the terminal device. That is, after the terminal device is woken up, the terminal device enters a "short receiving" state, and can respond to the instruction issued by the user through the voice within a relatively short period of time (such as 8 seconds).
[0068] In some implementations, when the terminal device is connected to the network, it supports entering a continuous dialogue scenario after being woken up, allowing users to engage in continuous conversations with the terminal device. After each broadcast, the terminal device will continue to pick up audio without needing to be woken up again, until the user exits the continuous dialogue using a command such as "exit."
[0069] What is visible can be said:
[0070] It can be seen that this is implemented locally by the terminal device and does not require a network connection.
[0071] The "Visible and Talkable" feature is controlled by a preset switch. Users can turn the preset switch on to enable "Visible and Talkable" or turn it off to disable it.
[0072] Taking a mobile phone as an example, for instance, FIG. 2A As shown, a user can access the settings function of mobile phone 100; for example, the user clicks the application icon of the "Settings" app on the desktop. In response to the user's click on the "Settings" app icon, mobile phone 100 displays the "Settings" interface 101. The "Settings" interface 101 includes a "Smart Voice" option 102, which is used to configure the smart voice function. For example, in response to the user's click on the "Smart Voice" option 102, mobile phone 100 displays the "Smart Voice" interface 103, which includes a "See and Speak" option 104. The user can click on the "See and Speak" option 104 to configure options related to the "See and Speak" function. For example, refer to... FIG. 2A In response to a user's click on the "Speak When You See It" option 104, the mobile phone 100 displays the "Speak When You See It" interface 105. Optionally, the "Speak When You See It" interface 105 includes a prompt message 106 to instruct the user on how to use the "Speak When You See It" function. The "Speak When You See It" interface 105 also includes a "Speak When You See It" switch 107. The user can click the "Speak When You See It" switch 107 to turn the "Speak When You See It" switch on or off. In one example, in response to receiving a user's click on the "Speak When You See It" switch 107, the mobile phone 100's "Speak When You See It" switch turns on, enabling the "Speak When You See It" function. Optionally, the "Speak When You See It" interface 105 displays a prompt message 108 to inform the user that the "Speak When You See It" function has been successfully enabled.
[0073] In one implementation, after the "See and Say" function is enabled, the mobile phone 100 displays a first recording icon, which indicates that the "See and Say" function is enabled. For example, as shown... FIG. 2A As shown, after the "See and Talk" switch 107 is turned on, the status bar of the mobile phone 100 displays the recording icon 10a, indicating that the "See and Talk" function has been enabled.
[0074] In one scenario, the corresponding preset switch of the terminal device is not turned on, and the microphone of the terminal device is not turned on. When the corresponding preset switch of the visible-to-speak function is turned on, the terminal device starts the microphone and starts the corresponding recording channel of the visible-to-speak function in the system and driver. In this way, the voice (audio stream) collected by the microphone can be sent to the visible-to-speak application through the corresponding recording channel of the visible-to-speak function for processing, so that the user can control the terminal device through voice.
[0075] In another scenario, the corresponding preset switch of the terminal device is not turned on, and the microphone of the terminal device works in a power-saving mode (such as searching for signals at a lower power) to pick up the surrounding environment. When the corresponding preset switch of the visible-to-speak function is turned on, the terminal device starts the corresponding recording channel of the visible-to-speak function in the system and driver. In this way, the voice (audio stream) collected by the microphone can be sent to the visible-to-speak application through the corresponding recording channel of the visible-to-speak function for processing, so that the user can control the terminal device through voice.
[0076] After the visible-to-speak function is turned on, the terminal device starts the corresponding recording channel of the visible-to-speak function in the system and driver, and the terminal device enters a "long recording" state to continuously collect the surrounding sound. The user can issue an instruction to the terminal device through voice at any time without the need to input a wake-up word to wake up the terminal device.
[0077] In one embodiment, after any function on the terminal device turns on the voice recording function of the terminal device (turns on the microphone and turns on the recording channel), the terminal device sends a prompt information to the user to prompt the user that the terminal device is in a voice collection state. In this way, the privacy of the user can be avoided from being disclosed. For example, after the visible-to-speak function is turned on, the terminal device enters a continuous voice collection state, and a second recording icon is displayed on the display interface of the terminal device, which indicates that the recording channel is turned on and is used to prompt the user that the microphone is collecting voice. As shown in FIG. 2A As shown in FIG. 10B, the status bar of the display interface of the mobile phone 100 displays a recording icon 10b, indicating that the recording channel is turned on.
[0078] After the visible-to-speak function is turned on, the terminal device turns on the recording channel and continuously collects the voice of the user through the microphone. The user can input voice to the terminal device at any time. The terminal device analyzes and recognizes the voice input by the user and executes the instruction corresponding to the voice of the user.
[0079] In one example, as shown in FIG. 2BAs shown, mobile phone 100 displays the desktop interface. The user inputs the voice command "Open video" into mobile phone 100. In response to receiving the user's voice command "Open video", mobile phone 100 launches the video application and displays the video application's user interface 120.
[0080] Hot words:
[0081] Terminal devices can be set with hot words. Once the "visible and speakable" setting is enabled, if the received user voice message matches at least one of the set hot words, the terminal device will execute the command corresponding to the user's voice message.
[0082] Hot words can be pre-configured on the terminal device, obtained from the content displayed on the terminal device interface, or entered by the user, etc.
[0083] In one implementation, each time an application switch or in-app display interface change is detected, the terminal device obtains the text information of the current foreground application interface and sets hot words based on the obtained text information. In this way, when the user inputs voice that matches the hot words, the terminal device can be controlled to execute the corresponding command.
[0084] For example, refer to FIG. 3 The user turns on the "Speak When You See It" switch on their terminal device. Then, application 1 runs in the foreground. The voice control module retrieves elements (controls, etc.) from the display interface of application 1, obtains interface information (text included in the display interface), and refreshes the saved interface information. Next, application 2 runs in the foreground. The voice control module retrieves elements (controls, etc.) from the display interface of application 2 and refreshes the saved interface information. The user inputs voice into the terminal device. The voice control module sets hotwords based on the saved interface information. If it determines that the received user voice matches at least one hotword, it executes the command corresponding to that user voice.
[0085] For example, refer to FIG. 2BThe phone displays a desktop interface in the illustrated scenario. After detecting that the desktop application switches to the foreground, the phone acquires and saves the text information in the desktop interface, including “5G”, “8:00”, “clock”, “calendar”, “gallery”, “memo”, “file management”, “email”, “music”, “calculator”, “application store”, “sports health”, “weather”, “browser”, “settings”, “AR application”, “video”, “camera”, “address book”, “phone”, “message”, and the like. The text information “5G”, “8:00”, “clock”, “calendar”, “gallery”, “memo”, “file management”, “email”, “music”, “calculator”, “application store”, “sports health”, “weather”, “browser”, “settings”, “AR application”, “video”, “camera”, “address book”, “phone”, “message”, and the like are set as hot words.
[0086] The user inputs a voice “open video” to the phone. After receiving the voice of the user, the phone analyzes and identifies the voice of the user, determines that the voice “open video” contains the hot word “video”, that is, determines that the voice “open video” matches the hot word “video”, and the phone executes the instruction corresponding to the voice “open video”. Exemplarily, as shown in FIG. 6, the phone displays a video application interface in response to receiving the voice “open video” input by the user. FIG. 2B As shown, in response to receiving the voice “open video” input by the user, the phone starts the video application and displays a user interface of the video application.
[0087] The video application switches to the foreground, and the phone acquires and saves the text information in the display interface of the video application, including “5G”, “8:00”, “media title”, “selection”, “1”, “2”, “3”, “4”, “5”, “6”, “7”, “8”, “9”, and the like. The “5G”, “8:00”, “media title”, “selection”, “1”, “2”, “3”, “4”, “5”, “6”, “7”, “8”, “9”, and the like are set as hot words. The user inputs a voice to the phone, and if the voice of the user matches any one of the hot words, the phone can execute the corresponding instruction.
[0088] In this implementation, the terminal device updates the hot words according to the text information in the current display interface every time the application switches or the display interface in the application switches. The hot words are set and updated according to the text in the current display interface of the terminal device. The text not included in the display interface will not be set as a hot word. The acquisition scenario of the hot words is relatively limited. In this way, the supported voice of the terminal device is relatively limited and the types are relatively few. Exemplarily, when the phone displays the display interface of the video application, if the user inputs a voice “pause”, the phone determines that the voice “pause” does not match all the hot words, and does not execute the corresponding instruction. FIG. 2B
[0089] The method for controlling a terminal device by voice provided in the embodiments of the present application not only supports setting a text in a display interface as a hot word, but also supports setting a text not included in the display interface as a hot word. The acquisition of the hot word does not depend on the text in the display interface. In this way, after the visible-to-say function is enabled, the terminal device can support the user to control the terminal device by more, more comprehensive and more rich voice.
[0090] After the visible-to-say function is enabled, the terminal device enters a continuous receiving state. The user can input voice to the terminal device. The terminal device analyzes and recognizes the voice input by the user, and if it is recognized that the voice of the user matches any hot word, the instruction corresponding to the voice is executed.
[0091] For example, referring to FIG. 1, FIG. 4A The mobile phone 100 displays a short video playing interface 121 of a short video application, and receives the voice “back to the desktop” input by the user. The text information in the short video playing interface 121 includes “5G”, “8:00”, “follow”, “recommendation”, “nearby”, “home page”, “friends”, “message”, “me” and the like, but does not include “back to the desktop”. The method for controlling a terminal device by voice provided in the embodiments of the present application does not depend on the text in the display interface for the acquisition of the hot word, and “back to the desktop” can be set as a hot word. The mobile phone 100 recognizes that the voice “back to the desktop” of the user matches the hot word “back to the desktop”, and then executes the instruction corresponding to the voice “back to the desktop”. In an example, in response to receiving the voice “back to the desktop” input by the user, the desktop application of the mobile phone 100 switches to the foreground, and the mobile phone 100 displays a desktop interface. In this example, the hot word “back to the desktop” is not acquired from the display interface of the mobile phone, the acquisition of the hot word does not depend on the text in the display interface, and the hot word can be pre-configured on the mobile phone.
[0092] The method for controlling a terminal device by voice provided in the embodiments of the present application pre-sets some hot words in the terminal device, and these hot words are not acquired from the display interface. In an implementation, the terminal device also acquires some hot words from the display interface. That is, on the terminal device, part of the hot words come from the text in the display interface, and another part of the hot words are pre-set. In this way, the terminal device supports more, more comprehensive and more rich hot words, and accordingly, the terminal device can support the user to control the terminal device by more, more comprehensive and more rich voice.
[0093] In some embodiments, a system-level hotword is pre-configured in the terminal device, and the system-level hotword is applicable to all applications on the terminal device. When any one of the applications on the terminal device is running in the foreground, if the voice input by the user matches at least one system-level hotword, the terminal device executes an instruction corresponding to the voice of the user. The system-level hotword is global and does not depend on the application or the application interface. In an example, the system-level hotword can include left swipe, right swipe, up swipe, back to desktop, return, and the like.
[0094] For example, referring to FIG. 1, the mobile phone 100 displays the interface of the short video application, and receives the voice input by the user “back to desktop”. The mobile phone 100 identifies that the voice “back to desktop” input by the user matches the system-level hotword “back to desktop”, and then executes an instruction corresponding to the voice “back to desktop”. FIG. 4A For example, referring to FIG. 1, the mobile phone 100 displays the interface of the short video application, and receives the voice input by the user “back to desktop”. The mobile phone 100 identifies that the voice “back to desktop” input by the user matches the system-level hotword “back to desktop”, and then executes an instruction corresponding to the voice “back to desktop”.
[0095] In another example, as shown in FIG. 2, the mobile phone 100 displays the application interface 122 of the e-book application, and receives the voice input by the user “return to desktop”. The mobile phone 100 identifies that the voice “return to desktop” input by the user matches the system-level hotword “back to desktop”, and then executes an instruction corresponding to the voice “return to desktop”. FIG. 4B For example, referring to FIG. 2, the mobile phone 100 displays the interface of the e-book application, and receives the voice input by the user “return to desktop”. The mobile phone 100 identifies that the voice “return to desktop” input by the user matches the system-level hotword “back to desktop”, and then executes an instruction corresponding to the voice “return to desktop”. FIG. 4B For example, referring to FIG. 2, the mobile phone 100 displays the interface of the e-book application, and receives the voice input by the user “return to desktop”. The mobile phone 100 identifies that the voice “return to desktop” input by the user matches the system-level hotword “back to desktop”, and then executes an instruction corresponding to the voice “return to desktop”.
[0096] It should be noted that the mobile phone 100 can perform semantic understanding on the voice input by the user. That is, the voice input by the user can not be completely consistent with the hotword, as long as the semantics are the same or close. That is, the voice input by the user matches the hotword, but does not have to be consistent. For example, in the scenario shown in FIG. 1, the voice input by the user is “return to desktop”, and the system-level hotword includes “back to desktop”. The semantics of “return to desktop” and “back to desktop” are the same, that is, they match. FIG. 4B
[0097] In another example, as shown in FIG. 3, the mobile phone 100 displays the desktop interface, and receives the voice input by the user “left swipe”. The mobile phone 100 identifies that the voice “left swipe” input by the user matches the system-level hotword “left swipe”, and then executes an instruction corresponding to the voice “left swipe”. FIG. 5 For example, referring to FIG. 3, the mobile phone 100 displays the interface of the e-book application, and receives the voice input by the user “return to desktop”. The mobile phone 100 identifies that the voice “return to desktop” input by the user matches the system-level hotword “back to desktop”, and then executes an instruction corresponding to the voice “return to desktop”. FIG. 5 For example, referring to FIG. 3, the mobile phone 100 displays the interface of the e-book application, and receives the voice input by the user “return to desktop”. The mobile phone 100 identifies that the voice “return to desktop” input by the user matches the system-level hotword “back to desktop”, and then executes an instruction corresponding to the voice “return to desktop”.
[0098] In some embodiments, a scene-level hotword is pre-configured in the terminal device, and the scene-level hotword is applicable to all applications in a scene. In an implementation, the terminal device pre-configures multiple scenes, and each scene corresponds to at least one application. For example, the pre-configured scenes can include an audio / video scene, an e-book scene, a home screen scene, a permission popup window scene, and the like.
[0099] In an example, the terminal device is preconfigured with a correspondence between an application and a scenario. For example, the correspondence between an application and a scenario is shown in Table 1. The terminal device can determine the scenario corresponding to the application according to the application name (package name of the application).
[0100] Table 1
[0101]
[0102]
[0103] In another example, the terminal device is preconfigured with a correspondence between an application type and a scenario. For example, the correspondence between an application type and a scenario is shown in Table 2. The terminal device can also be preconfigured with a correspondence between an application and an application type. For example, the correspondence between an application and an application type is shown in Table 3. The terminal device can determine the application type according to the application name, and determine the scenario corresponding to the application according to the correspondence between the application type and the scenario.
[0104] Table 2
[0105] Scenario Application type Audio and video scenario Audio, short video, long video E-book scenario E-book playing Home screen scenario Home screen management Permission popup scenario Permission management
[0106] Table 3
[0107]
[0108] It should be noted that the above Tables 1 to 3 are only illustrative of the correspondence between an application and a scenario. In other embodiments, the terminal device can obtain the correspondence between an application and a scenario in other ways; for example, the correspondence between an application and a scenario can not be in the form of a configuration table; for example, the correspondence between an application and a scenario can be different from the above examples. In other embodiments, the terminal device can not save the correspondence between an application and a scenario; for example, the terminal device can obtain the correspondence between an application and a scenario from a server. The embodiments of the present application do not limit the specific way in which the terminal device obtains the scenario to which the application belongs.
[0109] The scenario-level hotword is related to the scenario to which the application belongs, and does not depend on the text contained in the application interface. In an example, the scenario-level hotword corresponding to each scenario is shown in Table 4.
[0110] Table 4
[0111] Scenario Scenario-level hotword Audio and video scenario Play, pause, stop, fast forward, fast backward E-book scenario Previous page, next page, table of contents, next chapter Home screen scenario Application list information, settings menu item Permission popup scenario Allow, confirm, refuse, I know
[0112] It can be understood that in other embodiments, the division of scenarios can be different from the above examples; for example, more or fewer scenarios can be divided. The scenario-level hotword corresponding to each scenario can also be adjusted according to actual conditions.
[0113] When any application on a terminal device is running in the foreground, in response to receiving voice input from the user, the terminal device determines the scene to which the application belongs. If the voice input from the user matches at least one scene-level hot word corresponding to the scene to which the application belongs, the terminal device executes the instruction corresponding to the user's voice.
[0114] For example, FIG. 6 This diagram illustrates a scenario where a mobile phone receives user voice messages in an audio / video environment.
[0115] refer to FIG. 6 The phone displays short video applications (e.g., The short video playback interface 121 is shown. The short video application belongs to the audio and video scenario, and the corresponding scenario-level hot words include: play, pause, stop, fast forward, rewind, etc. The user inputs the voice "pause" into the mobile phone 100. The mobile phone 100 receives the user's voice input "pause", and this voice input "pause" matches the hot word "pause". In response to receiving the voice input "pause", the mobile phone 100 executes the command corresponding to the voice input "pause", pausing the playback of the short video. In one example, such as FIG. 6 As shown, the short video playback interface 121 displays a "Play" button 123 to indicate to the user that the short video has been paused. The user can click the "Play" button 123 to continue playing the short video. The user can also input the voice command "Play" into the phone 100 to continue playing the short video.
[0116] For example, FIG. 7 This diagram illustrates a scenario where a mobile phone receives a user's voice message in an e-book context.
[0117] refer to FIG. 7 The phone displays an e-book application (e.g., The application interface 122. The scenario to which the e-book application belongs is the e-book scenario, and the scenario-level hot words corresponding to the e-book scenario include: previous page, next page, table of contents, next chapter, etc. The user inputs the voice "next page" into the mobile phone 100. The mobile phone 100 receives the user's voice input "next page", and this voice "next page" matches the hot word "next page". In response to receiving the voice "next page", the mobile phone 100 executes the command corresponding to the voice "next page" and displays the e-book application (e.g., Application interface 124. Application interface 124 is page 20, and application interface 122 is page 19. Application interface 124 is the next page after application interface 122.
[0118] In some implementations, the terminal device also obtains hot keywords from the display interface of the foreground application, i.e., it obtains interface hot keywords. For example, see [reference]. FIG. 8 The phone displays short video applications (e.g., The text information in the short video playing interface 121 includes "5G", "8:00", "follow", "recommendation", "nearby", "home page", "friends", "message", "me", and the like. The phone 100 sets the text in the current display interface of the foreground application as the interface hotword, that is, the interface hotword includes "5G", "8:00", "follow", "recommendation", "nearby", "home page", "friends", "message", "me", and the like.
[0119] The phone 100 receives the voice "check message" input by the user, the voice "check message" matches the interface hotword "message", and the phone 100 executes the instruction corresponding to the voice "check message". For example, as shown in FIG. 1C, the phone 100 displays the message interface 125. FIG. 8
[0120] The method for controlling a terminal device through voice provided by the embodiments of the present application supports the pre-set system-level hotword, the scene-level hotword, and the interface hotword obtained from the display interface of the foreground application, and supports more and more comprehensive hotwords, so that the user can control the terminal device through more, more comprehensive, and more rich voice.
[0121] The method for controlling a terminal device through voice provided by the embodiments of the present application can be applied to a terminal device supporting voice input. The terminal device can include a phone, a tablet computer, a notebook computer, a personal computer (PC), an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a smart home device (for example, a smart television, a smart screen, a large screen, a smart speaker, a smart air conditioner, and the like), a personal digital assistant (PDA), a wearable device (for example, a smart watch, a smart bracelet, and the like), a vehicle-mounted device, a virtual reality device, and the like, and the embodiments of the present application do not make any limitation in this regard.
[0122] For example, FIG. 9 Fig. 1 shows a schematic diagram of a terminal device 100. The terminal device 100 can include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headset jack 270D, a sensor module 280, a key 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc. The sensor module 280 can include a temperature sensor, an ambient light sensor, a touch sensor, etc.
[0123] It can be understood that the structure shown in the embodiment does not constitute a specific limitation on the terminal device 100. In other embodiments, the terminal device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0124] The processor 210 can include one or more processing units, for example: the processor 210 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated into one or more processors.
[0125] The controller can be the nerve center and command center of the terminal device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0126] The processor 210 can also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. The memory can hold instructions or data that the processor 210 has just used or is using repeatedly. If the processor 210 needs to use the instructions or data again, it can call them directly from the memory. This avoids repeated access and reduces the waiting time of the processor 210, thus improving the efficiency of the system.
[0127] In some embodiments, the processor 210 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0128] It can be understood that the interface connection relationship between the modules shown in the embodiments is only illustrative and does not constitute a structural limitation on the terminal device. In other embodiments, the terminal device can also use different interface connection methods or combinations of multiple interface connection methods.
[0129] The charging management module 240 is configured to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 240 can receive charging input from a wired charger through the USB interface 230. In some wireless charging embodiments, the charging management module 240 can receive wireless charging input through a wireless charging coil of the terminal device. The charging management module 240 can charge the battery 242 while also supplying power to the terminal device through the power management module 241.
[0130] The power management module 241 is configured to connect the battery 242 and the charging management module 240 to the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240 to power the processor 210, the internal memory 221, the external memory, the display screen 294, the camera 293, and the wireless communication module 260, etc. The power management module 241 can also be configured to monitor parameters such as the battery capacity, the number of battery cycles, the state of health of the battery (leakage, impedance), etc. In some other embodiments, the power management module 241 can also be disposed in the processor 210. In some other embodiments, the power management module 241 and the charging management module 240 can also be disposed in the same device.
[0131] The wireless communication function of the terminal device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 250, the wireless communication module 260, the modem processor, and the baseband processor, etc.
[0132] The terminal device 100 can implement the display function by the GPU, the display screen 294, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is configured to perform mathematical and geometric calculations for graphics rendering. The processor 210 can include one or more GPUs that execute program instructions to generate or change display information.
[0133] The display screen 294 is configured to display images, videos, etc. The display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Mini-LED, a Micro-OLED, a quantum dot light emitting diode (QLED), etc.
[0134] The terminal device 100 can implement the photographing function by the ISP, the camera 293, the video codec, the GPU, the display screen 294, and the application processor, etc.
[0135] ISP is used to process the data feedback from the camera 293. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and conversion into a visible image. ISP can also optimize the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, ISP can be provided in the camera 293.
[0136] The camera 293 is used to capture still images or videos. Objects generate optical images through lenses and project them onto photosensitive elements. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, or other format image signal. In some embodiments, the terminal device can include one or N cameras 293, where N is a positive integer greater than 1.
[0137] The digital signal processor is used to process digital signals, in addition to processing digital image signals, it can also process other digital signals. For example, when the terminal device selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0138] The video codec is used for digital video compression or decompression. The terminal device 100 can support one or more video codecs. In this way, the terminal device can play or record videos in multiple encoding formats, such as moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0139] NPU is a neural-network (NN) computing processor that learns from biological neural network structures, such as the transmission mode between human brain neurons, to quickly process input information and continuously self-learn. Through NPU, the terminal device can achieve intelligent cognition applications, such as image recognition, face recognition, voice recognition, text understanding, etc.
[0140] The terminal device 100 can realize audio functions through the audio module 270, the speaker 270A, the receiver 270B, the microphone 270C, the earphone interface 270D, and the application processor, etc. For example, music playing, recording, etc.
[0141] The audio module 270 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 270 can also be configured to encode and decode audio signals. In some embodiments, the audio module 270 can be disposed in the processor 210, or some functional modules of the audio module 270 can be disposed in the processor 210. The speaker 270A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The receiver 270B, also referred to as a "earpiece", is configured to convert an audio electrical signal into a sound signal. The microphone 270C, also referred to as a "microphone", "microphone", is configured to convert a sound signal into an electrical signal. The earphone interface 270D is configured to connect a wired earphone. The earphone interface 270D can be a USB interface 230, or a 3.5mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface. In the embodiments of the present application, after the voice control function is enabled, the terminal device 100 continuously collects the surrounding sound through the microphone 270C.
[0142] The external memory interface 220 can be configured to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device. The external memory card communicates with the processor 210 through the external memory interface 220 to realize data storage functions. For example, audio, video and other files are saved in the external memory card.
[0143] The internal memory 221 can be configured to store computer executable program codes, and the executable program codes include instructions. The processor 210 executes various functional applications and data processing of the terminal device by running the instructions stored in the internal memory 221. For example, in the embodiments of the present application, the processor 210 can execute the instructions stored in the internal memory 221, and the internal memory 221 can include a storage program area and a storage data area. The storage program area can store an operating system and at least one application program required by a function (such as a sound playing function, an image playing function, etc.). The storage data area can store data (such as a video file) created during the use of the terminal device. In addition, the internal memory 221 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0144] The keys 290 include a power key, a volume key, and the like. The keys 290 can be mechanical keys. The keys 290 can also be touch keys. The motor 291 can generate a vibration prompt. The motor 291 can be used for incoming call vibration prompt, and can also be used for touch vibration feedback. The indicator 292 can be an indicator light, and can be used for indicating a charging state, a power change, and can also be used for indicating a message, a missed call, a notification, and the like. The SIM card interface 295 is used for connecting a SIM card. The SIM card can be inserted into or pulled out of the SIM card interface 295 to realize contact and separation with the terminal device. The terminal device can support one or N SIM card interfaces, and N is a positive integer greater than 1. The SIM card interface 295 can support a Nano SIM card, a Micro SIM card, a SIM card, and the like.
[0145] In the embodiments of the present application, the terminal device 100 is a terminal device that can run an operating system and install an application program. Optionally, the operating system running on the terminal device 100 can be an Android system, an iOS system, a Windows system, and the like.
[0146] In some embodiments, the software system of the terminal device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, or a cloud architecture. The embodiments of the present application take the layered architecture as an example to exemplarily describe the software structure of the terminal device 100.
[0147] The layered architecture divides software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through a software interface. In some embodiments, the Android system is divided into four layers, and from top to bottom, the four layers are an application layer, an application framework layer, an Android runtime and a system library, and a kernel layer.
[0148] The application layer can include a series of application packages APK. For example, camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, and the like. In the embodiments of the present application, the application layer includes a voice control APK, which is used to provide a voice control function of the terminal device 100. The voice control APK includes an interface listening module, a cache management module, an application interaction module, and the like. The interface listening module is used to listen to the starting, exiting, switching, and the like of an application interface. The cache management module is used to manage the acquisition and caching of hot words. The application interaction module is used to manage the process of executing an instruction by an application according to a voice.
[0149] The application program layer further comprises a voice processing engine for parsing, recognizing and processing voice. The voice parsing module is configured to convert voice into text; the voice recognition module is configured to understand and recognize the semantics of the text and determine whether the semantics match hot words; and the instruction mapping module is configured to convert the recognized semantics into machine executable instructions.
[0150] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the application programs of the application program layer. The application framework layer comprises a plurality of pre-defined functions.
[0151] As shown in FIG. 10 , the application framework layer can comprise a window manager, a content provider, a view system, a resource manager, an activity manager service (AMS), a package manager service (PMS), an interface acquisition and analysis module, a multi-mode control module, and the like.
[0152] The window manager is configured to manage window programs. The window manager can acquire the size of a display screen, determine whether there is a status bar, lock a screen, and capture a screen, and the like.
[0153] The content provider is configured to store and acquire data and make the data accessible to application programs. The data can comprise videos, images, audios, dialed and received calls, browsing history and bookmarks, phone books, and the like.
[0154] The view system comprises visual controls, such as a control for displaying text, a control for displaying pictures, and the like. The view system can be used to build application programs. A display interface can be composed of one or more views. For example, a display interface comprising a short message notification icon can comprise a view for displaying text and a view for displaying pictures.
[0155] The resource manager provides various resources for application programs, such as localized strings, icons, pictures, layout files, video files, and the like.
[0156] The AMS is mainly responsible for starting, switching, scheduling the four major components in the system, and managing and scheduling application processes, and the like. The responsibilities and operations of the AMS are similar to the process management and scheduling module in the operating system. When a process or component is initiated, a request is transmitted to the AMS through the Binder communication mechanism, and the AMS uniformly processes the request.
[0157] The PMS handles package management related work, such as application installation and uninstallation, and the like. The PMS is also configured to provide all information of an application program, such as starting and closing, and the like.
[0158] The interface obtaining and analyzing module is configured to analyze elements on the display interface of the terminal device 100, such as a control, text on the control, and the like.
[0159] The multi-mode control module is configured to manage execution of the voice instruction. For example, the instruction is sent to an application, so that the application executes the instruction.
[0160] The system library can include a plurality of functional modules. For example, a surface manager, media libraries, a three-dimensional graphics processing library (for example, OpenGL ES), a 2D graphics engine (for example, SGL), and the like.
[0161] The surface manager is configured to manage a display subsystem, and to provide fusion of 2D and 3D layers for a plurality of application programs.
[0162] The media library supports playback and recording of a plurality of commonly used audio, video formats, and static image files. The media library can support a plurality of audio and video encoding formats, for example, MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, and the like.
[0163] The three-dimensional graphics processing library is configured to implement three-dimensional graphics drawing, image rendering, synthesis, and layer processing, and the like.
[0164] The 2D graphics engine is a drawing engine for 2D drawing.
[0165] The Android runtime is responsible for scheduling and management of the Android system. The Android runtime includes a core library and a virtual machine.
[0166] The core library includes two parts: one part is a function function that needs to be called by the java language, and the other part is the core library of the Android.
[0167] The application program layer and the application program framework layer run in the virtual machine. The virtual machine executes the java files of the application program layer and the application program framework layer into binary files. The virtual machine is configured to perform functions of management of an object life cycle, stack management, thread management, security and exception management, and garbage collection.
[0168] The kernel layer is a layer between hardware and software. The kernel layer can include a display driver, a sensor driver, a microphone driver, a Wi-Fi driver, and the like.
[0169] In combination FIG. 10 , FIG. 11 Interaction between various modules of the terminal device in implementing the method of controlling the terminal device by voice according to the embodiments of the present application is shown.
[0170] The method for controlling a terminal device via voice provided in this application embodiment pre-sets system-level hot words and scene-level hot words for multiple scenarios (such as audio / video scenarios, e-book scenarios, home screen scenarios, and permission pop-up scenarios). For example, system-level hot words include: swipe left, swipe right, swipe up, return to desktop, return, etc.; the scene-level hot words corresponding to each scenario are shown in Figure 4.
[0171] System-level hot keywords are global and do not depend on applications or application interfaces. In one implementation, system-level hot keywords are stored in the terminal device and can be applied to various applications.
[0172] Scene-level hot keywords are applicable to all applications within a given scene. In other words, by determining the scene to which an application belongs, you can obtain the scene-level hot keywords corresponding to that application.
[0173] Terminal devices can also obtain text from the foreground application's display interface, i.e., obtain interface hot words.
[0174] In other words, both scenario-level and interface-level keywords are application-related. When an application starts or switches to the foreground, the terminal device obtains the scenario-level and interface-level keywords corresponding to that application.
[0175] refer to FIG. 11 In some embodiments, in response to user operations or inter-application calls, the terminal device switches between applications or within an application. The interface monitoring module of the voice control APK detects the application switch or the interface switch within the application.
[0176] In one implementation, such as FIG. 12 As shown, AMS can monitor the lifecycle of the Activity corresponding to each interface of each application on the terminal device, including onCreate, onStart, onResume, onPause, onStop, onRestart, and onDestroy.
[0177] onCreate: This event is executed when the Activity is first loaded. The onCreate event is executed the first time the interface is launched, and it is executed again after the Activity is destroyed and reloaded.
[0178] onStart: Executed after the onCreate event. If you switch out of this screen and then re-enter it after a period of time (without destroying the Activity), the onCreate event will be skipped and the onStart event will be executed directly.
[0179] onResume: executed after onStart event; or when the interface is switched to the background, and the user re-views the interface (Activity is not destroyed, and no onStop event is executed), the onCreate event and the onStart event are skipped, and the onResume event is directly executed.
[0180] onPause: called when the interface is switched to the background.
[0181] onStop: executed after the onPause event. If the interface is not returned for a period of time, the onStop event of the interface Activity will be executed. Or the user directly removes the interface from the current page, and the onStop event of the interface Activity will be executed.
[0182] onRestart: after the onStop event is executed, if the interface and the process of the application are not destroyed by the system, when the user re-enters the interface, the onRestart event of the interface Activity will be executed. After the onRestart event, the onCreate event is skipped and the onStart event is directly executed.
[0183] onDestroy: executed when the Activity is destroyed. After the onStop event is executed, if the interface is not returned, the Activity will be destroyed.
[0184] The interface monitoring module registers with the AMS. In this way, when the AMS detects a change in the life cycle of the Activity, the interface monitoring module can be called back to notify. The notification includes the component name corresponding to the Activity, which can include the application package name and the Activity name. According to the change in the life cycle of the Activity, the interface monitoring module can determine that the application switching or the display interface switching within the application has occurred.
[0185] In an example, the display interface 1 of the application 1 corresponds to the Activity 1, and the display interface 2 corresponds to the Activity 2. The AMS detects the onPause event of the Activity 1 and the onResume event of the Activity 2; the onPause event of the Activity 1 and the onResume event of the Activity 2 are notified to the interface monitoring module. The interface monitoring module determines that the display interface 1 is switched to the background according to the onPause event of the Activity 1, and determines that the display interface 2 is switched to the foreground display according to the onResume event of the Activity 2; that is, the display interface switching within the application is monitored.
[0186] In another example, the display interface 1 of the application 1 corresponds to Activity 1, and the display interface 1 of the application 2 corresponds to Activity 2. The AMS listens to the onPause event of Activity 1 and the onResume event of Activity 2, and notifies the interface monitoring module of the onPause event of Activity 1 of the application 1 and the onResume event of Activity 2 of the application 2. The interface monitoring module determines that the display interface 1 of the application 1 is switched to the background according to the onPause event of Activity 1 of the application 1, and determines that the display interface 1 of the application 2 is switched to the foreground display according to the onResume event of Activity 2 of the application 2; that is, the application switching is monitored.
[0187] With reference to FIG. 11 , the interface monitoring module listens to the application switching or the display interface switching in the application, and notifies the cache management module to refresh the cache of the scene-level hot words and the cache of the interface hot words; in an implementation manner, the component name of the display interface switched to the foreground is included in the notification message. The cache management module can obtain the application package name and the Activity name according to the component name.
[0188] In an example, the cache 1 of the terminal device is used to save preset system-level hot words, and the system-level hot words do not need to be updated.
[0189] The cache 2 is used to save preset scene-level hot words of each scene. The cache management module determines the corresponding scene of the application according to the application package name, and marks the scene-level hot words corresponding to the scene in the cache 2. In an implementation manner, the scene-level hot words of each scene are saved in the cache 2 in the form of a list. For example, list 1 is the scene-level hot words of the audio-video scene, list 2 is the scene-level hot words of the e-book scene, list 3 is the scene-level hot words of the home screen scene, and list 4 is the scene-level hot words of the permission pop-up window scene. The state of each list is in the default inactive state. The cache management module determines the corresponding scene according to the application package name, and sets the list corresponding to the scene in the cache 2 to the active state, and sets the remaining lists to the inactive state. For example, the cache management module determines that the application belongs to the audio-video scene according to the application package name, sets list 1 to the active state, and sets list 2, list 3 and list 4 to the inactive state. The list set to the active state is the marked list; that is, the scene-level hot words saved in the list are marked. In this way, the scene-level hot words corresponding to the interface can be conveniently found in the cache 2 when the hot word set corresponding to the interface is determined subsequently. It is not necessary to determine the corresponding scene and the scene-level hot words according to the application to which the interface belongs in real time every time the hot word set corresponding to the interface is determined, and the processing efficiency of obtaining the hot word set after the terminal device receives the user voice is improved.
[0190] The cache 3 is used to store interface hot words. The cache management module updates the cache 3 according to the component name (application package name and Activity name) of the display interface switched to the foreground.
[0191] In an implementation manner, at least one list item can be stored in the cache 3, wherein each list item is used to store a group of interface hot words corresponding to an interface.
[0192] For example, one list item includes the following parameters:
[0193] lastUseTime (last use time), indicating the time when the interface is last invoked;
[0194] totalUseTime (cumulative use time), indicating the total time length of cumulative invocation of the interface;
[0195] cacheId (component identifier), used to uniquely mark a component. For example, the cacheId is the component name, including the application package name and the Activity name;
[0196] hotword (hot word information), used to store the interface hot words corresponding to the interface.
[0197] In an example, the cache management module matches the component name (application package name and Activity name) corresponding to the display interface switched to the foreground with the cacheId (component identifier) of the list item in the cache 3, and finds the list item corresponding to the display interface switched to the foreground in the cache 3.
[0198] If the list item corresponding to the component name does not exist in the cache 3, the list item corresponding to the component name is created, and the parameters in the list item are updated accordingly. For example, lastUseTime = current time, totalUseTime = 0, and cacheId = the component name corresponding to the display interface switched to the foreground. The cache management module calls the interface acquisition and analysis module to acquire the text in the display interface switched to the foreground. In an implementation manner, all the acquired text in the display interface is set as the interface hot words. In another implementation manner, part of the text parsed from the display interface is set as the interface hot words. For example, in an e-book scenario, the text related to the control in the text included in the display interface is set as the interface hot words. Further, the hotword in the newly created list item is set as the interface hot words corresponding to the display interface switched to the foreground.
[0199] If the list item corresponding to the component name exists in the cache 3, the parameter lastUseTime in the list item is updated. For example, lastUseTime = current time. In addition, the component name of the previous display interface is obtained, and the corresponding list item in the cache 3 is searched according to the component name of the previous display interface. For example, the cache management module updates the parameter totalUseTime in the list item corresponding to the previous display interface, totalUseTime = current time - lastUseTime. Wherein, the previous display interface is the display interface displayed by the terminal device before the current display interface is displayed.
[0200] The cache 3 can save multiple list items. When the list item is created, the cacheId and hotword of the list item are recorded; when the display interface switches to the foreground, the lastUseTime in the list item corresponding to the display interface is updated; when the display interface exits the foreground, the totalUseTime in the list item corresponding to the display interface is updated. That is, in an implementation mode, the cache management module only needs to call the interface acquisition and analysis module to obtain the text in the display interface that switches to the foreground (i.e., obtain the interface hotword corresponding to the display interface that switches to the foreground) when creating a list item of a display interface, and does not need to call the interface acquisition and analysis module to analyze the display interface to obtain the interface hotword corresponding to the display interface every time the display interface switches to the foreground. Reduce the power consumption caused by analyzing the display interface when the display interface frequently switches.
[0201] In an implementation mode, the number of list items saved in the cache 3 is less than or equal to a preset threshold. In this way, storage space can be saved, and the cache 3 can be prevented from being too large due to list item abnormalities.
[0202] When a list item needs to be added in the cache 3, it is first determined whether the number of list items saved in the cache 3 is greater than or equal to a preset threshold. If the number of list items saved in the cache 3 is greater than or equal to the preset threshold, one or more list items saved in the cache 3 are cleared, so that the number of list items saved in the cache 3 is less than the preset threshold; and then the list item is created.
[0203] In an example, one or more list items to be cleared are determined according to totalUseTime in the list item. For example, N list items with the smallest totalUseTime value (i.e., the shortest total time length of the display interface accumulated by the call) are cleared; wherein N is greater than or equal to 1.
[0204] In the method, after the application switching or the display interface switching in the application is monitored, the corresponding scene-level hotword is obtained according to the scene to which the foreground application belongs, and is marked in the cache 2; the corresponding interface hotword is obtained according to the display interface of the foreground application, and the corresponding list item in the cache 3 is updated. In this way, through the cache 1, the cache 2 and the cache 3, the system-level hotword, the scene-level hotword and the interface hotword corresponding to the display interface switched to the foreground can be saved.
[0205] Further, when the terminal device displays the display interface switched to the foreground, the user can input the voice to the terminal device. The microphone of the terminal device receives the user voice and transmits the user voice to the application interaction module of the voice control APK. The application interaction module transmits the user voice to the voice processing engine for processing.
[0206] In an implementation manner, if the user voice is received for the first time after the current display interface is switched, the voice processing engine calls the hotword obtaining module to update the hotword set. In an example, the hotword obtaining module reads the system-level hotword from the cache 1, reads the scene-level hotword corresponding to the scene to which the foreground application belongs from the cache 2, and reads the interface hotword corresponding to the current display interface of the foreground application from the cache 3. The obtained system-level hotword, scene-level hotword and interface hotword are merged into the hotword set corresponding to the current display interface of the foreground application, and the hotword set is saved in the cache 4. Optionally, in another example, since the system-level hotword is applicable to any application, after the hotword set in the cache 4 is generated for the first time, the system-level hotword in the hotword set does not need to be updated subsequently. If the interface switching in the application or the application switching belongs to the same scene, the scene-level hotword in the hotword set does not need to be updated.
[0207] In an implementation manner, the scene-level hotword is obtained from the list of active states in the cache 2, and the interface hotword is obtained from the cache 3 according to the component name. Since the scene-level hotword has been cached in the cache 2 and the interface hotword has been cached in the cache 3 when the display interface is switched to the foreground, when the user voice is received, only the system-level hotword, the scene-level hotword and the interface hotword need to be read from the cache, and the scene does not need to be judged in real time and the display interface does not need to be parsed, so that the response speed when the voice is input is improved.
[0208] After the hotword set is updated, the voice parsing module in the voice processing engine parses the user voice to obtain the parsed text corresponding to the user voice. The voice parsing module sends the parsed text to the voice recognition module. The voice recognition module performs semantic understanding and recognition on the parsed text to obtain the recognition content; that is, the parsed text is converted into natural language which is easier to understand. For example, the voice recognition module uses a natural language processing (NLP) algorithm to perform semantic parsing on the parsed text.
[0209] In an implementation, when performing semantic understanding and recognition, the voice recognition module compares the parsed text with the hotword set corresponding to the current display interface of the foreground application saved in the cache 4, and if the parsed text matches any one of the hotword set saved in the cache 4, the voice recognition module succeeds in recognition and sends the recognized content to the instruction mapping module.
[0210] When performing semantic understanding and recognition, the parsed text matching the hotword can be that the semantics of the two are the same or the semantics of the two are close or the recognized content contains the hotword, and the parsed text does not need to be exactly the same as the hotword. For example, in the scenario shown in FIG. 6, the voice input by the user is “return to the desktop”, and the system-level hotword includes “return to the desktop”. The semantics of “return to the desktop” and “return to the desktop” are the same, that is, the two are matched. Because the hotword is used when performing semantic understanding and recognition, the semantic recognition has a reference standard, and the accuracy of voice recognition and semantic understanding is improved. FIG. 4B
[0211] In an implementation, the parsed text is first matched with the system-level hotword in the cache 4, and if the matching succeeds (that is, the parsed text matches any one of the system-level hotword in the cache 4), the recognized content is returned; if the parsed text fails to match the system-level hotword in the cache 4, the parsed text is matched with the scenario-level hotword in the cache 4. If the parsed text matches any one of the scenario-level hotword in the cache 4, the recognized content is returned; if the parsed text fails to match the scenario-level hotword in the cache 4, the parsed text is matched with the interface hotword in the cache 4. If the parsed text matches any one of the interface hotword in the cache 4, the recognized content is returned; if the parsed text fails to match the interface hotword in the cache 4, the voice recognition fails.
[0212] That is, when performing voice recognition, the parsed text is preferentially matched with the system-level hotword; if the matching fails in the system-level hotword, the parsed text is matched with the scenario-level hotword; if the matching fails in the scenario-level hotword, the parsed text is matched with the interface hotword. The priority of the system-level hotword is higher than that of the scenario-level hotword, and the priority of the scenario-level hotword is higher than that of the interface hotword. In this way, the voice parsing and recognition can be closer to the real intention of the user, and misoperation can be avoided. For example, in the e-book scenario, the text in the display interface of the e-book application contains “come to the desktop”. When the terminal device receives the voice “return to the desktop” input by the user, the parsed text “return to the desktop” is first matched with the system-level hotword, and the system-level hotword includes “return to the desktop”, so that the terminal device can execute the instruction corresponding to “return to the desktop” and display the desktop interface. Instead of matching the parsed text “return to the desktop” with the text “come to the desktop” in the display interface of the e-book application.
[0213] Further, with reference to FIG. 6,FIG. 11 The voice recognition module sends the recognized content to the instruction mapping module. The instruction mapping module receives the recognized content and converts the recognized content into an instruction that can be recognized and executed by the terminal device. The instruction mapping module sends the instruction to the multi-modal control module, and the multi-modal control module processes the instruction.
[0214] In an implementation, the instruction mapping module simulates a corresponding control event according to the recognized content, and sends the control event as the instruction to the multi-modal control module. The multi-modal control module sends the control event to the corresponding application for processing. For example, referring to FIG. 7 The mobile phone 100 displays the application interface 122 of the e-book application, and the current foreground application is the e-book application. No control "next page" is set on the application interface 122. After receiving the user voice "next page", the instruction mapping module generates a click event of the control "next page" according to the recognized content "next page". The instruction mapping module sends the click event of the control "next page" to the multi-modal control module, and the multi-modal control module sends the click event of the control "next page" to the e-book application for processing.
[0215] In an implementation, the instruction mapping module sends the instruction to the multi-modal control module, and the multi-modal control module calls a software development kit (SDK) of the related application to implement the logic of executing the instruction. For example, still referring to FIG. 7 The mobile phone 100 displays the application interface 122 of the e-book application, and the current foreground application is the e-book application. After receiving the user voice "next page", the instruction mapping module generates a machine instruction to display the next page according to the recognized content "next page". The instruction mapping module sends the machine instruction to display the next page to the multi-modal control module. The multi-modal control module calls the SDK of the e-book application to implement the logic of switching the display interface of the e-book application to the next page of the current display interface.
[0216] The method for controlling a terminal device by voice provided by the embodiments of the present application comprises: when the terminal device receives a user voice, obtaining a hotword set corresponding to a foreground application from a cache, the hotword set comprising a system-level hotword, a scene-level hotword corresponding to the foreground application, and an interface hotword corresponding to a display interface of the foreground application. The hotword not only comprises text in the display interface, but also comprises the system-level hotword and the scene-level hotword, so that the terminal device supports more and richer hotwords. Moreover, the hotword is obtained from the cache, so that it is not necessary to analyze the text from the current display interface in real time, the speed of voice recognition is faster, the terminal device responds to the user voice faster, and the user experience is improved. Especially, when the display interface is frequently switched back and forth, since the interface hotword corresponding to multiple display interfaces is saved in the cache, it is not necessary to frequently analyze the text from the display interface again when the display interface is switched, the power consumption generated by analyzing the text is reduced, and the response speed of the terminal device is improved. If the display interface is switched, and the scene corresponding to the application does not change, it is not necessary to update the scene-level hotword, and the processing efficiency is improved.
[0217] Exemplarily, FIG. 13 A flowchart of the method for controlling a terminal device by voice provided by the embodiments of the present application is shown. As shown in the figure, FIG. 13 The method comprises:
[0218] The interface monitoring module of the terminal device monitors application switching or intra-application display interface switching. The interface monitoring module notifies the cache management module to update the cache. The cache management module obtains the list item corresponding to the previous display interface in the cache 3 according to the component name of the previous display interface of the current display interface, and updates the accumulated use time (totalUseTime) in the list item corresponding to the previous display interface in the cache (cache 3). The cache management module identifies the scene corresponding to the foreground application. If the scene changes, the scene-level hotword corresponding to the foreground application is marked in the cache 2. If the scene does not change, the cache 2 does not need to be updated. The cache management module determines whether the list item corresponding to the current display interface exists in the cache (cache 3). If the list item corresponding to the current display interface exists in the cache (cache 3), the last use time (lastUseTime) in the list item corresponding to the current display interface in the cache (cache 3) is updated. If the list item corresponding to the current display interface does not exist in the cache (cache 3), it is first determined whether the number of list items in the cache (cache 3) reaches a preset threshold. If the number of list items in the cache (cache 3) reaches the preset threshold, one or more list items in the cache (cache 3) are cleared. For example, according to the accumulated use time (totalUseTime) in the list item, the N list items with the shortest accumulated use time are cleared. If the number of list items in the cache (cache 3) does not reach the preset threshold, the cache management module obtains the text in the current display interface, generates an interface hotword according to the obtained text, and creates a list item corresponding to the current display interface in the cache (cache 3). The interface hotword is saved in the list item. In this way, the cache updating process is completed, the scene-level hotword corresponding to the current interface (current application) is marked in the cache 2, and the interface hotword is saved in the cache 3.
[0219] The user can input a voice to the terminal device. The application interaction module of the terminal device receives the user voice, and determines whether the display interface of the terminal device is switched during the period from the last time the user voice is received to the current time the user voice is received. If the display interface of the terminal device is switched, the voice processing engine obtains the hotword set (including the system-level hotword, the scene-level hotword, and the interface hotword) corresponding to the current display interface from the cache 1, the cache 2, and the cache 3. Optionally, if the cache 3 is empty or the list item corresponding to the current display interface does not exist in the cache 3, the application interaction module notifies the cache management module to update the cache. Further, the voice processing engine analyzes and identifies the user voice. If the display interface of the terminal device is not switched, that is, the user voice is not received for the first time on the current display interface, the voice processing engine directly analyzes and identifies the user voice, and does not need to obtain the hotword set again. If the recognition content of the user voice matches any hotword in the hotword set successfully, the instruction corresponding to the user voice is executed. Optionally, if the recognition content of the user voice fails to match the hotword set, the cache management module is notified to update the cache.
[0220] It can be understood that, in order to implement the above functions, the terminal device includes hardware structures and / or software modules corresponding to the functions. Those skilled in the art can easily understand that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed in the present document, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of the present application.
[0221] The embodiments of the present application can divide the terminal device into function modules according to the above method examples. For example, each function module can be divided according to each function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in the form of hardware or software function module. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical function division. Actual implementation can have another division manner.
[0222] In the case of using integrated units, FIG. 14 A possible structural schematic diagram of the terminal device involved in the above embodiments is shown. The terminal device 1400 includes a processing unit 1401, a storage unit 1402, an audio unit 1403, and a display unit 1404.
[0223] The processing unit 1401 is configured to control and manage the actions of the terminal device 1400. For example, it can be used to update the scene-level hot words, interface hot words, etc. in the cache when application switching or intra-application interface switching occurs; it can be used to analyze and recognize the user voice when the user voice is received, and / or for other processes of the technologies described herein.
[0224] The storage unit 1402 is configured to save the program codes and data of the terminal device 1400. For example, it saves the system-level hot words, scene-level hot words, interface hot words, etc. The processing unit 1401 invokes the program codes stored in the storage unit 1402 to execute each step in the above method embodiments.
[0225] The audio unit 1403 is configured to collect the voice uttered by the user, and play the voice, etc.
[0226] The display unit 1404 is configured to display the interface of the terminal device 1400.
[0227] Of course, the unit modules in the aforementioned terminal device 1400 include, but are not limited to, the processing unit 1401, storage unit 1402, audio unit 1403, and display unit 1404. For example, the terminal device 1400 may also include a power supply unit, a communication unit, etc. The power supply unit is used to supply power to the terminal device 1400. The communication unit is used to support communication between the terminal device 1400 and other devices.
[0228] The processing unit 1401 may be a processor or controller, such as a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The processor may include an application processor and a baseband processor. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The storage unit 1402 may be a memory. The audio unit 1403 may include a microphone, speaker, receiver, etc. The display unit 1404 may be a display screen, etc. The communication unit may be a transceiver, transceiver circuit, or communication interface, etc.
[0229] For example, processing unit 1401 is a processor (such as...) FIG. 9 The processor 110 shown can be a memory (such as a storage unit 1402). FIG. 9 The internal memory 121 shown. Audio unit 1403 may include a microphone (e.g., internal memory 121). FIG. 9 The microphone 170C shown), speaker (such as...) FIG. 9 The speaker 170A shown), receiver (such as...) FIG. 9 The receiver 170B shown is shown. The display unit 1404 can be a display screen (such as...). FIG. 9 The display screen 194 shown is an example. The communication unit includes a mobile communication module (such as...). FIG. 9 The mobile communication module 150 and wireless communication module shown are shown. FIG. 9 The wireless communication module 160 shown is a mobile communication module. The mobile communication module and the wireless communication module can be collectively referred to as a communication interface. The terminal device 1400 provided in this embodiment can be... FIG. 15The terminal device 100 is shown. Among them, the above-mentioned processor, memory, microphone, display screen and communication interface and the like can be coupled together, for example, connected through a bus. The processor calls the program code stored in the memory to execute the steps in the above method embodiments.
[0230] The embodiments of the present application also provide a chip system (for example, a system on a chip (SoC)), as shown in the figure. The chip system includes at least one processor 1501 and at least one interface circuit 1502. The processor 1501 and the interface circuit 1502 can be interconnected through a line. For example, the interface circuit 1502 can be used to receive signals from other devices (for example, the memory of the terminal device). For another example, the interface circuit 1502 can be used to send signals to other devices (for example, the processor 1501 or the touch screen of the terminal device). Illustratively, the interface circuit 1502 can read the instructions stored in the memory and send the instructions to the processor 1501. When the instructions are executed by the processor 1501, the terminal device can be caused to execute the steps in the above embodiments. Of course, the chip system can also include other discrete devices, and the embodiments of the present application do not make specific limitations thereto.
[0231] The embodiments of the present application also provide a computer readable storage medium, which includes computer instructions, when the computer instructions run on the above-mentioned terminal device, make the terminal device execute the functions or steps in the above-mentioned method embodiments.
[0232] The embodiments of the present application also provide a computer program product, when the computer program product runs on the computer, makes the computer execute the functions or steps executed by the mobile phone in the above-mentioned method embodiments.
[0233] Among them, the terminal device 1400, the chip system, the computer readable storage medium or the computer program product provided by the embodiments of the present application are all used to execute the corresponding methods provided above, so the beneficial effects they can achieve can refer to the beneficial effects in the corresponding methods provided above, which will not be repeated here.
[0234] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of functional modules is taken as an example, and in actual application, the above-mentioned functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0235] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiments described above are merely illustrative, for example, the division of the modules or units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.
[0236] The units described as separate components can or can not be physically separate, and the components shown as units can be one physical unit or a plurality of physical units, that is, can be located in one place, or can be distributed to a plurality of different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0237] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0238] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application essentially or the parts that make contributions to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for making a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the method described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk and various program code storage media.
[0239] The above is merely a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto, and any change or replacement within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of controlling a terminal device by voice, applied to a terminal device, the terminal device comprising a microphone, characterized in that, The method comprises: The terminal device displays a first interface, and the first interface comprises a first switch; In response to a user opening the first switch, the terminal device starts a voice control function; starting the voice control function comprises starting a recording channel corresponding to the voice control function, and the recording channel is used to acquire user voice collected by a microphone; In response to inter-application switching or interface switching within an application, the terminal device displays a second interface, and the terminal device acquires hot words corresponding to the second interface; the hot words corresponding to the second interface comprise system-level hot words, scene-level hot words corresponding to the second interface, and interface hot words; the system-level hot words are preconfigured and applicable to any application on the terminal device; the scene-level hot words are applicable to all applications in a scene; the interface hot words are associated with text included in the second interface; The terminal device receives first voice input by a user when the terminal device displays the second interface, acquires parsed text of the first voice, and matches the parsed text of the first voice with the hot words corresponding to the second interface; The matching of the parsed text of the first voice with the hot words corresponding to the second interface comprises: the terminal device matches the parsed text of the first voice with the system-level hot words; if the parsed text of the first voice fails to match the system-level hot words, the terminal device matches the parsed text of the first voice with the scene-level hot words corresponding to the second interface; if the parsed text of the first voice fails to match the scene-level hot words corresponding to the second interface, the terminal device matches the parsed text of the first voice with the interface hot words corresponding to the second interface; If the parsed text of the first voice matches at least one of the hot words corresponding to the second interface, the terminal device executes a first instruction corresponding to the first voice; In response to a user closing the first switch, the terminal device closes the voice control function; closing the voice control function comprises closing the recording channel corresponding to the voice control function.
2. The method of claim 1, wherein, The terminal device starts the voice control function comprises: The terminal device displays a first recording icon and a second recording icon, and the first recording icon indicates that the voice control function is started, and the second recording icon indicates that the recording channel is started.
3. The method of claim 1, wherein, The hot words comprise at least one of the following: Left swipe, right swipe, up swipe, return to desktop, and return.
4. The method of claim 1, wherein, The application to which the second interface belongs belongs to an audio category, a short video category, or a long video category, and the hot words comprise at least one of the following: Play, pause, stop, fast forward, and fast backward.
5. The method of claim 1, wherein, The application to which the second interface belongs belongs to an e-book playing category, and the hot words comprise at least one of the following: Previous page, next page, table of contents, and next chapter.
6. The method of any one of claims 1-5, wherein, Before the receiving of the first voice input by the user when the terminal device displays the second interface, the method further comprises: After detecting that the second interface is switched to foreground display, the terminal device updates hot words corresponding to the second interface in a cache; After receiving the first voice input by the user when the terminal device displays the second interface, the method further comprises: The terminal device acquires the hot word corresponding to the second interface from the cache.
7. The method of claim 6, wherein, The terminal device updates the hot word corresponding to the second interface in the cache, comprising: The terminal device acquires the scene-level hot word corresponding to the second interface; the scene-level hot word is applicable to all applications in a scene; wherein the scene is pre-configured according to the type of the application, and the correspondence between the scene and the scene-level hot word is pre-configured; The terminal device marks the scene-level hot word corresponding to the second interface in the cache.
8. The method of claim 7, wherein, The cache of the terminal device comprises at least one list item, and each list item is used to save the interface hot word corresponding to an interface, and the interface hot word is applicable to an interface, The terminal device updates the hot word corresponding to the second interface in the cache, comprising: If the list item corresponding to the second interface does not exist in the cache of the terminal device, the terminal device creates the list item corresponding to the second interface in the cache; the list item corresponding to the second interface is used to save the interface hot word corresponding to the second interface.
9. The method of claim 8, wherein, After the terminal device creates the list item corresponding to the second interface in the cache, the method further comprises: The terminal device acquires the text included in the second interface; The terminal device saves the interface hot word corresponding to the second interface in the list item corresponding to the second interface according to the text included in the second interface.
10. The method of claim 9, wherein, Before the terminal device creates the list item corresponding to the second interface in the cache, the method further comprises: If the number of list items in the cache is greater than or equal to a preset threshold, one or more list items in the cache are cleared.
11. The method of claim 10, wherein, The clearing of the one or more list items in the cache comprises: The list items corresponding to one or more interfaces with the shortest total accumulated calling time are cleared from the cache.
12. The method of claim 6, wherein, The first voice is the first voice received by the user after the second interface switches to the foreground display.
13. The method of claim 8, wherein, The method further comprises: The terminal device acquires the parsed text of the first voice; The terminal device performs voice recognition on the parsed text of the first voice according to the hot word corresponding to the second interface.
14. A terminal device, comprising: The terminal device comprises a processor, a memory, a microphone and a display screen, the microphone is used to collect user voice, the display screen is used to display the interface of the terminal device, and the processor is coupled with the memory; the memory is used to store computer program code; the computer program code comprises computer instructions, and when the processor executes the above computer instructions, the terminal device executes the method according to any one of claims 1-13.
15. A computer readable storage medium characterized by: The computer readable storage medium comprises computer instructions, and when the computer instructions run on the electronic device, the electronic device executes the method according to any one of claims 1-13.
Citation Information
Patent Citations
Speech control method, device and terminal equipment
CN105957530A
Voice control method and device of terminal, storage medium and terminal
CN110865755A
Voice control method and electronic equipment
CN113794800A