Voice control method and terminal equipment

CN119968674APending Publication Date: 2025-05-09HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070142.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-14
Filing Date
2023-09-25
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When existing terminal devices are controlled by voice in standby mode, they need to issue voice commands multiple times to trigger functions such as powering on and playing videos, resulting in a slowdown in user operation speed.

Method used

In the standby state of the terminal device, the voice information is obtained to determine whether it contains a wake-up word. If the wake-up word is included, it triggers startup and determines whether it contains voice information indicating media resources, thereby directly displaying or playing the corresponding media resources.

Benefits of technology

It simplifies the steps for users to operate the terminal device in standby mode, improves the speed at which the user controls the terminal device to display media resources, and reduces the need for additional voice commands after waiting for power on.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119968674A_ABST
    Figure CN119968674A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a voice control method and terminal equipment, and the terminal equipment comprises a display which is configured to display media resources; the memory is configured to store computer instructions and data associated with the display, and the at least one processor is connected with the display and the memory and is configured to run the computer instructions to enable the terminal equipment to execute the following operations: the first processor is configured to stop working when the terminal equipment is in a standby state; the second processor is configured to obtain voice information when the terminal equipment is in a standby state; if the voice information comprises the first voice information, triggering a first processor to start, and determining whether the voice information comprises second voice information except the first voice information; the first voice information is used for indicating a wake-up word; and the first processor is configured to control the display to display the media resource indicated by the second voice information according to the second voice information if the voice information comprises the second voice information.
Need to check novelty before this filing date? Find Prior Art

Description

Voice control method and terminal device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed on December 12, 2022, with application number 202211597149.3; and filed on December 14, 2022, with application number 202211611894.9, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of intelligent terminal technology, and in particular to a voice control method and terminal device. Background Art

[0004] Currently, terminal devices (such as mobile phones and televisions) are becoming increasingly intelligent. For example, terminal devices are beginning to provide far-field voice functions. With this far-field voice function, users can control terminal device functions through voice without performing physical operations. For example, voice control can be used to turn on the terminal device or play a video.

[0005] Among them, when the terminal device is in standby mode, it only responds to the power-on voice that triggers the terminal device to power on. Therefore, if the user wants to control the terminal device to play a video when the terminal device is in standby mode, the user needs to first issue the power-on voice that triggers the terminal device to power on and wait for the terminal device to power on. After the terminal device is successfully powered on, the user then issues a display command voice that triggers the terminal device to play the video. The terminal device receives and responds to the display command voice and displays the video. It can be seen that the user needs to issue voice commands multiple times to control the terminal device to play the video. This undoubtedly increases the number of steps for the user to control the terminal device to play the video in the standby mode, thereby reducing the speed at which the user can control the terminal device to play the video in the standby mode.

[0006] Summary of the Invention

[0007] An embodiment of the present application provides a terminal device, the terminal device comprising:

[0008] A display configured to display media resources; a memory configured to store computer instructions and data associated with the display; at least one processor connected to the display and the memory, configured to execute the computer instructions so that the terminal device performs:

[0009] a first processor, configured to stop operating when the terminal device is in a standby state;

[0010] The second processor is configured to: when the terminal device is in a standby state, obtain voice information; if the voice information includes first voice information, trigger the first processor to start, and determine whether the voice information includes second voice information other than the first voice information; the first voice information is used to indicate a wake-up word;

[0011] The first processor is configured to: if the voice information includes the second voice information, control the display to display the media resource indicated by the second voice information according to the second voice information.

[0012] The present application also provides a voice control method for a terminal device, wherein the terminal device includes a second processor, and the second processor operates when the terminal device is in a standby state. The method includes: when the terminal device is in the standby state, the second processor obtains voice information; if the voice information includes a first voice information, the terminal device is powered on and determines whether the voice information includes a second voice information in addition to the first voice information; the first voice information is used to indicate a wake-up word; if the voice information includes a second voice information, the terminal device displays the media resource indicated by the second voice information based on the second voice information.

[0013] The present application also provides a terminal device that has the function of implementing the method described above for voice control. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0014] An embodiment of the present application also provides a terminal device, including: a processor and a memory; the memory is used to store computer instructions, and when the terminal device is running, the processor executes the computer instructions stored in the memory to enable the terminal device to perform the voice control method as described above.

[0015] The embodiment of the present application also provides a computer-readable non-volatile storage medium, which stores instructions. When the computer-readable non-volatile storage medium is run on a terminal device, the terminal device can execute the voice control method described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] FIG1 is a flow chart of a voice control method provided by the related art;

[0017] FIG2 is a schematic diagram of a scenario of a voice control method provided in some embodiments of the present application;

[0018] FIG3 is a first structural diagram of a terminal device provided in some embodiments of the present application;

[0019] FIG4 is a second structural diagram of a terminal device provided in some embodiments of the present application;

[0020] FIG5 is a flowchart 1 of a voice control method provided in some embodiments of the present application;

[0021] FIG6 is a second flowchart of a voice control method provided by some embodiments of the present application;

[0022] FIG7 is a schematic diagram of a user voice-controlled television video playback method according to some embodiments of the present application;

[0023] FIG8 is a hardware schematic diagram of a terminal device provided in some embodiments of the present application;

[0024] FIG9 is a third flowchart of a voice control method provided by some embodiments of the present application;

[0025] FIG10 is a third structural diagram of a terminal device provided in some embodiments of the present application;

[0026] FIG11 is a schematic diagram of the system architecture of a speech recognition method and a speech recognition device provided in some embodiments of the present application;

[0027] FIG12 is a schematic diagram of a configuration of a terminal device provided in some embodiments of the present application;

[0028] FIG13 is a schematic diagram of a voice interaction network architecture provided by some embodiments of the present application;

[0029] FIG14 is a flow chart of a voice wake-up method provided in some embodiments of the present application;

[0030] FIG15 is a flowchart illustrating speech signal processing according to some embodiments of the present application;

[0031] FIG16 is a flowchart illustrating wake-up word recognition according to some embodiments of the present application;

[0032] FIG17 is a schematic diagram of a process for detecting a wake-up state according to some embodiments of the present application;

[0033] FIG18 is a diagram illustrating the interaction between the first processor and the second processor during secondary verification according to some embodiments of the present application;

[0034] FIG19 is a schematic diagram of a process for verifying speech feature values ​​according to some embodiments of the present application;

[0035] FIG20 is a schematic structural diagram of a chip system provided in some embodiments of the present application. DETAILED DESCRIPTION

[0036] In order to make the purpose and implementation of this application clearer, the exemplary implementation of this application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of this application. Obviously, the described exemplary embodiments are only part of the embodiments of this application, not all of the embodiments.

[0037] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.

[0038] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.

[0039] The terms "comprises" and "comprising" and any variations thereof in this application are intended to cover but not exclude inclusion. For example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed, but may include other components not expressly listed or inherent to such products or devices.

[0040] Televisions are increasingly offering a wide range of features, such as far-field voice control. This feature allows users to control the TV by speaking within a certain range of the TV, without having to operate the controls. For example, users can voice-control the TV to turn on or play a video.

[0041] Among them, since all modules in the TV (for example, controller, communication module) basically stop working when the TV is powered on and in standby mode. Therefore, in order to ensure that the user can control the TV to turn on by voice when the TV is in standby mode, a power-on control module is set in the TV, and the power-on control module still works normally when the TV is in standby mode. The power-on control module can obtain the power-on voice of the user to trigger the TV to turn on, and trigger the TV to turn on according to the power-on voice. Thus, the user voice control of the TV to turn on in standby mode is realized.

[0042] However, in addition to turning the TV on and off, the TV also provides many other functions. For example, the TV can display an application, a channel, a soccer game video, a singing video, and so on. In other words, the TV provides a wide variety of other functions. Therefore, the content of the functional voice commands used by users to trigger other TV functions is also rich and varied. Therefore, it is quite difficult to recognize the functional voice commands that trigger and control other TV functions. The power-on control module in the TV has limited processing power and can only recognize the power-on voice commands, but cannot recognize the rich and diverse functional voice commands. The controller in the TV needs to control the communication module to communicate with the server to request the server to recognize the functional voice commands.

[0043] However, when the TV is in standby mode, the controller and communication module in the TV stop working. Therefore, when the TV is in standby mode, the only way to control the TV to turn on is by voice control through the active power-on control module. If the user wants to control the TV to perform another TV function while the TV is in standby mode, for example, to play a football game video, the user must first issue a power-on voice command to trigger the TV to turn on and wait for the TV to turn on. After the TV successfully turns on, the user then issues a display command voice command to trigger the TV to play the football game video. The TV receives and responds to the display command voice command to display the football game video. It can be seen that the user needs to issue voice commands multiple times to control the TV to play the football game video. This undoubtedly increases the number of steps for the user to control the TV to play the football game video while in standby mode, thereby reducing the speed at which the user can control the TV to play the football game video while in standby mode.

[0044] Exemplarily, the voice control method provided by the related technology as shown in Figure 1 includes the following steps: S11. When the TV is in standby mode, it obtains a voice message 1 sent by the user; S12. It determines whether the voice message 1 includes a wake-up word; S13. If the voice message 1 includes a wake-up word, the TV is turned on and enters the power-on state; S14. If the voice message 1 does not include a wake-up word, the TV remains in standby mode; S15. When the TV is in the power-on state, it obtains a voice message 2 sent by the user; S16. It determines whether the voice message 2 includes a wake-up word; S17. If the voice message 2 includes a wake-up word, the voice message 2 is continued to be recognized to determine the media resources indicated by the voice message 2; S18. If the voice message 2 does not include a wake-up word, the TV remains in the power-on state; S19. The TV displays the media resources indicated by the voice message 2.

[0045] It's understandable that the user needs to send two separate voice messages, both of which include the wake-up word. Furthermore, after sending the first voice message, the user must wait for the TV to power on before sending the second voice message. This undoubtedly increases the number of steps required for the user to control media resources played by the TV in standby mode, thereby reducing the speed at which the user can control media resources played by the TV in standby mode.

[0046] In response to the above problems, some embodiments of the present application provide a voice control method, in which after the terminal device obtains voice information in standby mode, it first determines whether the voice information includes a wake-up word. If the voice information includes a wake-up word, the terminal device is powered on, and the terminal device further continues to determine whether the voice information includes a second voice information indicating a media resource. If the voice information also includes a second voice information indicating a media resource, the terminal device then displays the media resource. In other words, the voice information sent by the user when the terminal device is in standby mode can control the terminal device to power on and also control the terminal device to display media resources. The user does not need to wait for the terminal device to power on before controlling the terminal device to display media resources, which simplifies the steps for the user to control the terminal device in standby mode to display media resources, thereby increasing the speed at which the user controls the terminal device in standby mode to display media resources.

[0047] The following describes the voice control methods provided in some embodiments of the present application.

[0048] The terminal device provided in the embodiments of this application can have various implementation forms. For example, it can be a terminal device with a display, such as a tablet computer, a PC, a television, a smart TV, a laser projection device, an electronic table, etc. Some embodiments of this application do not limit the specific form of the terminal device. In some embodiments of this application, a television is used as an example for illustration.

[0049] Figure 2 is a schematic diagram of a scenario in which a user controls a terminal device according to an embodiment. As shown in Figure 2 , a user can operate terminal device 200, i.e., a television, through control device 100 or smart device 300. Alternatively, the user can control terminal device 200 by speaking voice within a certain range of terminal device 200.

[0050] In some embodiments, the control device 100 may be a remote controller, and the communication between the remote controller and the terminal device 200 may include infrared protocol communication and other short-range communication methods, and the terminal device 200 may be controlled wirelessly or wired. The user may control the terminal device 200 by inputting user commands through buttons on the remote controller, voice input, control panel input, etc.

[0051] In some embodiments, the user may also use a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) to control the terminal device 200. For example, the user may use an application running on the smart device 300 to control the terminal device 200.

[0052] In some embodiments, the terminal device 200 may not use the above-mentioned smart device 300 or the control device 100 to receive instructions, but may receive user instructions through touch or gestures.

[0053] In some embodiments, the terminal device 200 can also be controlled in a manner other than the control apparatus 100 and the smart device 300. For example, the user's voice can be directly received through a voice acquisition module (e.g., a microphone) configured within the terminal device 200, or the user's voice can be received through a voice acquisition device provided external to the terminal device 200. The following describes the methods provided in some embodiments of the present application using an example of receiving the user's voice through a voice acquisition module configured within the terminal device 200.

[0054] In some embodiments, the terminal device 200 also communicates data with the server 400. The terminal device 200 may be connected to a local area network (LAN), a wireless local area network (WLAN), or other networks. The server 400 may provide various content and interactions to the terminal device 200. The server 400 may be a single cluster or multiple clusters, and may include one or more types of servers.

[0055] For example, FIG3 shows a schematic structural diagram of a television provided in some embodiments of the present application.

[0056] As shown in FIG3 , the terminal device 200 includes at least one of a tuner 210 , a communicator 220 , a detector 230 , an external device interface 240 , a controller 250 , a display 260 , an audio output interface 270 , a memory, a power supply, and a user interface 280 .

[0057] In some embodiments, the controller 250 includes: a CPU, a video processor, an audio processor, a graphics processing unit (GPU), a random access memory (RAM), a read-only memory (ROM), a first interface to an nth interface for input / output, a communication bus (Bus), etc.

[0058] The tuner-demodulator 210 receives broadcast and television signals by wired or wireless reception, and demodulates audio and video signals from multiple wireless or wired broadcast and television signals. The detector 230 is used to collect signals from the external environment or interact with the outside. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environmental scenes, user attributes or user interaction gestures; or, the detector 230 includes a sound collector, such as a microphone, for receiving external sounds. The controller 250 and the tuner-demodulator 210 can be located in different split devices, that is, the tuner-demodulator 210 can also be in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0059] Display 260 includes a display screen component for displaying images, a driver component for driving the image display, and a component for receiving image signals output by the controller to display video content, image content, a menu control interface, and a user control UI interface. Display 260 can be at least one of a liquid crystal display, an organic light-emitting diode (OLED) display, a touch display, and a projection display. It can also be a projection device and projection screen.

[0060] Communicator 220 is a component used to communicate with external devices or servers using various communication protocols. For example, the communicator may include at least one of a Wi-Fi module, a Bluetooth module, a wired Ethernet module, or other network communication protocol chip, a near-field communication protocol chip, and an infrared receiver. Terminal device 200 can use communicator 220 to send and receive control signals and data signals.

[0061] The user interface 280 may be configured to receive external control signals.

[0062] In some embodiments, the controller 250 controls the operation of the terminal device 200 and responds to user operations through various software control programs stored in the memory. Exemplarily, the controller includes at least one of a central processing unit (CPU), an audio processor, a graphics processing unit (GPU), random access memory (RAM), read-only memory (ROM), first to nth interfaces for input / output, and a communication bus. The controller 250 controls the overall operation of the terminal device 200. The user can enter user commands through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input commands through the graphical user interface (GUI). Alternatively, the user can enter user commands by inputting specific sounds or gestures, and the user input interface recognizes the sounds or gestures through sensors to receive the user input commands.

[0063] In some embodiments, the sound collector can be a microphone, also known as a "microphone" or "microphone", which is used to convert sound signals into electrical signals. When performing voice interaction, the user can speak by putting their mouth close to the microphone to input the sound signal into the microphone. The terminal device 200 can be provided with at least one microphone. In other embodiments, the terminal device 200 can be provided with two microphones, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the terminal device 200 can also be provided with three, four or more microphones to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.

[0064] Among them, the microphone may be built into the terminal device 200, or the microphone may be connected to the terminal device 200 by wire or wireless means. For example, the microphone may be arranged at the lower edge of the display 260 of the terminal device 200. Of course, some embodiments of the present application do not limit the position of the microphone on the terminal device 200. Alternatively, the terminal device 200 may not include a microphone, that is, the above-mentioned microphone is not arranged in the terminal device 200. The terminal device 200 may be connected to an external microphone (also referred to as a microphone) through an interface (such as a USB interface 130). The external microphone may be fixed to the terminal device 200 by an external fixing member (such as a camera holder with a clip). For example, the external microphone may be fixed to the edge of the display 260 of the terminal device 200, such as the upper edge, by an external fixing member.

[0065] In some embodiments, a "user interface" is a medium interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. A common form of user interface is a graphical user interface (GUI), which refers to a user interface related to computer operations that is displayed in a graphical manner. It can be an interface element such as an icon, window, or control displayed on the display screen of an electronic device, where a control can include at least one of a visual interface element such as an icon, button, menu, tab, text box, dialog box, status bar, navigation bar, or widget.

[0066] In some examples, taking the operating system of the terminal device 200 as the Android system, as shown in Figure 4, the terminal device 200 can be logically divided into an application layer (abbreviated as "application layer") 21, a kernel layer 22 and a hardware layer 23.

[0067] As shown in Figure 4 , the hardware layer may include the controller 250, communicator 220, detector 230, and display 260 shown in Figure 3 . The application layer 21 includes one or more applications. Applications can be system applications or third-party applications. For example, the application layer 21 includes a far-field voice application that can provide far-field voice functionality. The far-field voice application can be specifically used to obtain voice information emitted by the user when the terminal device 200 is powered on, and request the server 400 to recognize the obtained voice information; then, based on the recognition result, the terminal device 200 is controlled to display the media resources indicated by the recognition result.

[0068] The kernel layer 22 serves as a software middleware between the hardware layer and the application layer 21 and is used to manage and control hardware and software resources.

[0069] The server 400 includes a communication control module 201 and a speech recognition module 202. The communication control module 201 is used to establish a communication connection with the terminal device 200. For example, the far-field speech application in the terminal device 200 establishes a communication connection with the communication control module 201 of the server 400 by calling the communicator 220.

[0070] In some examples, the core layer 22 includes a detector driver, which is configured to send voice information collected by the sound collector in the detector 230 to a far-field voice application. When the terminal device 200 is powered on, the far-field voice application and the communicator 220 in the terminal device 200 are activated. The communicator 220 establishes a communication connection with the communication control module 201 in the server 400. The detector driver is configured to send the user input voice information collected by the sound collector in the detector 230 to the far-field voice application. The far-field voice application then sends the voice information to the speech recognition module 202 in the server 400. After receiving the voice information sent by the terminal device 200, the speech recognition module 202 determines the voice text corresponding to the voice information. The speech recognition module 202 sends the voice text corresponding to the voice information to the far-field voice application in the terminal device 200. After receiving the voice text from the server 400, the far-field voice application controls the display 260 to display the media resource indicated by the voice text.

[0071] The voice information involved in this application may be data authorized by the user or fully authorized by all parties.

[0072] The methods in the following embodiments can all be implemented in a terminal device having the above hardware structure.

[0073] It should be noted that the terminal device includes at least one processor, which is connected to the display and the memory and is configured to run computer instructions to enable the terminal device to execute the corresponding program. In some embodiments of the present application, the main control module can be the first processor and the power-on control module can be the second processor.

[0074] The following describes in detail the voice control method provided by some embodiments of the present application in conjunction with Figure 5. As shown in Figure 5, continuing to use the terminal device as the terminal device 200 as an example for schematic description, the voice control method provided by some embodiments of the present application may include the following S501-S503.

[0075] S501: The terminal device 200, i.e., the television, is in standby mode, and the power-on control module obtains voice information.

[0076] Terminal device 200 may include a power-on control module and a main control module. When terminal device 200 is in standby mode, the power-on control module remains operational, while multiple modules within terminal device 200, including the main control module and communication module, cease operation. At this point, if a user speaks within a certain range of terminal device 200, terminal device 200 can receive the user's voice information through the power-on control module and cache the information.

[0077] Among them, the power-on control module in the terminal device 200 is used to obtain voice information and recognize the voice information. This power-on control module is mainly used to recognize the voice information for triggering the TV to power on. This power-on control module can also be called a far-field voice module.

[0078] Exemplarily, the power-on control module in the terminal device 200 can be implemented by Digital Signal Processing (DSP). The main control module in the terminal device 200 can be implemented by a System on Chip (SOC).

[0079] It should be noted that when the terminal device 200 is in the standby state, multiple modules such as the main control module, the communication module, and the display stop working, and the power consumption of the terminal device 200 is reduced. Therefore, this standby state can also be called a low-power state. Secondly, when the terminal device 200 is in the standby state, the display also stops working, and the display is in a black screen state.

[0080] In some embodiments, when the terminal device 200 is in the standby state, the power-on control module can continuously obtain voice information. After the power-on control module obtains the voice information, it can recognize the voice information to determine whether the voice information includes the first voice information. If the voice information includes the first voice information, it indicates that the user successfully wakes up the TV and the TV powers on, that is, S502 is executed. If the voice information does not include the first voice information, the terminal device 200 can remain in the standby state, and the power-on control module can recognize the next voice information obtained after this voice information to determine whether the next voice information includes the first voice information.

[0081] Among them, the first voice information is used to indicate a wake-up word. The wake-up word can be a specified word. For example, the wake-up word can include "Hello **".

[0082] It should be noted that the processing process of the power-on control module for the next voice information can refer to the processing process of the power-on control module for this voice information.

[0083] Exemplarily, when the terminal device 200 is in the standby state, the detector driver and the detector 230 are still in the working state. If the user emits a voice, the sound collector in the detector 230 collects the voice information emitted by the user. The detector driver then sends the voice information collected by the sound collector to the power-on control module. The power-on control module obtains the voice information from the sound collector and then determines whether the voice information includes the wake-up word.

[0084] For example, the power-on control module may obtain voice information of a first duration each time, and then process the obtained voice information of the first duration (e.g., voice recognition). After obtaining a voice message of a first duration, the power-on control module may continue to obtain voice information of the next first duration.

[0085] S502: If the voice information includes the first voice information, the terminal device 200 is powered on and determines whether the voice information includes the second voice information in addition to the first voice information; the first voice information is used to indicate the wake-up word.

[0086] If the power-on control module determines that the voice information includes the first voice information (or includes the wake-up word), indicating that the voice wake-up is successful, the power-on control module can trigger the main control module in the terminal device 200 to start, and the power-on control module determines whether the voice information includes the second voice information in addition to the first voice information. Further, when the main control module is started, the main controller can control the display to display an interface (e.g., the main interface) in the TV.

[0087] In some embodiments, if the voice information includes the first voice information, in addition to the startup control module triggering the main control module to start up, other modules in the terminal device 200 that have stopped working, such as the communication module, may also be started up. For example, the startup control module triggers the startup of other modules in the terminal device 200 that have stopped working, or the main control module triggers the startup of other modules in the terminal device 200 that have stopped working.

[0088] In some embodiments, after the startup control module in terminal device 200 obtains the voice information, it may cache the voice information. Then, after determining that the voice information includes the first voice information, the startup control module further determines whether the voice information includes other voice information in addition to the first voice information, i.e., the second voice information. If the voice information includes the second voice information, terminal device 200 may recognize the second voice information, i.e., execute S503. If the voice information does not include the second voice information, the startup control module may delete the voice information.

[0089] For example, the power-on control module may use Voice Activity Detection (VAD) to determine whether the voice information includes the second voice information in addition to the first voice information.

[0090] S503: If the voice information includes the second voice information, the terminal device 200 displays the media resource indicated by the second voice information according to the second voice information.

[0091] If the startup control module determines that the voice information includes the second voice information, the startup control module may send the second voice information to the main control module in the terminal device 200. The main control module may then display the media resource indicated by the second voice information based on the second voice information. For example, if the second voice information may include a live broadcast of a football game, the main control module may control the display to play the video of the football game.

[0092] Exemplarily, the main control module may display the media resources indicated by the second voice information based on the second voice information, including the following steps: first, identifying the second voice information to obtain the voice text corresponding to the second voice information (which may be called a voice recognition result); and then, based on the voice text, controlling the display to display the media resources indicated by the voice text (that is, the media resources indicated by the second voice information).

[0093] For example, the main control module can send the second voice information to the server through the communication module. The server recognizes the second voice information, obtains the voice text corresponding to the second voice information, and sends the voice text to the main control module.

[0094] In some embodiments, the terminal device 200 may locally store the media resources indicated by the second voice information, or the terminal device 200 obtains the media resources indicated by the second voice information from a server.

[0095] In some embodiments, the terminal device 200 may be installed with a far-field voice application. The main control module may implement the above-mentioned "displaying the media resources indicated by the second voice information according to the second voice information" through the far-field voice application.

[0096] For example, let's take a scenario where a user wants to watch a fitness video while the TV is in standby mode, and the terminal device 200 includes a power-on control module and a main control module. The following describes the voice control methods provided by some embodiments of the present application. As shown in FIG6 , S501 in this method includes S601. The method also includes S602. S502 in this method includes S603, and S503 includes S604-S607.

[0097] S601: The terminal device 200, i.e., the television, is in standby mode, and the power-on control module obtains voice information.

[0098] For example, as shown in FIG7 (a), terminal device 200 is in standby mode and the display of terminal device 200 is black. A user, within a certain range of terminal device 200, says "Hello, I want to exercise." The startup control module can obtain the user's voice message and cache the voice message. The voice message includes "Hello, I want to exercise."

[0099] S602: The power-on control module determines whether the voice information includes first voice information, where the first voice information is used to indicate a wake-up word.

[0100] If it is determined that the voice information includes the first voice information, the power-on control module executes S603 to S604. If it is determined that the voice information does not include the first voice information, the terminal device 200 remains in a standby state.

[0101] For example, continuing to take the example that the voice information obtained by the power-on control module includes "Hello, I want to exercise", the power-on control module can determine that the voice information includes the wake-up word, that is, "Hello", and then execute S603.

[0102] S603: The startup control module triggers the main control module to start.

[0103] For example, as shown in (b) of FIG7 , the main control module is started, and controls the display of the terminal device 200 to display the main interface 701 .

[0104] S604: The startup control module obtains second voice information other than the first voice information in the voice information.

[0105] For example, continuing to take the example that the voice information obtained by the power-on control module includes "Hello, I want to exercise", the power-on control module can obtain the second voice information in the voice information including "I want to exercise".

[0106] It should be noted that, in addition to the boot control module executing S603 first and then S604 as shown in FIG6 , the boot control module may also execute S603 and S604 simultaneously. Some embodiments of the present application do not limit the order of S603 and S604.

[0107] S605: The startup control module sends the second voice information to the main control module.

[0108] S606: The main control module recognizes the second voice information and obtains a voice text corresponding to the second voice information.

[0109] For example, continuing to take the second voice information including "I want to exercise" as an example, the main control module recognizes the second voice information and can obtain the voice text corresponding to the second voice information including "exercise".

[0110] S607: The main control module displays the media resources indicated by the voice text according to the voice text corresponding to the second voice information.

[0111] For example, as shown in FIG7(c), taking the case where the voice text corresponding to the second voice information includes "fitness", the main control module can determine a fitness video indicated by "fitness" and control the display to display the fitness video 702. The fitness video can be a fitness video that the user has previously browsed, or a fitness video customized by the terminal device 200.

[0112] It should be noted that after the user sends "Hello, I want to exercise", in addition to the TV display jumping from a black screen to the main interface 701 and then jumping from the main interface 701 to the interface including the exercise video 702 as shown in Figure 7, the TV display can also jump directly from a black screen to the interface including the exercise video 702. Some embodiments of the present application are not limited to this.

[0113] It is understandable that when the terminal device 200 is in standby mode, after the user utters a voice including a wake-up word, not only can the terminal device 200 be controlled to power on, but also the media resources indicated by the voice can be controlled to be played after the terminal device 200 is powered on. When the terminal device 200 is in standby mode, the user utters a voice once to control the terminal device 200 to display media resources, thereby realizing the control of the terminal device 200 in standby mode to display media resources by uttering a voice once. The process of uttering a voice once to control the terminal device 200 in standby mode to display media resources can be called oneshot. The user does not need to wait for the terminal device 200 to power on before uttering a second voice, which improves the speed of controlling the terminal device 200 to display media resources.

[0114] 8 , the power-on control module in the terminal device 200 is implemented by a DSP, and the main control module is implemented by a SOC. The DSP may include a wake-up module 811 , a VAD module 812 , and an audio interaction module 813 .

[0115] In conjunction with the structure of the terminal device 200 shown in FIG8 , the voice control method provided in some embodiments of the present application is introduced.

[0116] First, regardless of whether the terminal device 200 is in the power-on state or the standby state, the DSP can always obtain voice information through the sound collector. The wake-up module 811 can identify the voice information obtained by the DSP to determine whether the voice information includes the first voice information. If the voice information includes the first voice information, the DSP can send a startup notification message to the SOC. The startup notification message is used to trigger the setting of the general-purpose input / output (GPIO) pin of the SOC to a low level. When the GPIO pin of the SOC is at a low level, the SOC starts. The GPIO pin of the SOC is used to control the pause / start of the SOC.

[0117] Secondly, while determining whether the voice information includes the first voice information, the DSP still obtains voice information through the sound collector. After obtaining the voice information of the first duration, the DSP can determine whether the obtained voice information of the first duration includes the second voice information in addition to the first voice information through the VAD module 812.

[0118] After the SOC is started, the SOC can trigger other modules that have stopped working in the terminal device 200 to start, such as triggering the communicator 220 to start. The SOC can also trigger at least one application (including a far-field voice application) installed in the terminal device 200 to start. The far-field voice application in the SOC can obtain the second voice information from the audio interaction module 813 in the DSP. Then, the far-field voice application in the SOC can communicate with the server through the communicator 220 and request the server to recognize the second voice information. The far-field voice application in the SOC can receive the voice text corresponding to the second voice information sent by the server. The SOC can control the display to display the media resources indicated by the voice text based on the voice text.

[0119] The first duration may be the duration required for SOC startup, such as 2 seconds (s).

[0120] In some embodiments, after determining the second voice information, the startup control module in the terminal device 200 waits for the main control module to send a query request (i.e., the second query request) that triggers the acquisition of the second voice information. After receiving the query request sent by the main control module, the startup control module responds to the query request and sends the second query request to the main control module.

[0121] For example, the voice control method provided by some embodiments of the present application is described using the example of a terminal device 200 including a power-on control module and a main control module, and a far-field voice application installed on the terminal device 200. As shown in FIG9 , S502 in the method may further include S901, and S606 in S503 may further include S902-S906. The method may further include S907-S908.

[0122] S901: The power-on control module determines whether the voice information includes second voice information other than the first voice information.

[0123] The power-on control module may perform VAD on the information in the voice information other than the first voice information to determine whether the information in the voice information other than the first voice information includes the second voice information. If the information in the voice information other than the first voice information includes the second voice information, i.e., the voice information includes the second voice information, then S604 and S902 are executed. If the information in the voice information other than the first voice information does not include the second voice information, i.e., the voice information does not include the second voice information, then the power-on control module may generate a second identifier indicating that the voice information does not include the second voice information, and then S909 is executed.

[0124] S902: The startup control module generates a first identifier; the first identifier indicates that the voice information includes the second voice information.

[0125] If the voice information includes the second voice information, the power-on control module may generate a first identifier. The first identifier indicates that the voice information includes the second voice information. The second voice information is a voice information other than the one used to trigger the power on of the terminal device 200. In other words, the second voice information is a functional voice information. The functional voice information is used to trigger other television functions besides the power on / off function.

[0126] S903: The main control module sends a first query request to the startup control module.

[0127] After the main control module is started, it can send a first query request to the startup control module, where the first query request is used to query whether the user inputs a functional voice.

[0128] S904: The startup control module sends a first identifier to the main control module in response to the first query request.

[0129] After generating the first identifier, the startup control module sends the first identifier to the main control module in response to the first query request sent by the main control module.

[0130] S905: The main control module sends a second query request to the startup control module according to the first identifier.

[0131] The main control module determines that the user inputs the second voice information (ie, functional voice) based on the first identifier sent by the startup control module, and then sends a second query request to the startup control module. The second query request is used to request to obtain the second voice information.

[0132] S906: The startup control module sends a second voice message to the main control module in response to the second query request.

[0133] The startup control module sends a second voice message to the main control module in response to the second query request sent by the main control module.

[0134] S907: The power-on control module generates a second identifier; the second identifier indicates that the voice information does not include the second voice information.

[0135] If the voice information does not include the second voice information, the startup control module may generate a second identifier, wherein the second identifier indicates that the voice information does not include the second voice information, that is, indicates that the user has not input a functional voice.

[0136] S908: The startup control module sends a second identifier to the main control module in response to the first query request.

[0137] After generating the second identifier, the startup control module sends the second identifier to the main control module in response to the first query request sent by the main control module. The main control module can then send the first query request to the startup control module again based on the second identifier sent by the startup control module, thereby re-inquiring the startup control module to determine whether the user has input a functional voice message.

[0138] Exemplarily, continuing with the structure of the terminal device 200 shown in FIG8 as an example, the specific process of the main control module obtaining the second voice information from the power-on control module in some embodiments of the present application is introduced. First, after the SOC is started, a first query request can be sent to the audio interaction module 813 through the far-field voice application. In response to the first query request, the audio interaction module 813 can obtain the first identifier generated by the VAD module 812; and then send the first identifier to the far-field voice application in the SOC. Among them, the VAD module 812 generates the first identifier when it determines that the acquired voice information includes the second voice information.

[0139] Then, based on the first identifier, the far-field voice application in the SOC may send a second query request to the audio interaction module 813. In response to the second query request, the audio interaction module 813 may send the second voice information to the far-field voice application in the SOC.

[0140] The audio interaction module 813 may divide the second voice information into multiple audio blocks (e.g., audio blocks of 512 bytes each) and send these multiple audio blocks sequentially to the far-field voice application in the SOC. Alternatively, the audio interaction module 813 may send the entire second voice information to the far-field voice application in the SOC at once.

[0141] For example, the SOC and the DSP may be connected via a Universal Serial Bus (USB) interface. In this case, the far-field voice application in the SOC and the audio interaction module 813 in the DSP may interact via the USB interface.

[0142] The above mainly introduces the solutions provided by some embodiments of the present application from the perspective of methods. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0143] In some embodiments of the present application, the terminal device (such as a television) can be divided into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in some embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0144] The embodiment of the present application further provides a terminal device. As shown in FIG10 , the terminal device 1000 includes: a display 1001 , a power-on control module 1002 , and a main control module 1003 .

[0145] Among them, the display 1001 is configured to display media resources. The main control module 1003 is configured to stop working when the terminal device is in standby mode. The power-on control module 1002 is configured to: obtain voice information when the terminal device 1000 is in standby mode; if the voice information includes the first voice information, trigger the main control module 1003 to start, and determine whether the voice information includes the second voice information in addition to the first voice information; the first voice information is used to indicate the wake-up word. The main control module 1003 is configured to: if the voice information includes the second voice information, then according to the second voice information, control the display 1001 to display the media resources indicated by the second voice information.

[0146] In one possible implementation, the power-on control module 1002 is further configured to: determine whether the voice information includes the first voice information; if the voice information does not include the first voice information, determine whether the next voice information includes the first voice information; the next voice information is obtained by the power-on control module after the voice information.

[0147] In another possible implementation, the terminal device 1000 includes a communicator 1004. The main control module 1003 is specifically configured to: obtain the second voice information from the power-on control module 1002; control the communicator 1004 to send the second voice information to the server; control the communicator 1004 to receive the speech recognition result of the second voice information sent by the server; determine the media resource indicated by the second voice information based on the speech recognition result, and control the display 1001 to display the media resource indicated by the second voice information.

[0148] In another possible implementation, the power-on control module 1002 is further configured to generate a first identifier if the voice information includes the second voice information; the first identifier indicates that the voice information includes the second voice information. The main control module 1003 is further configured to send a first query request to the power-on control module 1002. The power-on control module 1002 is further configured to send the first identifier to the main control module 1003 in response to the first query request sent by the main control module 1003. The main control module 1003 is further configured to send a second query request to the power-on control module 1002 based on the first identifier sent by the power-on control module 1002. The power-on control module 1002 is further configured to send the second voice information to the main control module 1003 in response to the second query request sent by the main control module 1003.

[0149] Of course, the terminal device 1000 provided in some embodiments of the present application includes but is not limited to the above modules. For example, the terminal device 1000 may also include a memory. The memory may be used to store executable instructions for the terminal device 1000 and may also be used to store data generated during operation of the terminal device 1000, such as acquired voice information.

[0150] In some embodiments of the present application, a terminal device refers to an electronic device with a sound collection function, and may be a smart TV, mobile phone, smart speaker, computer, robot, or other electronic device. To meet the diverse and personalized needs of users, the terminal device has voice recognition technology, allowing users to interact with the terminal device through voice. For example, when the smart TV is in standby mode, the user can use voice recognition technology to wake up the smart TV, that is, wake up the smart TV through a far-field voice command, and turn the smart TV on from standby mode.

[0151] Usually, the wake-up process of a smart TV is to collect user voice and recognize the wake-up word of the user voice. In order to reduce power consumption, the wake-up word recognition usually uses a low-power small model simple network to perform wake-up calculations. When a wake-up word is determined, it rolls back a fixed time, saves the corresponding audio, and transmits the audio to a large model complex network for wake-up calculations, that is, a secondary check to see if it is a real wake-up. If so, it enters the normal wake-up, recognition, semantic understanding and user command response process.

[0152] However, in order to reduce power consumption and cost, there is less space available to independently store audio during voice wake-up. Currently, the memory space for far-field voice audio caching is only 80K-100K. Calculated at a sampling rate of 16000bit / s and a sampling accuracy of 16bit, it can only cache 2.5-3.2s of audio at most. Calculated based on the user's average speaking speed of 2 words / second, it can only accommodate wake-up words within five words. For wake-up words with a slower speaking speed or more than 5 words, the wake-up will fail. In addition, after the wake-up word is recognized in a low-power, small-model, simple network, the audio is rolled back and saved. When it is transmitted to a large-model, complex network for secondary verification, the audio transmission efficiency is slow, and signal processing and feature extraction will be performed again, resulting in a longer wake-up response time, which reduces the user experience.

[0153] Figure 11 shows an exemplary system architecture to which the speech recognition method and speech recognition apparatus of the present application can be applied. As shown in Figure 11, 400 is a server, and terminal devices exemplarily include a smart TV 200a, a mobile device 200b, and a smart speaker 200c.

[0154] In this application, the server 400 and the terminal device 200 communicate data via various communication methods. The terminal device 200 may be connected to a local area network (LAN), a wireless local area network (WLAN), or other networks. The server 400 may provide various content and interactions to the terminal device 200. For example, the terminal device 200 and the server 400 may send and receive information, as well as receive software program updates.

[0155] Server 400 can be a server that provides various services, such as a backend server that supports audio data collected by terminal device 200. The backend server can analyze and process the received audio data and other data, and feed back the processing results (such as endpoint information) to the terminal device. Server 400 can be a server cluster or multiple server clusters, and can include one or more types of servers.

[0156] The terminal device 200 can be hardware or software. When the terminal device 200 is hardware, it can be various electronic devices with sound collection functions, including but not limited to smart speakers, smart phones, TVs, tablets, e-book readers, smart watches, players, computers, AI devices, robots, smart vehicles, etc. When the terminal device 200 is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules (for example, for providing sound collection services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0157] It should be noted that the method for speech recognition provided in some embodiments of the present application can be executed by the server 400, or by the terminal device 200, or by both the server 400 and the terminal device 200, and this application does not limit this.

[0158] In some embodiments, the operating system of the terminal device 200, taking the Android system as an example, as shown in Figure 12, the terminal device 200 can be logically divided into an application layer (abbreviated as "application layer") 21, a kernel layer 22 and a hardware layer 23.

[0159] As shown in Figure 12 , the hardware layer 23 may include the controller 250, communicator 220, and detector 230 shown in Figure 3 . The application layer 21 includes one or more applications. These applications can be system applications or third-party applications. For example, the application layer 21 may include a speech recognition application that provides a speech interaction interface and services for connecting the terminal device 200 to the server 400.

[0160] The kernel layer 22 serves as a software middleware between the hardware layer and the application layer 21 and is used to manage and control hardware and software resources.

[0161] In some embodiments, the kernel layer 22 includes a detector driver, which is used to send the voice data collected by the detector 230 to the voice recognition application. Exemplarily, when the voice recognition application in the terminal device 200 is started and the terminal device 200 establishes a communication connection with the server 400, the detector driver is used to send the voice data of the user input collected by the detector 230 to the voice recognition application. Afterwards, the voice recognition application sends the query information containing the voice data to the intent recognition module 102 in the server 400. The intent recognition module 102 is used to input the voice data sent by the terminal device 200 into the intent recognition model. Among them, the server 400 includes a communication control module 101, an intent recognition module 102 and a data storage module 103, wherein the communication control module 101 is used to control communication and the data storage module 103 is used to store relevant data.

[0162] To clearly illustrate the embodiments of the present application, a speech recognition network architecture provided in some embodiments of the present application is described below in conjunction with Figure 13.

[0163] See Figure 13, which is a schematic diagram of a voice interaction network architecture provided by some embodiments of the present application. In Figure 13, the terminal device 200 is used to receive input information and output the processing results of the information. The voice recognition module is deployed with a voice recognition service for recognizing audio as text; the semantic understanding module is deployed with a semantic understanding service for semantically parsing text; the business management module is deployed with a business instruction management service for providing business instructions; the language generation module is deployed with a language generation service (NLG) for converting instructions for the terminal device to execute into text language; the speech synthesis module is deployed with a speech synthesis (TTS) service for processing the text language corresponding to the instruction and sending it to the speaker for broadcast. In one embodiment, the architecture shown in Figure 13 may have multiple physical service devices deployed with different business services, or one or more functional services may be collected in one or more physical service devices.

[0164] In some embodiments, the following describes an example of a process for processing information input into a terminal device based on the architecture shown in FIG13 , taking the information input into a smart device as a query statement input via voice as an example:

[0165] [Speech Recognition]

[0166] After receiving a query statement input via voice, the smart device may perform noise reduction processing and feature extraction on the audio of the query statement. The noise reduction processing here may include steps such as removing echoes and ambient noise.

[0167] [Semantic Understanding]

[0168] Using acoustic and language models, the natural language understanding module performs natural language understanding on the identified candidate text and associated contextual information, parsing the text into structured, machine-readable information, including business domain, intent, word slots, and other information to express semantics. The executable intent is obtained and the intent confidence score is determined. The semantic understanding module selects one or more candidate executable intents based on the determined intent confidence score.

[0169] [Business Management]

[0170] Based on the semantic analysis results of the query statement text, the semantic understanding module sends query instructions to the corresponding business management module to obtain the query results given by the business service, and executes the actions required to "complete" the user's final request, and feeds back the device execution instructions corresponding to the query results.

[0171] It should be noted that the architecture shown in Figure 13 is only an example and does not limit the scope of protection of this application. In some embodiments of this application, other architectures can also be used to implement similar functions. For example, all or part of the above process can be completed by the terminal device, which will not be described in detail here.

[0172] Based on the above-mentioned terminal device 200, the user can perform voice interaction with the terminal device 200. The terminal device 200 can perform voice recognition on the voice input by the user, wherein voice wake-up, i.e., keyword detection, is a branch of the voice recognition task. The terminal device 200 can perform keyword detection on the voice input by the user and detect a limited number of pre-defined activation words or keywords from a string of voice streams without having to recognize all the voices. The wake-up word can be pre-set. Usually, the Chinese wake-up word is four characters. The more syllable coverage and the greater the syllable difference, the better the relative wake-up and false wake-up performance. For example, in order to reduce power consumption, the terminal device 200 can be set to a low power consumption mode, that is, the terminal device 200 is in standby mode. When the terminal device 200 is in standby mode, the user can use a pre-set wake-up word, such as "Xiao A Xiao A", to wake up the voice interaction function of the terminal device 200, so that the terminal device 200 enters the power-on state from the standby state.

[0173] In some embodiments, the basic principle of voice wake-up is based on keyword spotting technology, also known as wake-up word recognition technology. Its core components mainly include two parts: the wake-up model and the decoding algorithm. The wake-up model needs to be trained before use. By collecting a large amount of pronunciation data of pre-defined wake-up words, the collected data is input into the algorithm model for training, thereby generating a model similar to a compressed package. When voice wake-up is required, the wake-up calculation is performed using this wake-up model.

[0174] The wakeup model can be understood as a sound model that converts speech input into an acoustic representation. More precisely, it provides the probability that a speech belongs to a specific acoustic symbol. The wakeup model parameters are trained based on the feature parameters of the training speech database. During recognition, the feature parameters of the speech to be recognized are matched with the wakeup model to obtain the recognition result.

[0175] In some embodiments, the wake-up model can be compressed through different training methods. During recognition, the required wake-up model can be obtained by calling the pre-set model path (setting), and the wake-up model is transmitted to the recognition channel through the model channel, and the corresponding wake-up threshold (threshold) is given. That is to say, through the model path (setting), the wake-up model (model) and the wake-up threshold (threshold) correspond to the recognition module one by one, and the recognition result is obtained through the recognition module, that is, the probability of recognizing the wake-up word. For example, if the probability of recognizing the wake-up word is 0.85 and the wake-up threshold (threshold) is 0.8, then the wake-up recognition is considered successful.

[0176] In some embodiments, the terminal device 200 performs wake-up word recognition based on the controller 250. The controller 250 includes a first sub-processor 500 and a second sub-processor 600. The first sub-processor 500 uses a low-power, small-model simple network to perform wake-up calculations. When it is determined that there is a wake-up word, it rolls back a fixed time, saves the corresponding audio, and transmits the audio to the second sub-processor 600. The second sub-processor 600 uses a large-model complex network to perform wake-up calculations, that is, a secondary check to see if it is a true wake-up. If so, it enters the normal wake-up, recognition, semantic understanding, and user command response process.

[0177] Since the computing power and power consumption of the processor are strongly correlated, the first sub-processor 500 is used to recognize the wake-up word, so the computing power requirement is low and the power consumption is also low. The second sub-processor 600 needs to use a larger wake-up model to wake up the voice data, so the computing power requirement is high and the power consumption is also high. For example, the first sub-processor 500 can be a DSP (Digital Signal Processing, digital signal processing technology) chip, and the second sub-processor 600 can be a SOC chip (System on a Chip, system-on-chip). In some embodiments of the present application, the first sub-processor 500 can be integrated on the second sub-processor 600, or it can be separated from the second sub-processor 600, and some embodiments of the present application are not limited.

[0178] However, in order to reduce power consumption and cost, the DSP chip has less space to store audio independently. For example, the memory space for far-field voice audio caching of chips such as MTK848 and MT9652 is only 80K-100K. Calculated at a sampling rate of 16000bit / s and a sampling accuracy of 16bit, it can only cache 2.5-3.2s of audio at most. Calculated based on the user's average speaking speed of 2 words / second, it can only meet the needs of wake-up words within five words. For wake-up words with a slower speaking speed or more than 5 words, the wake-up will fail. After the wake-up word is recognized in the DSP chip, it rolls back and saves the audio. When it is transmitted to the SOC chip for secondary verification, the audio transmission efficiency is slow, and signal processing and feature extraction will be performed again, resulting in a longer wake-up response time, which reduces the user experience.

[0179] In order to reduce the occupied storage space and improve the wake-up response speed, some embodiments of the present application also provide a voice wake-up method, which can be applied when performing voice control, and the method can be applied to the terminal device 200. In an embodiment of the present application, the method can be used when determining whether the voice information includes the first voice information. Among them, the terminal device 200 that can apply the voice wake-up method includes: a sound collector 700 and a processor, which may include a first sub-processor and a second sub-processor, wherein the second sub-processor includes a first sub-processor 500 and a second sub-processor 600. The sound collector 700 is used to collect voice information. The first sub-processor 500 and the second sub-processor 600 are configured to execute the above-mentioned voice wake-up method. Figure 14 is a flow chart of the voice wake-up method provided in some embodiments of the present application. As shown in Figure 14, the first sub-processing module, i.e., the first sub-processor 500, is configured to perform the following steps:

[0180] S100 , in response to voice information input by a user, extracting voice feature values ​​from the voice information.

[0181] Among them, the voice feature value is a spectral feature containing a wake-up word, and the spectral feature is obtained by processing the voice information through a voice signal. The terminal device 200 can receive the voice information input by the user through the sound collector 700, and input the voice information into the first sub-processor 500. The first sub-processor 500 receives the voice information input by the user, performs voice signal processing and feature extraction on the voice information, obtains the spectral feature of the voice information, and then performs wake-up word recognition on the spectral feature to determine whether the spectral feature contains the wake-up word, thereby obtaining the voice feature value.

[0182] When collecting voice information input by the user, the sound collector 700 can adopt streaming reception, that is, the audio stream of the voice information is segmented and sent to the first sub-processor 500. For example, every time the sound collector 700 collects 20ms of audio, it sends 20ms of audio to the first sub-processor 500 for voice signal processing, thereby realizing simultaneous reception and voice signal processing, and improving the efficiency of voice signal processing.

[0183] The voice information is processed by voice signal processing, that is, the voice signal corresponding to the voice information is analyzed to extract feature parameters that can represent the essence of the voice information as the basis for wake-up word recognition. Therefore, the first sub-processor 500 receives the voice information input by the user, processes the voice signal, and extracts feature parameters that represent the essence of the voice information, that is, spectral features, such as Mel-frequency cepstral coefficients (MFCCs) and filter bank-based features Fbank (Filter bank).

[0184] FIG15 is a flow chart of voice signal processing in some embodiments of the present application. When the first sub-processor 500 performs voice signal processing on voice information, as shown in FIG15 , it is configured to perform the following steps:

[0185] The voice information is preprocessed, and the voice signal corresponding to the voice information is split into multiple frame audio data segments.

[0186] Among them, preprocessing includes pre-emphasis processing, framing processing and windowing processing. First, the voice signal is pre-emphasized, and the pre-emphasis processing is used to amplify the high frequency band of the voice signal, weaken the influence of low frequency interference, and balance the spectrum. Specifically, pre-emphasis processing can be performed through a high-pass filter. For example, the high-pass filter can be y(t) = x(t) - ax(t-1), and the voice signal is input into the high-pass filter for pre-emphasis processing, where a is the filter coefficient, which can be 0.97, x is the input voice signal, and y is the voice signal after pre-emphasis.

[0187] After the pre-emphasis processing, the voice signal is subjected to frame processing, and the voice signal is split into a plurality of frame audio data segments arranged in sequence according to the formation sequence of the voice signal, wherein two adjacent frame audio data segments contain overlapping areas.

[0188] For example, the speech signal can be split into multiple 20-40ms frame audio data segments. Taking the 25ms frame audio data segment as an example, the frame length of the 16kHz frame audio data segment is 0.025×16000=400 sampling points. The frame shift is usually 10ms, that is, 0.01×16000=160 sampling points. In order to avoid excessive changes between two adjacent frames of audio data segments, there is an overlapping area between two adjacent frames of audio data segments. The overlapping area can be 1 / 2, 1 / 3, etc. of each frame of audio data segment. For example, if the overlapping area is set to 15ms, the overlapping area contains 0.015×16000=240 sampling points.

[0189] Since speech signals fluctuate over a long range, in order to eliminate signal discontinuities at both ends of each audio data frame and reduce spectral leakage, the speech signal is framed and then windowed. This involves multiplying each audio data frame by a window function, such as a Hamming window or Hanning window. This windowing process increases the continuity between multiple audio data frames and increases the continuity between the left and right ends of the frame.

[0190] Since it is usually difficult to see the characteristics of the voice signal in the time domain, it is necessary to convert the voice signal from the time domain to the energy distribution in the frequency domain for analysis. Different energy distributions can represent the characteristics of different voices. Therefore, after preprocessing, it is necessary to perform time-frequency conversion on the frame audio data segment and calculate the power spectrum of the frame audio data segment, wherein the power spectrum is the energy of the spectral line corresponding to the conversion of the frame audio data segment from the time domain to the frequency domain. Specifically, the power spectrum of the frame audio data segment is calculated by the following formula:

[0191] Among them, x i The i-th frame audio data segment after the speech signal is framed is processed. The N-point fast Fourier transform (FFT) of the frame audio data segment after the windowing processing is performed to obtain the frequency spectrum of each frame audio data segment, and the frequency spectrum of the frame audio data segment is squared to obtain the power spectrum of the speech signal.

[0192] In order to simulate the human ear's perception characteristics of speech of different frequencies, after obtaining the power spectrum of the speech signal, it is necessary to use a filter group to convert the linear spectrum. That is, the power spectrum is input into a preset filter to obtain a spectrogram. Specifically, the power spectrum is input into a set of Mel scale triangular filters to extract frequency bands, that is, the power spectrum is multiplied and accumulated with each filter to obtain the energy value of the frame audio data segment in the frequency band corresponding to the filter. Among them, the number of triangular filters can be 20-40. If the number of triangular filters is 22, 22 energy values ​​will be obtained through the triangular filter group. After obtaining the spectrogram, the spectrum features can be obtained according to the spectrogram.

[0193] When obtaining spectral features, a logarithmic operation is performed on the spectrum graph, that is, the logarithmic energy output by the filter, to obtain a logarithmic spectrum domain, and then a discrete cosine transform is performed on the logarithmic spectrum domain, that is, the logarithmic energy is subjected to a discrete cosine transform to obtain spectral features.

[0194] It can be understood that the above is an exemplary method for processing speech signals. Different spectral features correspond to different speech signal processing methods. The spectral features can be Mel-frequency cepstral coefficients (MFCC) or filter bank-based features Fbank (Filter bank). For Mel-frequency cepstral coefficients (MFCC), preprocessing, time-frequency conversion, filters, logarithmic operations, discrete cosine transforms and other steps are required; for filter bank-based features Fbank, after performing logarithmic operations, a set of features can be generated.

[0195] After acquiring the spectral features of the voice information, it is necessary to perform wake-up word recognition on the spectral features to detect whether the wake-up word is included in the recognized spectral features. If the spectral features include the wake-up word, it is determined that the wake-up word recognition is successful. If the spectral features do not include the wake-up word, it is determined that the wake-up word recognition fails. Figure 16 is a flowchart of wake-up word recognition in some embodiments of the present application. As shown in Figure 16, the first sub-processor 500 is further configured to perform the following steps:

[0196] Acquire frequency spectrum features of the voice information, and cache the frequency spectrum features.

[0197] Detecting the wake-up state. The wake-up state is a recognition result of a wake-up word recognition performed on the spectrum feature; the wake-up state includes a success or failure of the wake-up word recognition.

[0198] That is, the first sub-processor 500 is further configured to perform the following steps:

[0199] S1601: Perform voice signal processing on the wake-up voice.

[0200] S1602: Extracting spectral features of the wake-up speech.

[0201] S1603: Cache the extracted spectrum features.

[0202] S1604: Perform wake-up word recognition on the wake-up speech based on the first wake-up model.

[0203] S1605: If the wake-up word recognition fails, execute S1601; if the wake-up word recognition succeeds, execute S1606.

[0204] S1606: Roll back the localized speech feature value.

[0205] As shown in FIG. 17 , FIG. 17 is a flow chart of detecting a wake-up state in some embodiments of the present application. The first sub-processor 500 may have a built-in first wake-up model and perform the following steps when performing wake-up word recognition:

[0206] S1701: Acquire spectrum characteristics.

[0207] S1702: Input the spectral feature into a first wake-up model, calculate the wake-up word on the spectral feature using the first wake-up model, and determine whether the spectral feature contains the wake-up word, so as to obtain a wake-up value of the spectral feature output by the first wake-up model. The wake-up value is used to represent the probability of recognizing the wake-up word.

[0208] S1703: If the wake-up value is greater than or equal to the wake-up threshold, execute S1704; if the wake-up value is less than the wake-up threshold, execute S1705.

[0209] S1704: Determine that the wake-up state is successful in wake-up word recognition.

[0210] S1705: Determine that the wake-up state is a wake-up word recognition failure.

[0211] Exemplarily, the wake-up threshold is 0.8. If the first wake-up model built into the first sub-processor 500 calculates the wake-up word on the spectrum feature and determines that the spectrum feature contains the wake-up word, and the wake-up value is 0.85, then the wake-up state is determined to be a successful wake-up word recognition. If the wake-up value is 0.7, then the wake-up state is determined to be a failed wake-up word recognition.

[0212] After obtaining the wake-up state, if the wake-up state indicates successful wake-up word recognition, roll back and locate the spectrum features containing the wake-up word to obtain the speech feature value. If the wake-up state indicates failed wake-up word recognition, filter the spectrum features.

[0213] In this embodiment, the first sub-processor 500 receives voice information input by the user, performs voice signal processing on the voice information to obtain spectral features, caches the spectral features, and calculates the wake-up words on the spectral features through the first wake-up model. If the calculated wake-up value is greater than or equal to the wake-up threshold, it is determined that the wake-up word recognition is successful, that is, the spectral features contain the wake-up word, and the spectral features containing the wake-up word are rolled back to obtain the voice feature value. If the calculated wake-up value is less than the wake-up threshold, it is determined that the wake-up word recognition fails, that is, the spectral features do not contain the wake-up word, and the spectral features are filtered.

[0214] It is understandable that after the voice information is processed and feature extracted, it is converted from stored wake-up audio into stored feature values, which will reduce the cache space of the overall data. For example, using a 16k, 16-bit sampling format for voice recording, approximately 16000×16 / 2=32Kbytes of storage space is required per second. Recording a 4-word wake-up word at a normal speaking speed requires 2S, and redundancy for slow speaking requires approximately 3S, i.e. 96Kbyte. In addition to other content such as data packets, each transmission requires approximately 100K of storage space. Using a method of storing feature values, for example, a 40-dimensional Mel-frequency cepstral coefficient (MFCC) feature value, 1S of audio is calculated as 100 frames, and each frame of data is 2 bytes. Then, only 100×40×2×3=24K is required, which is 3 / 4 less than the original audio's 100K of storage space. As a result, the first sub-processor 500, such as the MT9652 chip, can support 10S of storage space and can support the recognition of longer wake-up words, such as "Hisense Xiaoju, please turn on the phone."

[0215] S200: Send the speech feature value to the second sub-processor 600.

[0216] To reduce power consumption and cost, as shown in FIG18 , FIG18 is an interaction diagram of the first sub-processor and the second sub-processor during secondary verification in some embodiments of the present application. The first wake-up word recognition is performed by the first wake-up model built into the first sub-processor 500. The first wake-up model is a wake-up model based on a small model and a simple network. After the first wake-up word recognition, if the wake-up value calculated by the first wake-up model for the wake-up word recognition of the spectral features exceeds the wake-up threshold, that is, the presence of the wake-up word in the voice information is detected, the voice information is rolled back for a fixed time, the corresponding voice feature value is saved, and the voice feature value is sent to the second sub-processor 600 for the second wake-up word recognition, that is, to verify whether the wake-up word exists. The second wake-up word recognition is performed by the second wake-up model built into the second sub-processor 600. The second wake-up model is a wake-up model based on a large model and a complex network. After the second wake-up word recognition, if the verification is successful, the terminal device 200 is controlled to enable the voice interaction function, that is, to enter the normal wake-up, recognition, semantic understanding and user command response process. If the verification fails, the terminal device 200 is controlled not to enable the voice interaction function.

[0217] Specifically, the interaction between the first sub-processor 500 and the second sub-processor 600 includes the following steps:

[0218] S1801: Perform voice signal processing on the wake-up voice.

[0219] S1802: Extracting spectral features of the wake-up speech.

[0220] S1803: Cache the extracted spectrum features.

[0221] S1804: Perform wake-up word recognition on the wake-up speech based on the first wake-up model.

[0222] S1805: If the wake-up word recognition fails, execute S1801; if the wake-up word recognition succeeds, execute S1806.

[0223] S1806: Roll back the located speech feature value and execute S1807.

[0224] S1807: Send a power-on command.

[0225] S1808: The second sub-processor 600 is powered on and sends a power-on receipt signal to the first sub-processor 500.

[0226] S1809 : The first sub-processor 500 sends the speech feature value to the second sub-processor 600 .

[0227] S1810: The second sub-processor 600 performs wake-up word recognition based on the second wake-up model.

[0228] S1811: Generate verification results.

[0229] It is understandable that the first sub-processor 500 performs wake-up word recognition on the spectrum features. After recognizing the wake-up word, it rolls back and completes the saving of the voice feature value, and sends the voice feature value to the second sub-processor 600 for verification. The transmission time of the audio determines the response time of the entire wake-up. The transmission of the wake-up audio is replaced by the transmission of the voice feature value. Compared with the transmission format of the audio, the voice feature value can reduce the transmission time. For example, if the audio requires a transmission time of 500-1500ms, the voice feature value may only require 125-250ms. When the second sub-processor 600 performs the second wake-up word recognition calculation, there is no need to perform voice signal processing and feature extraction. The voice feature value obtained from the first sub-processor 500 can be directly used for the wake-up word recognition calculation, thereby reducing the wake-up response time.

[0230] The second sub-processor 600 is configured to perform the following steps:

[0231] S300 , in response to the speech feature value sent by the first sub-processing module, namely the first sub-processor 500 , verify the speech feature value.

[0232] In this embodiment, after the first sub-processor 500 performs voice signal processing and feature extraction on the voice information input by the user, it performs wake-up word recognition on the extracted spectral features. When the wake-up word recognition is successful, the voice feature value is rolled back and the voice feature value is sent to the second sub-processor 600 for secondary wake-up verification. As shown in Figure 19, Figure 19 is a schematic diagram of the process of verifying voice feature values ​​in some embodiments of the present application. The second sub-processor 600 obtains the voice feature value sent by the first sub-processor 500 and inputs the voice feature value into the second wake-up model to obtain the verification result of the voice feature value output by the second wake-up model. The verification result includes verification success and verification failure.

[0233] Specifically, the process of verifying voice features includes the following steps:

[0234] S1901: Acquire speech feature values.

[0235] S1902: Input the speech feature value into the second wake-up model.

[0236] S1903: Obtain the verification result of the speech feature value output by the second wake-up model. If the verification result is successful, execute S1904; if the verification result is failed, execute S1905.

[0237] S1904: Enable the voice interaction function.

[0238] S1905: The voice interaction function remains disabled.

[0239] S400: When the verification succeeds, determine that the voice information includes the first voice information.

[0240] After obtaining the verification result of the voice feature value, the terminal device 200 is controlled to perform subsequent operations based on the verification result. If the verification result is successful, it is determined that the voice information includes the first voice information. Of course, if the verification result is successful, the terminal device 200 can also be controlled to enable a voice interaction function; if the verification result is a failure, the voice interaction function of the terminal device 200 is controlled to remain disabled.

[0241] It is understandable that the method for the second wake-up word recognition is different from the method for the first wake-up word recognition. The difference is that the second wake-up word recognition will be re-confirmed based on the wake-up model of the large model and the complex network, using a more layered network architecture and more accurate classification, thereby achieving a lower false wake-up rate. In other words, the first sub-processor 500 uses a simple wake-up model with lower computing power to perform wake-up word recognition calculations, while the second sub-processor 600 uses a complex wake-up model with higher computing power and higher accuracy to perform wake-up confirmation.

[0242] Exemplarily, the wake-up word is "Xiao A, please turn on the phone". The first sub-processor 500 performs voice signal processing and feature extraction on the voice information in response to the voice information input by the user to obtain spectral features of the voice information, inputs the spectral features into the first wake-up model, and performs wake-up word recognition through the first wake-up model. If the wake-up value calculated by the first wake-up model exceeds the first wake-up threshold (the wake-up threshold of the first wake-up model), it is determined that the wake-up word recognition is successful, that is, it is considered that the voice information input by the user may contain the wake-up word "Xiao A, please turn on the phone". The feature can be rolled back to locate the voice feature value containing the wake-up word, and the voice feature value is sent to the second sub-processor 600 for secondary wake-up verification. The second sub-processor 600 receives the voice feature value, inputs the voice feature value into the second wake-up model, and performs wake-up word recognition through the second wake-up model. If the wake-up value calculated by the second wake-up model exceeds the second wake-up threshold (the wake-up threshold of the second wake-up model), it is determined that the verification is successful, that is, it is considered that the voice information input by the user contains the wake-up word "Xiao A, please turn on the phone", and the terminal device 200 is controlled to turn on.

[0243] It should be noted that in this embodiment, the purpose of the secondary wake-up check is to more accurately identify the voice input signal and avoid false wake-ups. Therefore, in actual applications, the second wake-up threshold should be greater than or equal to the first wake-up threshold. For example, the first wake-up threshold is 0.7 and the second wake-up threshold is 0.9.

[0244] In some embodiments, when the first sub-processor 500 does not detect the presence of a wake-up word in the voice information, the second sub-processor 600 is in a low-power working state or a sleep state. When the first sub-processor 500 detects the wake-up word in the voice information, it can send a start instruction to the second sub-processor 600, triggering the second sub-processor 600 to enter a start-up state, such as a power-on state, from the low-power working state or the sleep state. When the first sub-processor 500 does not detect the wake-up word in the voice information, the second sub-processor 600 is not triggered to enter the start-up state from the low-power working state or the sleep state.

[0245] Therefore, after the first sub-processor 500 recognizes the wake-up word based on the first wake-up model, if the wake-up word recognition is successful, it sends a start instruction to the second sub-processor 600. In response to the start instruction sent by the first sub-processor 500, the second sub-processor 600 enters the startup state and feeds back a start receipt signal to the first sub-processor 500 after the startup is completed. In response to the start receipt signal fed back by the second sub-processor 600, the first sub-processor 500 sends the voice feature value to the second sub-processor 600.

[0246] As shown in Figure 18, the initial state of the terminal device 200 is that the second sub-processor 600 is in a dormant state. The low-power first sub-processor 500 can be kept in an on state. After receiving the voice information input by the user, the first sub-processor 500 recognizes the wake-up word of the voice information and determines whether the voice information contains the wake-up word. If the voice information contains the wake-up word, while rolling back the positioning voice feature value, a power-on instruction is sent to the second sub-processor 600. The second sub-processor 600 receives the power-on instruction and enters the power-on state from the dormant state. At the same time, after the power-on is successful, a power-on receipt signal is fed back to the first sub-processor 500 to notify the first sub-processor 500 that the second sub-processor 600 is currently turned on. The first sub-processor 500 receives the power-on receipt signal fed back by the second sub-processor 600 and sends the voice feature value to the second sub-processor 600. The second sub-processor 600 receives the voice feature value and performs a secondary wake-up check.

[0247] In some embodiments, if the first sub-processor 500 fails to recognize the wake-up word based on the first wake-up model, that is, the voice information does not contain the wake-up word and no secondary wake-up check is required, the second sub-processor 600 is kept in a sleep state.

[0248] For example, when the first sub-processor 500 receives a voice message input by a user and determines that the voice message contains a wake-up word, it sends a power-on instruction to the second sub-processor 600, causing the second sub-processor 600 to enter the power-on state from the sleep state. If it determines that the voice message does not contain the wake-up word, it does not send a power-on instruction to the second sub-processor 600, and the second sub-processor 600 remains in the sleep state.

[0249] In this embodiment, after the first sub-processor 500 performs voice signal processing and feature extraction on the voice information, the feature values ​​are directly saved, replacing the saved audio data. A first wake-up model with lower computing power is then used to perform preliminary wake-up word recognition on the voice information. If the wake-up word recognition is successful, the voice feature values ​​of the voice information are directly sent to the second sub-processor 600, where a second wake-up word recognition is performed on the voice feature values ​​using a second wake-up model with higher computing power. Based on the result of the second wake-up word recognition, a decision is made as to whether to enable the voice interaction function. Because the first wake-up model with lower computing power is used during the first wake-up word recognition process, power consumption during the first wake-up word recognition process is low. Only after the first wake-up word recognition is successful is the second wake-up model with higher computing power enabled for the second wake-up word recognition process. By combining the two models, the low wake-up recognition accuracy and high false alarm rate associated with using the first wake-up model with lower computing power alone are avoided, while the high power consumption associated with using the second wake-up model with higher computing power alone is also avoided, thereby achieving a balance between power consumption and a low false alarm rate.

[0250] Based on the above terminal device 200, some embodiments of the present application further provide a voice wake-up method, which includes the following steps:

[0251] The first sub-processor 500 extracts voice feature values ​​from the voice information in response to the voice information input by the user, and sends the voice feature values ​​to the second sub-processor.

[0252] The speech feature value is a spectrum feature containing the wake-up word, and the spectrum feature is obtained by processing the speech information through speech signal processing.

[0253] The second sub-processor 600 verifies the voice feature value in response to the voice feature value sent by the first sub-processor, and when the verification succeeds, determines that the voice information includes the first voice information.

[0254] It can be seen from the above technical solution that the terminal device and voice wake-up method provided in the above embodiment include a sound collector, a first sub-processor and a second sub-processor. The first sub-processor can respond to the voice information input by the user, perform feature extraction on the voice information, extract the voice feature value of the voice information, and send the voice feature value to the second sub-processor, wherein the voice feature value is a spectral feature containing the wake-up word, and the spectral feature is obtained by voice information through voice signal processing. The second sub-processor can respond to the voice feature value sent by the first sub-processor, verify the voice feature value, and when the verification is successful, determine that the voice information includes the first voice information. The method can cache the voice feature value extracted from the voice information, and when the wake-up word is verified for the second time, directly transmit the voice feature value for the second verification to reduce the occupied storage space and improve the wake-up response speed.

[0255] Some embodiments of the present application also provide a terminal device, including: a processor and a memory; the memory is used to store computer instructions, and when the terminal device is running, the processor executes the computer instructions stored in the memory to enable the terminal device to execute the voice control method provided by some embodiments of the present application.

[0256] Some embodiments of the present application also provide a computer-readable non-volatile storage medium, which stores computer instructions. When the computer instructions are executed on a terminal device, the terminal device can execute the voice control method provided in some embodiments of the present application.

[0257] For example, the computer readable storage medium may be a ROM, a RAM, a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0258] Some embodiments of the present application also provide a device (for example, a chip system) that includes a processor for supporting a terminal device in implementing the voice control method provided in some embodiments of the present application. In one possible design, the device also includes a memory for storing program instructions and data necessary for the terminal device. When the device is a chip system, it can be composed of a chip or include a chip and other discrete components.

[0259] For example, as shown in FIG20 , the chip system provided in some embodiments of the present application may include at least one processor 1101 and at least one interface circuit 1102. The processor 1101 may be the processor in the terminal device 200 described above. The processor 1101 and the interface circuit 1102 may be interconnected via a line. The processor 1101 may receive and execute computer instructions from the memory of the terminal device 200 described above via the interface circuit 1102. When the computer instructions are executed by the processor 1101, the terminal device 200 may execute the various steps performed by the terminal device 200 in the above-described embodiments. Of course, the chip system may also include other discrete components, which are not specifically limited in some embodiments of the present application.

[0260] For ease of explanation, the above description has been made with reference to specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations are possible. The above embodiments are selected and described to better explain the principles and practical applications, so that those skilled in the art can better utilize the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. A terminal device, comprising: a display configured to display a media resource; A memory is configured to store computer instructions and data associated with the display, and at least one processor is connected to the display and the memory and is configured to execute the computer instructions so that the terminal device performs: a first processor, configured to stop operating when the terminal device is in a standby state; The second processor is configured to: when the terminal device is in a standby state, obtain voice information; if the voice information includes first voice information, trigger the first processor to start, and determine whether the voice information includes second voice information other than the first voice information; the first voice information is used to indicate a wake-up word; The first processor is configured to: if the voice information includes the second voice information, control the display to display the media resource indicated by the second voice information according to the second voice information.

2. According to the terminal device according to claim 1, the second processor is further configured to: determine whether the voice information includes the first voice information; if the voice information does not include the first voice information, determine whether the next voice information includes the first voice information; the next voice information is obtained by the second processor after the voice information.

3. The terminal device according to claim 2, wherein the terminal device comprises a communicator; The first processor is specifically configured to: acquiring the second voice information from the second processor; controlling the communicator to send the second voice information to the server; controlling the communicator to receive a voice recognition result of the second voice information sent by the server; According to the voice recognition result, the media resource indicated by the second voice information is determined, and the display is controlled to display the media resource indicated by the second voice information.

4. The terminal device according to claim 3, wherein the second processor is further configured to: generate a first identifier if the voice information includes the second voice information; the first identifier indicating that the voice information includes the second voice information; The first processor is further configured to: send a first query request to the second processor; The second processor is further configured to: send the first identifier to the first processor in response to the first query request sent by the first processor; The first processor is further configured to: send a second query request to the second processor according to the first identifier sent by the second processor; The second processor is further configured to: send the second voice information to the first processor in response to the second query request sent by the first processor.

5. The terminal device according to claim 1, further comprising: a sound collector for collecting the voice information; the second processor includes a first sub-processor and a second sub-processor; The first sub-processor is configured to: In response to voice information input by a user, extracting voice feature values ​​from the voice information, and Sending the voice feature value to the second sub-processor, where the voice feature value is a spectrum feature containing the wake-up word, and the spectrum feature is obtained by processing the voice information through voice signal processing; The second sub-processor is configured to: In response to the voice feature value sent by the first sub-processor, the voice feature value is verified, and when the verification succeeds, it is determined that the voice information includes the first voice information.

6. The terminal device according to claim 5, after the first sub-processor performs the step of extracting speech feature values ​​from the speech information, further configured to: sending an opening instruction to the second sub-processor, and sending the voice feature value to the second sub-processor in response to an opening receipt signal fed back by the second sub-processor; The second sub-processor is further configured to: In response to the start instruction sent by the first sub-processor, the system enters a start state, and feeds back a start receipt signal to the first sub-processor after the start is completed.

7. The terminal device according to claim 5, wherein the first sub-processor is further configured to: Acquiring frequency spectrum features of the voice information, and caching the frequency spectrum features; Detecting a wake-up state, where the wake-up state is a recognition result of a wake-up word recognition performed on the spectrum feature; the wake-up state includes a success or failure of the wake-up word recognition; If the wake-up state is that the wake-up word recognition is successful, rolling back and locating the spectrum features containing the wake-up word to obtain the speech feature value; If the wake-up state indicates that the wake-up word recognition fails, the spectrum feature is filtered.

8. The terminal device according to claim 7, wherein the first sub-processor performs wake-up word recognition on the spectrum feature, and is further configured to: Inputting the spectral feature into a first wake-up model to obtain a wake-up value of the spectral feature output by the first wake-up model, wherein the wake-up value is used to represent the probability of recognizing the wake-up word; If the wake-up value is greater than or equal to the wake-up threshold, determining that the wake-up state is a successful wake-up word recognition; If the wake-up value is less than the wake-up threshold, it is determined that the wake-up state is a wake-up word recognition failure.

9. The terminal device according to claim 5, wherein the first sub-processor is further configured to: Acquiring a voice signal of the voice information; Splitting the speech signal into a plurality of frame audio data segments; Calculating a power spectrum of the frame audio data segment, where the power spectrum is the energy of a spectrum line corresponding to the frame audio data segment converted from the time domain to the frequency domain; Inputting the power spectrum into a preset filter to obtain a spectrum graph; According to the spectrum diagram, spectrum features are obtained.

10. The terminal device according to claim 9, wherein the first sub-processor is further configured to: acquire spectrum features according to the spectrum graph; Performing a logarithmic operation on the spectrum graph to obtain a logarithmic spectrum domain; A discrete cosine transform is performed on the logarithmic spectrum domain to obtain a spectrum feature.

11. The terminal device according to claim 9, wherein the first sub-processor performs the step of splitting the voice signal into a plurality of frame audio data segments, and is further configured to: Performing pre-emphasis processing on the voice signal, wherein the pre-emphasis processing is used to amplify the high frequency band of the voice signal; The voice signal is split into a plurality of frame audio data segments arranged in sequence according to the formation timing of the voice signal, wherein: Two adjacent frame audio data segments contain overlapping areas; A windowing process is performed on the frame audio data segments, where the windowing process is used to increase the continuity between the plurality of frame audio data segments.

12. The terminal device according to claim 5, wherein the second sub-processor verifies the voice feature value and is further configured to: Obtaining the speech feature value; The speech feature value is input into the second wake-up model to obtain a verification result of the speech feature value output by the second wake-up model, where the verification result includes verification success and verification failure.

13. A voice control method, applied to a terminal device, wherein the terminal device includes a second processor, and the second processor operates when the terminal device is in a standby state; the method comprising: When the terminal device is in the standby state, the second processor acquires voice information; If the voice information includes first voice information, the terminal device is powered on and determines whether the voice information includes second voice information other than the first voice information; the first voice information is used to indicate a wake-up word; If the voice information includes the second voice information, the terminal device displays the media resource indicated by the second voice information according to the second voice information.

14. The method according to claim 13, further comprising: The second processor determines whether the voice information includes the first voice information; If the voice information does not include the first voice information, the second processor determines whether the next voice information includes the first voice information; the next voice information is obtained by the second processor after the voice information.

15. The method according to claim 14, wherein the terminal device further comprises a first processor; the first processor stops working when the terminal device is in a standby state; The terminal device is powered on, and determining second voice information other than the first voice information from the voice information, including: The second processor triggers the first processor to start, and the second processor determines whether the voice information includes second voice information other than the first voice information.

16. The method according to claim 15, wherein the terminal device displays the media resource indicated by the second voice information according to the second voice information, comprising: The first processor acquires the second voice information from the second processor; The first processor sends the second voice information to the server; The first processor receives a speech recognition result of the second voice information sent by the server; The first processor determines the media resource indicated by the second voice information based on the voice recognition result and displays the media resource.

17. The method according to claim 16, wherein the first processor obtains the second voice information from the second processor, comprising: If the voice information includes the second voice information, the second processor generates a first identifier; The first identifier indicates that the voice information includes the second voice information; The first processor sends a first query request to the second processor; The second processor sends the first identifier to the first processor in response to the first query request sent by the first processor; The first processor sends a second query request to the second processor according to the first identifier sent by the second processor; The second processor sends the second voice information to the first processor in response to the second query request sent by the first processor; The first processor receives the second voice information sent by the second processor.

18. The method according to claim 13, wherein the terminal device further comprises a sound collector; the sound collector is used to collect voice information; the second processor comprises a first sub-processor and a second sub-processor; after acquiring the voice information and before the voice information includes the first voice information, the method further comprises: The first sub-processor extracts a voice feature value from the voice information in response to voice information input by the user, and sends the voice feature value to the second sub-processor, wherein the voice feature value is a spectrum feature containing the wake-up word, and the spectrum feature is obtained by processing the voice information through voice signal processing; The second sub-processor verifies the voice feature value in response to the voice feature value sent by the first sub-processor, and determines that the voice information includes the first voice information when the verification succeeds.