A detection method for terminal devices and custom wake words.

By generating and scaling spectrograms to train a wake word recognition model, the problem of low recognition rate of custom wake words in terminal devices is solved, achieving more efficient wake word detection and improving user experience.

CN116259309BActive Publication Date: 2025-10-28HISENSE VISUAL TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211610721.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-14
Publication Date
2025-10-28
Estimated Expiration
2042-12-14

AI Technical Summary

Technical Problem

The low recognition rate of custom wake words in terminal devices leads to a degraded user experience, as existing wake word recognition models have failed to effectively learn user-defined wake word data.

Method used

By collecting user-input voice data, analyzing the voice signal and extracting the acoustic features of the wake word, a spectrogram is generated. The spectrogram is then scaled based on a preset imaging focal length to obtain wake word training data, and a wake word recognition model is trained.

Benefits of technology

It improves the detection efficiency of wake words, enhances the wake word recognition capability of terminal devices in an inactive state, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116259309B_ABST
    Figure CN116259309B_ABST
Patent Text Reader

Abstract

This application provides a terminal device and a method for detecting custom wake words in some embodiments. The method responds to user-inputted voice interaction commands and acquires voice data. It then parses the voice signal in the voice data and extracts wake word target features from the voice signal. The target features are acoustic features of the voice signal containing the wake word. The method generates a spectrogram of the voice signal based on the target features and scales the spectrogram based on a preset imaging focal length to obtain wake word training data. The wake word training data is then used to train a wake word recognition model, enabling the wake word recognition model to learn more acoustic features and improve wake word detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and in particular to a terminal device and a method for detecting a custom wake word. Background Technology

[0002] Terminal devices refer to electronic devices with sound acquisition capabilities, such as smart TVs, mobile phones, smart speakers, computers, and robots. Taking smart TVs as an example, smart TVs are television products based on Internet application technology, equipped with open operating systems and chips, and featuring voice recognition modules, enabling two-way human-computer interaction to meet diverse and personalized user needs.

[0003] Users can also wake up their terminal devices from standby mode using voice commands, thus bringing the device from standby to running mode. Voice wake-up, also known as keyword detection, is an important branch of speech recognition technology. Voice wake-up refers to the terminal device detecting specific keywords from a continuous stream of speech, emitting a signal when a specific keyword is detected, and thus waking the device. The specific keyword is called the wake-up word. Users can customize wake-up words in their terminal devices and wake them up by speaking with the wake-up word.

[0004] However, since the wake word is user-defined, the wake word recognition model built into the terminal device has not learned from the data of the user-defined wake word. Therefore, when the terminal device uses a general wake word recognition model to recognize the user's voice input, it may fail to detect the wake word due to mismatches in some acoustic features, resulting in a low wake word recognition rate and a reduced user experience. Summary of the Invention

[0005] This application provides a method for detecting terminal devices and custom wake words to solve the problem of low wake word recognition rate in terminal devices.

[0006] In a first aspect, some embodiments of this application provide a terminal device, including a detector and a controller. The detector is configured to collect voice data input by a user, and the controller is configured to execute the following program steps:

[0007] Responding to user-inputted voice interaction commands, acquire voice data;

[0008] The speech signal in the speech data is analyzed, and the wake word target feature in the speech signal is extracted. The target feature is the acoustic feature of the speech signal containing the wake word.

[0009] Generate a spectrogram of the speech signal based on the target features;

[0010] The spectrogram is scaled based on a preset imaging focal length to obtain wake word training data;

[0011] The wake word recognition model is trained using the wake word training data, and the wake word recognition model is used to recognize the wake word in the voice data when the terminal device is in an unwake-up state.

[0012] Secondly, some embodiments of this application also provide a method for detecting a custom wake word, including:

[0013] Responding to user-inputted voice interaction commands, acquire voice data;

[0014] The speech signal in the speech data is analyzed, and the wake word target feature in the speech signal is extracted. The target feature is the acoustic feature of the speech signal containing the wake word.

[0015] Generate a spectrogram of the speech signal based on the target features;

[0016] The spectrogram is scaled based on a preset imaging focal length to obtain wake word training data;

[0017] The wake word recognition model is trained using the wake word training data. The wake word recognition model is used to recognize the wake word in the voice data when the terminal device is in an unwakeable state.

[0018] As can be seen from the above technical solutions, the terminal device and custom wake-up word detection method provided in some embodiments of this application can respond to user-inputted voice interaction commands and acquire voice data. The voice signal in the voice data is then parsed, and wake-up word target features are extracted from the voice signal. The target features are the acoustic features of the voice signal containing the wake-up word. The method can generate a spectrogram of the voice signal based on the target features, and scale the spectrogram based on a preset imaging focal length to obtain wake-up word training data. The wake-up word training data is then used to train a wake-up word recognition model, enabling the wake-up word recognition model to learn more acoustic features and improve the detection efficiency of wake-up words. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A schematic diagram illustrating interactive scenarios for voice interaction on a terminal device provided in some embodiments of this application;

[0021] Figure 2 This is a schematic diagram of the hardware configuration of a terminal device provided in some embodiments of this application;

[0022] Figure 3 This is a schematic diagram of the software configuration of the control device provided in some embodiments of this application;

[0023] Figure 4 A schematic diagram of a network architecture for voice interaction provided for some embodiments of this application;

[0024] Figure 5 This application provides schematic diagrams illustrating the process of waking up a terminal device via a wake word voice, as shown in some embodiments.

[0025] Figure 6 A schematic flowchart illustrating the calculation of the registration wake-up score provided for some embodiments of this application;

[0026] Figure 7 A schematic diagram illustrating the effect of a prompt screen on a terminal device when a custom wake word is used, as provided in some embodiments of this application;

[0027] Figure 8 This is a schematic diagram illustrating the effect of another prompt screen on a terminal device when a custom wake word is used, as provided in some embodiments of this application.

[0028] Figure 9 A schematic diagram illustrating the process of generating feature vectors by combining identity authentication vectors for some embodiments of this application;

[0029] Figure 10 A flowchart illustrating a custom wake word detection method provided in some embodiments of this application;

[0030] Figure 11 A schematic diagram illustrating a scenario of performing a short-time Fourier transform on a terminal device provided in some embodiments of this application;

[0031] Figure 12 A schematic flowchart illustrating the generation of spectrograms for speech signals provided in some embodiments of this application;

[0032] Figure 13 Example diagrams of spectrograms provided for some embodiments of this application;

[0033] Figure 14a A schematic diagram showing the spectrogram and lens distance provided in some embodiments of this application, which is greater than one lens focal length and less than two lens focal lengths;

[0034] Figure 14b A schematic diagram showing a spectrogram provided for some embodiments of this application, with the lens distance equal to twice the focal length of the lens;

[0035] Figure 14cA schematic diagram showing a spectrogram and a lens distance greater than twice the focal length provided in some embodiments of this application;

[0036] Figure 15 A schematic diagram illustrating the effect of waking up a terminal device in operation, provided in some embodiments of this application;

[0037] Figure 16 This is a schematic diagram of the structure of a terminal device provided in some embodiments of this application. Detailed Implementation

[0038] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.

[0039] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0040] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0041] Figure 1 An exemplary system architecture is shown that can be used by the terminal device 200 of this application to perform voice interaction. Figure 1 As shown, 10 is a server and 200 is a terminal device, including (smart TV 200a, mobile device 200b, smart speaker 200c).

[0042] In this application, server 10 and terminal device 200 communicate data through various communication methods. Terminal device 200 can be connected via local area network (LAN), wireless local area network (WLAN), and other networks. Server 10 can provide various content and interactive features to terminal device 200. For example, terminal device 200 and server 10 can send and receive information, and receive software updates.

[0043] Server 10 can be a server that provides various services, such as a backend server that supports audio data collected by terminal device 200. The backend server can analyze and process the received audio and other data, and feed back the processing results (such as endpoint information) to the terminal device. Server 10 can be a server cluster or multiple server clusters, and can include one or more types of servers.

[0044] The terminal device 200 can be either hardware or software. When the terminal device 200 is hardware, it can be various electronic devices with sound acquisition capabilities, including but not limited to smart speakers, smartphones, televisions, tablets, e-book readers, smartwatches, media players, computers, AI devices, robots, smart vehicles, etc. When the terminal device 200 is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software programs or software modules (e.g., used to provide sound acquisition services) or as a single software program or software module. No specific limitations are imposed here.

[0045] It should be noted that the custom wake word detection method provided in this application embodiment can be executed by server 10, by terminal device 20, or by both server 10 and terminal device 20. This application does not limit this.

[0046] Figure 2 A hardware configuration block diagram of a terminal device 200 according to an exemplary embodiment is shown. For example... Figure 2 The terminal device 200 shown includes at least one of the following: a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface 280. The controller includes a central processing unit, an audio processor, a graphics processor, RAM, ROM, and a first to an nth interface for input / output.

[0047] The communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of the following: a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The terminal device 200 can establish the transmission and reception of control signals and data signals through the communicator 220 and the server 10.

[0048] The user interface can be used to receive external control signals.

[0049] Detector 230 is used to collect signals from the external environment or to interact with the external environment. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0050] The sound acquisition device can be a microphone, also known as a "microphone" or "voice transducer," which can be used to receive the user's voice and convert the sound signal into an electrical signal. The terminal device 200 can be equipped with at least one microphone. In some embodiments, the terminal device 200 can be equipped with two microphones, which, in addition to acquiring sound signals, can also perform noise reduction. In other embodiments, the terminal device 200 can also be equipped with three, four, or more microphones, enabling sound signal acquisition, noise reduction, sound source identification, and directional recording functions, etc.

[0051] Furthermore, the microphone can be built into the terminal device 200, or it can be connected to the terminal device 200 via wired or wireless means. Of course, this embodiment does not limit the location of the microphone on the terminal device 200. Alternatively, the terminal device 200 may not include a microphone, meaning the microphone is not located within the terminal device 200. The terminal device 200 can connect an external microphone (also called a microphone) via an interface (such as a USB interface). This external microphone can be fixed to the terminal device 200 using an external fastener (such as a camera bracket with a clip).

[0052] The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the terminal device 200.

[0053] For example, the controller includes at least one of a central processing unit (CPU), an audio processor, RAM (random access memory), ROM (read-only memory), a first to an nth interface for input / output, a communication bus, etc.

[0054] In some examples, the operating system of terminal device 200 is Android, such as... Figure 3 As shown, the terminal device 200 can be logically divided into an application layer (referred to as "application layer") 21, a kernel layer 22, and a hardware layer 23.

[0055] Among them, such as Figure 3As shown, the hardware layer may include Figure 2 The controller 250, communicator 220, detector 230, etc., are shown. Application layer 21 includes one or more applications. These applications can be system applications or third-party applications. For example, application layer 21 includes a voice recognition application, which can provide a voice interaction interface and services to enable the connection between the smart TV 200-1 and the server 10.

[0056] The kernel layer 22 serves as a software middleware between the hardware layer and the application layer 21, and is used to manage and control hardware and software resources.

[0057] In some examples, kernel layer 22 includes a detector driver that sends voice data collected by detector 230 to a speech recognition application. For instance, when the speech recognition application in terminal device 200 starts and a communication connection is established between terminal device 200 and server 10, the detector driver sends user-input voice data collected by detector 230 to the speech recognition application. The speech recognition application then sends query information containing this voice data to intent recognition module 202 in the server. Intent recognition module 202 inputs the voice data sent by terminal device 200 into the intent recognition model.

[0058] To clearly illustrate the embodiments of this application, the following description is provided in conjunction with... Figure 4 This application describes a speech recognition network architecture provided in its embodiments.

[0059] See Figure 4 , Figure 4 This is a schematic diagram of a voice interaction network architecture provided in an embodiment of this application. Figure 4 In this system, the intelligent device receives input information and outputs the processing results of that information. The speech recognition module deploys a speech recognition service to convert audio into text; the semantic understanding module deploys a semantic understanding service to perform semantic parsing on the text; the business management module deploys a business instruction management service to provide business instructions; the language generation module deploys a language generation service (NLG) to convert instructions instructing the intelligent device to execute into text language; and the speech synthesis module deploys a text-to-speech (TTS) service to process the text language corresponding to the instructions and send it to a speaker for playback. In one embodiment, Figure 4 The architecture shown can contain multiple entity service devices with different business services deployed, or one or more entity service devices can combine one or more functional services.

[0060] In some embodiments, the following describes the basis Figure 4 The process of processing input information from smart devices in the architecture shown is illustrated with an example, taking a query statement input via voice as an example:

[0061] [Speech Recognition]

[0062] After receiving a query statement input via voice, the smart device can perform noise reduction and feature extraction on the audio of the query statement. The noise reduction process may include steps such as removing echoes and environmental noise.

[0063] [Semantic understanding]

[0064] Using acoustic and language models, natural language understanding is performed on the identified candidate text and associated contextual information. The text is parsed into structured, machine-readable information, including business domain, intent, slots, and other semantic information. An executable intent confidence score is obtained, and the semantic understanding module selects one or more candidate executable intents based on the determined intent confidence score.

[0065] [Business Management]

[0066] Based on the semantic parsing results of the query statement text, the semantic understanding module sends query instructions to the corresponding business management module to obtain the query results provided by the business service, as well as the actions required to "complete" the user's final request, and feeds back the device execution instructions corresponding to the query results.

[0067] It should be noted that, Figure 4 The architecture shown is merely an example and is not intended to limit the scope of protection of this application. Other architectures can also be used to achieve similar functions in the embodiments of this application. For example, all or part of the above process can be completed by the terminal device 200, which will not be elaborated here.

[0068] Based on the aforementioned terminal device 200, users can interact with the terminal device 200 via voice commands. Furthermore, in some embodiments, the terminal device 200, while in standby mode, also receives and parses voice data input by the user. When the voice data contains a wake-up word, the standby state of the terminal device 200 is switched to the running state. That is, the user can also wake up the terminal device 200 from standby mode using a wake-up word.

[0069] For example, taking terminal device 200 as a display device, and the display device is equipped with voice interaction function, the wake-up word is "Hi ABC". The display device is powered on but the screen is off. At this time, the display device is in standby mode. Figure 5 As shown, the user says "Hi ABC" to the display device, which then receives and recognizes the voice prompt. Upon recognizing the wake-up word "Hi ABC," the display device switches from standby mode to running mode to wake up.

[0070] Obviously, a user can only wake up the terminal device 200 by using a wake word if a wake word has been configured in the terminal device 200. In some embodiments, the wake word can be a fixed keyword configured in the terminal device 200, or a keyword customized by the user in the terminal device 200. If the wake word is a fixed keyword, a wake word recognition model is trained based on the fixed keyword, and the user's input voice data is recognized by the trained wake word recognition model; if the wake word is a customized keyword, the user's input voice data is recognized by a general wake word recognition model.

[0071] Furthermore, since custom wake words are keywords freely set by the user, the user needs to pre-register the custom wake word in the terminal device 200. Therefore, in some embodiments, such as Figure 6 As shown, the terminal device 200 receives a user-inputted custom wake-up word command and, in response, acquires wake-up word voice data within a preset time period. It then calculates a registration wake-up score for the wake-up word voice data using a general wake-up model. This general wake-up model is the initial wake-up word recognition model configured in the terminal device 200. If the registration wake-up score is greater than or equal to a preset score threshold, the wake-up word recognition model is trained based on the wake-up word voice data; otherwise, the wake-up word voice data is filtered out.

[0072] In other words, users can input a custom wake-up word command into terminal device 200 to trigger it to enter custom wake-up word mode. Once in custom wake-up word mode, terminal device 200 collects voice data within a preset time period to gather a registration template for the wake-up word. However, if the user's speech during this time period is unclear, the registration template collected by terminal device 200 will also have issues. Therefore, because the wake-up word is problematic from the outset, users will experience low wake-up word recognition rates when using terminal device 200. Consequently, terminal device 200 also calculates a registration wake-up score for the wake-up word voice data using a general wake-up model. If the wake-up score is greater than or equal to the wake-up threshold, it indicates that the collected voice data has a high signal-to-noise ratio and can be used as a wake-up word template. If the wake-up score is less than the wake-up threshold, it indicates that the collected voice data has a low signal-to-noise ratio, and this segment of voice data is filtered out.

[0073] Understandably, the scoring threshold can be set based on the number of false wake-ups. That is, the higher the scoring threshold, the fewer times the terminal device will be falsely woken up.

[0074] In some embodiments, when collecting wake-word voice data over a preset time period, the terminal device 200 also monitors user operation events. If the user performs a confirmation operation on the terminal device 200 within the preset time period, the voice data collection is stopped. That is, the user can confirm and stop the terminal device 200's voice collection at any time within the preset time period, thereby reducing the acquisition of unnecessary voice data by the terminal device 200 and preventing resource waste.

[0075] For example, taking terminal device 200 as the display device, the preset time period is 60 seconds. Figure 7 As shown, the user sends a custom wake-up word command to the display device via the remote control. The display device then responds to the button presses on the remote control and displays... Figure 7 The prompt screen shown indicates that the user has 45 seconds to input a wake-up word into the display device. This means that when the user finishes speaking the keyword, there are 15 seconds remaining in the preset time slot on the display device. At this point, the user can click... Figure 7 The confirmation control in the display device ends the acquisition of the wake-up word voice data. Therefore, when processing the wake-up word voice data, the display device reduces the processing of unnecessary voice segments, thereby reducing the consumption of system resources within the display device.

[0076] It is understandable that the above example exemplifies the execution process of terminal device 200 when the user performs a confirmation operation within a preset time period. Obviously, if the user does not perform a confirmation operation within the preset time period, terminal device 200 will still automatically end the collection of wake-word voice data according to the preset time period.

[0077] Furthermore, the more syllables a wake word covers and the greater the syllable differences, the better the wake word detection performance of the terminal device 200. Therefore, in some embodiments, after entering the custom wake word mode, the terminal device 200 can also prompt the user with wake word examples, so that the user can set a better wake word in the terminal device 200 based on the wake word examples.

[0078] For example, let's take terminal device 200 as a display device. Figure 7 As shown, after the user enters the custom wake word mode of the display device, the display device also displays a wake word example "hi ABC" to the user, who can refer to the wake word example "hi ABC" to set the corresponding wake word.

[0079] It should be noted that the above embodiments use a display device as an example for explanation, but the same applies to other electronic devices with voice interaction functions. When the terminal device is an electronic device without a display function, the displayed screen information can be output in the form of voice to prompt the user to perform a response.

[0080] To optimize the wake-up word registration template collected by the terminal device 200, in some embodiments, when training the wake-up word recognition model based on the wake-up word speech data, the terminal device 200 further decomposes the wake-up word speech data into multiple speech segments. Furthermore, it extracts target features from each speech segment and generates a registration speech feature template based on the target features. The registration speech feature template is then input into the wake-up word recognition model as the wake-up word template. Here, the target features are the acoustic features of the speech data. By segmenting the collected wake-up word speech, the quality of the wake-up word registration template can be improved.

[0081] For example, the initial wake-up word recognition model configured in terminal device 200 is a model that determines whether speech and text match. When a user defines a wake-up word, terminal device 200 can first receive the wake-up word text set by the user, and then collect the wake-up word speech input by the user. Simultaneously, it calculates the corresponding score threshold based on the wake-up word text. The wake-up word speech is then input into the initial wake-up word recognition model to calculate the registered wake-up score. When the registered wake-up score is greater than or equal to the score threshold, the wake-up word speech is processed. That is, terminal device 200 aligns the wake-up word speech data with the wake-up word text, and then extracts the speech segments corresponding to each character in the wake-up word text for target feature extraction.

[0082] Understandably, in order to improve the recognition rate of wake words, in some embodiments, when the terminal device 200 collects the wake word registration voice input by the user, it can remind the user to input the wake word registration voice multiple times, so as to generate more training data and improve the recognition rate of the wake word recognition model.

[0083] For example, let's take terminal device 200 as a display device. Figure 7 As shown, the user sends a custom wake-up word command to the display device via the remote control. The display device then responds to the button presses on the remote control and displays... Figure 7 The prompt screen shown. After the user speaks the wake-up word once according to the prompt information on the display device and clicks confirm, the display device completes the first collection of wake-up word voice data. Then it displays... Figure 8 The prompt screen shown prompts the user to repeat the wake word speech so that the display device can extract more training data for the wake word recognition model.

[0084] Furthermore, to reduce the false wake-up rate of the terminal device 200, the terminal device 200 can also add the user's voice characteristics to the wake-up word recognition model. That is, in some embodiments, such as Figure 9As shown, terminal device 200 extracts wake-up word speech data based on a speech endpoint detection algorithm and performs noise reduction processing on the wake-up word speech data. Then, it extracts acoustic features and authentication vectors from the wake-up word speech data. The acoustic features and authentication vectors are then concatenated into a feature vector, and the wake-up word recognition model is trained using this feature vector. Therefore, the trained wake-up word recognition model carries speaker information and has more detailed feature information, which can reduce the misrecognition by terminal device 200.

[0085] For example, when a user inputs a wake-up word, the terminal device 200 first extracts the wake-up word speech data through a VAD (Voice Endpoint Detection) during training data extraction. Then, it performs noise reduction on the wake-up word speech data. The acoustic layer features and authentication vector of the denoised data are extracted separately, and then concatenated into a single feature vector. Based on this feature vector, an HMM-DNN (Hidden Markov Model-Depth Neural Network) is trained to obtain the user-related SD-WUW (Wake-up Model).

[0086] Clearly, by combining authentication vectors, terminal device 200 can recognize only the speaker's voice during the wake word registration phase. However, some terminal devices 200 are shared electronic devices used by multiple users, such as smart TVs. Therefore, authenticating only one authentication vector makes it inconvenient to use the wake word function of terminal device 200. Thus, in some embodiments, when customizing the wake word, terminal device 200 can provide an authentication selection function, allowing users to choose multiple authentication methods. This enables terminal device 200 to extract multiple authentication vectors and recognize the voices of users with multiple identities.

[0087] Furthermore, to facilitate user interaction with the terminal device 200, in some embodiments, the terminal device 200 can also acquire various control commands input by the user. Among these, some control commands that enable the terminal device 200 to enter a custom wake-up word mode are called custom wake-up word commands. Custom wake-up commands can be input via physical buttons or voice commands. For example, when the terminal device 200 is a display device, a custom wake-up word command can be input to the display device via a remote control or other control device, or via a preset voice command.

[0088] As can be seen from the above embodiments, users can wake up terminal device 200 based on a custom wake-up word. However, since the custom wake-up word is randomly set by the user, terminal device 200 has not pre-trained a wake-up word recognition model for the custom wake-up word. Furthermore, the amount of wake-up word speech input by the user when creating a custom wake-up word is limited, resulting in insufficient training data for the wake-up word recognition model. Consequently, when terminal device 200 recognizes the user's input speech data, it is prone to misrecognition or recognition failure, leading to a low wake-up word recognition rate and a reduced user experience.

[0089] Based on the above application scenarios, in order to improve user experience and alleviate the problem of low wake-up word recognition rate in terminal devices 200, some embodiments of this application provide a method for detecting custom wake-up words. For example... Figure 10 As shown, the method specifically includes the following steps:

[0090] S100: Responds to user-inputted voice interaction commands and acquires voice data.

[0091] Voice interaction commands refer to the process where a user speaks to terminal device 200, and terminal device 200 receives the voice data from the user to complete the voice data acquisition, and then generates the corresponding voice interaction command. In other words, terminal device 200 can collect the sounds emitted by the user within its acquisition range and convert the collected sounds into electrical signals for storage, thus generating voice data. For example, if the user says "Hi ABC" within the acquisition range of terminal device 200, terminal device 200 will convert "Hi ABC" into electrical signal data for storage.

[0092] S200: Analyzes speech signals in speech data and extracts wake word target features from speech signals.

[0093] After acquiring the voice data, the terminal device 200 parses the voice signal in the voice data and extracts the wake-up word target feature from the voice signal. The target feature is the acoustic feature of the voice signal containing the wake-up word.

[0094] In some embodiments, the acoustic features include FFT (fast Fourier transform), pitch (fundamental frequency), MFCC (Mel Frequency Cepstrum Coefficient), Fbank (Filter Bank), and PCEN (Per-channel energy normalization), etc.

[0095] However, as Figure 11As shown, standard Fourier analysis is required for stationary global signals. However, the terminal device 200 typically acquires non-stationary local speech signals. Therefore, standard Fourier analysis cannot be used to extract target features.

[0096] Therefore, in some embodiments, such as Figure 11 As shown, when extracting wake-word target features from a speech signal, the terminal device 200 performs frame segmentation and windowing function processing, such as Hamming windowing, on the speech signal. Then, it extracts target features for each frame of the speech signal based on cepstral domain features, including Mel-frequency cepstral coefficients (MFCCs) and linear predictive cepstral coefficients (LMCCs). A short-time Fourier transform is then performed on the target features. The analysis of Mel-frequency cepstral coefficients (MFCCs) is based on human auditory mechanisms. That is, the MFCCs can be used to analyze the spectrum of the speech signal based on human auditory experimental results to obtain good speech characteristics. Therefore, the features extracted through MFCCs and LMCCs have uniqueness, which can reduce the time and frequency domain losses when the terminal device 200 processes the speech signal.

[0097] Since speech signals are one-dimensional signals based on the time axis, for the purpose of speech signal analysis, it is assumed that the speech signal is in a stable state for a short period of time at the millisecond level, and frame segmentation is performed on this basis. In some embodiments, frame segmentation can be carried out by continuous segmentation, but in order to make the transition between frames smooth and thus maintain the continuity of the speech signal, overlapping segmentation can be used. Here, frame segmentation is implemented by weighting with movable, finite-length windows, that is, by multiplying a certain window function w(n) with the speech signal s(n) to form a windowed speech signal.

[0098] In some embodiments, the formula for the Short-Time Fourier Transform (STFT) is as follows:

[0099]

[0100] Where n is time, W(nk) is the window sequence, and f is the frequency of the speech signal. Terminal device 200 can convert the speech signal based on the above formula.

[0101] S300: Generates a spectrogram of the speech signal based on the target features.

[0102] After extracting the target features of the speech signal, the terminal device 200 can process the speech signal based on the target features to generate a spectrogram corresponding to the speech signal.

[0103] For example: Figure 12As shown, the terminal device 200 preprocesses the parsed speech signal, namely, it segments the speech signal into frames, applies windows, and parses the time-domain signal of the speech signal. Then, it performs a short-time Fourier transform on the processed time-domain signal to obtain the spectrogram corresponding to the speech signal, thereby reducing the time and frequency domain losses in the process of processing the speech signal by the terminal device 200.

[0104] In some embodiments, after performing a short-time Fourier transform on the target features, the terminal device 200 also calculates the energy density of the target features and plots the axes of a spectrogram. The horizontal axis of the axes represents the time of the speech signal, and the vertical axis represents the frequency of the speech signal. The speech signal is then plotted on the axes based on the energy density to obtain the spectrogram. The energy density is the ratio of signal power to signal frequency.

[0105] In other words, a spectrogram can also be understood as a spectrum analysis view. For example... Figure 13 As shown, Figure 13 This is an example of a spectrogram. (From...) Figure 13 As can be seen, the horizontal axis of a spectrogram represents time, the vertical axis represents frequency, and the coordinate point value represents the energy of the speech data. Since a spectrogram represents three-dimensional information in a two-dimensional plane, the energy value of the coordinate point is represented by color. That is, the darker the color, the stronger the speech energy at that point; the lighter the color, the weaker the speech energy at that point.

[0106] It should be noted that the spectrogram generation method provided in this application embodiment is merely illustrative, and other methods can also be used to generate spectrograms from speech data. This application does not impose any restrictions on this.

[0107] S400: Obtain wake word training data by scaling the spectrogram based on the preset imaging focal length.

[0108] After generating the spectrogram, the terminal device 200 can use it as training data for the wake-word recognition model. Since a single spectrogram results in limited training data, the terminal device 200 can also scale the spectrogram based on the principle of image augmentation to increase the amount of training data. In other words, each time the terminal device 200 scales the spectrogram, it generates a new spectrogram, thereby acquiring more training data.

[0109] In some embodiments, when scaling the spectrogram based on a preset imaging focal length, the terminal device 200 adjusts the distance between the spectrogram and the lens according to the principle of convex lens imaging to generate sub-spectral maps, which are spectrograms of different scales. The scale of the sub-spectral maps is then normalized to serve as the training data. That is, the terminal device 200 can adjust the distance between the spectrogram and the lens according to the principle of convex lens imaging. Different distances between the spectrogram and the lens result in different sizes of the imaged spectrograms.

[0110] A convex lens converges light, and the image it produces on a screen is called a real image. By placing the spectrogram beyond the focal length of the convex lens, an inverted real image can be formed on the other side of the lens. Real images can be categorized into three types: reduced, same-size, and magnified. When the distance between the spectrogram and the lens is twice the focal length of the lens, the image formed by the convex lens is the same size as the original spectrogram.

[0111] It is understood that convex lens imaging requires the distance between the spectrogram and the lens to be greater than one focal length in order to form an image. Therefore, the sub-spectral graphs in the embodiments of this application are all real images formed when the distance between the original spectrogram and the convex lens is greater than one focal length.

[0112] To acquire spectrogram images at different scales, in some embodiments, the terminal device 200, while adjusting the distance between the spectrogram and the lens according to the imaging principle of a convex lens, also detects the focal length of the lens. Then, it calculates the imaging focal length based on the lens focal length, which includes focal lengths of twice the lens focal length, less than twice the lens focal length, and greater than twice the lens focal length. Finally, it adjusts the distance between the spectrogram and the lens according to the imaging focal length.

[0113] For example: Figure 14a , Figure 14b and Figure 14c As shown, the focal length of lens F in the figure. The preprocessed spectrogram is placed at point P. (As shown...) Figure 14a As shown, when the distance of point P from the lens is greater than F and less than 2F, the obtained sub-spectral map is larger in scale than the original spectrogram; for example... Figure 14b As shown, when point P is 2F away from the lens, the obtained sub-spectral is the same size as the original spectrogram; as Figure 14c As shown, when the distance between point P and the lens is greater than 2F, the obtained sub-spectral map is smaller in scale than the original spectrogram. Based on the three transformation methods described above, the terminal device 200 adjusts the distance between the spectrogram and the lens to obtain sub-spectral maps of different scales.

[0114] S500: Train the wake word recognition model using wake word training data.

[0115] After generating sub-spectral maps at multiple scales, terminal device 200 normalizes all sub-spectral maps. These sub-spectral maps are then used as training data to train the wake-word recognition model within terminal device 200. This wake-word recognition model is used to identify wake-words in the speech data when terminal device 200 is in an inactive state. Using the normalized sub-spectral maps as input to a deep neural network allows the deep neural network to achieve better convergence, enabling the wake-word recognition model to learn rich acoustic features, such as spectrum, pitch, and formants, thereby improving the detection rate of voice wake-up.

[0116] Based on the above embodiments, the terminal device 200 can recognize the user-inputted voice data through a wake-up word recognition model, thereby waking up the terminal device 200 from its standby state. Therefore, in some embodiments, the terminal device 200 recognizes the voice data through the wake-up word recognition model. If the similarity between the acoustic features of the voice data and the target features of the wake-up word is within a preset similarity range, the terminal device 200 is woken up, and the voice signal in the voice data is parsed to train the wake-up word recognition model; if the similarity between the acoustic features of the voice data and the target features of the wake-up word is not within the preset similarity range, the terminal device 200 remains in an unwakeable state, and the voice data is filtered.

[0117] In other words, the wake-word recognition model of terminal device 200, after being trained on a large amount of data, automatically learns rich acoustic features. Therefore, when recognizing user-inputted speech using the wake-word recognition model, it can determine whether the user-inputted speech data includes a wake-word by calculating the similarity between the currently collected speech data and the speech data in the wake-word recognition model. Furthermore, if it is determined that the user's currently input speech data includes a wake-word, terminal device 200 can also use the currently collected speech data as training data to train the wake-word recognition model, thereby improving the detection rate of the wake-word recognition model.

[0118] It should be noted that the similarity calculation of acoustic features in the embodiments of this application can be done in various ways, and the determination criteria for wake words can also be done in various ways. This application does not impose any restrictions on these methods.

[0119] Furthermore, in some embodiments, the wake-up mode of the terminal device 200 can be either a standby state or a running state. When in standby state, the terminal device 200 responds to the user's input wake-up word voice, switches the terminal device 200 to the running state, and starts the voice question-and-answer function; when in the running state, the terminal device 200 responds to the user's input wake-up word voice, and starts the voice question-and-answer function.

[0120] For example, taking terminal device 200 as the display device, the wake-up word is "Hi ABC". Figure 5 As shown, when the display device is in standby mode, the user says "Hi ABC" to the display device, which then receives and recognizes the voice prompt "Hi ABC". Upon recognizing the wake word "Hi ABC", the display device switches from standby mode to running mode and activates its built-in voice question-and-answer function. Figure 5 The voice assistant shown allows users to interact with the terminal device 200 via voice by answering the voice assistant's question, "How can I help you?"; for example... Figure 15 As shown, when the display device is running, the user says "Hi ABC" to the display device, which then receives and recognizes the voice prompt "Hi ABC". After recognizing the wake word "Hi ABC", the display device activates the voice question-and-answer function. Figure 15 The voice assistant shown allows users to interact with the terminal device 200 by answering the voice assistant's question, "Is there anything I can help you with?".

[0121] Furthermore, the terminal device 200 can also disable the voice wake-up function. That is, after the user sets a custom wake-up word, the voice wake-up function of the terminal device 200 can be disabled, preventing the terminal device 200 from being woken up by the wake-up word. Therefore, in some embodiments, the terminal device 200 also monitors the on / off status of the wake-up word wake-up program. If the wake-up word wake-up program is off, it does not acquire current voice data, making the terminal device 200 more personalized and improving the user experience.

[0122] Based on the above-described method for detecting custom wake words, some embodiments of this application also provide a terminal device 200, such as... Figure 16 As shown, it includes: a detector 230 and a controller 250. The detector 230 is configured to collect voice data input by the user; as... Figure 10 As shown, the controller 250 is configured to perform the following program steps:

[0123] S100: Responds to user-inputted voice interaction commands and acquires voice data;

[0124] S200: Analyze the speech signal in the speech data and extract the wake word target feature in the speech signal, wherein the target feature is the acoustic feature of the speech signal containing the wake word;

[0125] S300: Generate a spectrogram of the speech signal based on the target features;

[0126] S400: Zoom the spectrogram based on a preset imaging focal length to obtain wake word training data;

[0127] S500: Train a wake-up word recognition model using the wake-up word training data. The wake-up word recognition model is used to recognize the wake-up word in the voice data when the terminal device is in an unwake-up state.

[0128] As can be seen from the above technical solutions, the terminal device and custom wake-up word detection method provided in some embodiments of this application can respond to user-inputted voice interaction commands and acquire voice data. The voice signal in the voice data is then parsed, and wake-up word target features are extracted from the voice signal. The target features are the acoustic features of the voice signal containing the wake-up word. The method can generate a spectrogram of the voice signal based on the target features, and scale the spectrogram based on a preset imaging focal length to obtain wake-up word training data. The wake-up word training data is then used to train a wake-up word recognition model, enabling the wake-up word recognition model to learn more acoustic features and improve the detection efficiency of wake-up words.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0130] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

Claims

1. A terminal device, characterized in that, include: The detector is configured to collect voice data input by the user; The controller is configured as follows: Responding to user-inputted voice interaction commands, acquire voice data; The speech signal in the speech data is analyzed, and the wake word target feature in the speech signal is extracted. The target feature is the acoustic feature of the speech signal containing the wake word. Generate a spectrogram of the speech signal based on the target features; The spectrogram is scaled based on a preset imaging focal length to obtain wake word training data; The wake word training data is used to train a wake word recognition model, which is used to recognize wake words in the voice data when the terminal device is in an unwake state. The controller performs scaling of the spectrogram based on a preset imaging focal length and is configured as follows: The distance between the spectrogram and the lens is adjusted according to the imaging principle of a convex lens to generate sub-spectral maps, which are spectrogram maps of different scales. The scale of the sub-language spectrogram is normalized to serve as the training data; The controller is configured to adjust the distance between the spectrogram and the lens according to the imaging principle of a convex lens, and is configured as follows: Detect the focal length of the lens; The imaging focal length is calculated based on the lens focal length, which includes the focal length of a 2x lens, the focal length of a lens less than 2x, and the focal length of a lens greater than 2x. The distance between the spectrogram and the lens is changed according to the imaging focal length.

2. The terminal device according to claim 1, characterized in that, The controller is configured to: Receive user-inputted custom wake word commands; In response to the custom wake word command, acquire wake word voice data within a preset time period; The registration wake-up score of the wake-up word voice data is calculated using a general wake-up model, wherein the general wake-up model is the initial wake-up word recognition model configured in the terminal device; If the registered wake-up score is greater than or equal to a preset score threshold, then the wake-up word recognition model is trained based on the wake-up word speech data; If the registered wake-up score is less than a preset score threshold, the wake-up word voice data is filtered.

3. The terminal device according to claim 2, characterized in that, The controller is configured to train the wake-up word recognition model based on the wake-up word speech data. The wake word speech data is decomposed into multiple speech segments; Extract the target features from the speech segments respectively, and generate a registered speech feature template based on the target features; The registered voice feature template is input into the wake word recognition model.

4. The terminal device according to claim 2, characterized in that, The controller is configured to: The wake word speech data is extracted based on the speech endpoint detection algorithm, and noise reduction processing is performed on the wake word speech data; Extract the acoustic features and authentication vector from the wake word speech data; The acoustic features are concatenated with the identity authentication vector to form a feature vector, and the wake word recognition model is trained using the feature vector.

5. The terminal device according to claim 1, characterized in that, The controller is configured to extract wake word target features from the speech signal. The speech signal is subjected to frame segmentation and windowing function processing; Target features of each frame of the speech signal are extracted based on cepstral domain features, which include Mel frequency cepstral coefficients and linear prediction cepstral coefficients. Perform a short-time Fourier transform on the target features.

6. The terminal device according to claim 5, characterized in that, After the controller performs a short-time Fourier transform on the target feature, it is configured as follows: Calculate the energy density of the target feature and lay out the axes of the spectrogram, where the horizontal axis represents the time of the speech signal and the vertical axis represents the frequency of the speech signal; The speech signal is plotted on the axis based on the energy density to obtain the spectrogram.

7. The terminal device according to claim 1, characterized in that, The controller is configured to: The voice data is identified using the wake word recognition model; If the similarity between the acoustic features of the speech data and the target features of the wake word is within a preset similarity range, then the terminal device is woken up, and the speech signal in the speech data is parsed to train the wake word recognition model. If the similarity between the acoustic features of the speech data and the target features of the wake word is not within a preset similarity range, the terminal device remains in an unwake state, and the speech data is filtered.

8. A method for detecting a custom wake word, characterized in that, include: Responding to user-inputted voice interaction commands, acquire voice data; The speech signal in the speech data is analyzed, and the wake word target feature in the speech signal is extracted. The target feature is the acoustic feature of the speech signal containing the wake word. Generate a spectrogram of the speech signal based on the target features; The spectrogram is scaled based on a preset imaging focal length to obtain wake word training data; The wake word training data is used to train a wake word recognition model, which is used to recognize wake words in the voice data when the terminal device is in an unwake-up state. The scaling of the spectrogram based on a preset imaging focal length includes: The distance between the spectrogram and the lens is adjusted according to the imaging principle of a convex lens to generate sub-spectral maps, which are spectrogram maps of different scales. The scale of the sub-language spectrogram is normalized to serve as the training data; The adjustment of the distance between the spectrogram and the lens based on the imaging principle of a convex lens includes: Detect the focal length of the lens; The imaging focal length is calculated based on the lens focal length, which includes the focal length of a 2x lens, the focal length of a lens less than 2x, and the focal length of a lens greater than 2x. The distance between the spectrogram and the lens is changed according to the imaging focal length.

Citation Information

Patent Citations

  • Speech emotion recognition method based on parameter migration and spectrogram

    CN108597539A

  • Voice wake-up method and electronic equipment

    CN109979438A

  • Keyword detection method capable of supporting self-defined wake-up words

    CN111933124A

  • Law enforcement instant evidence fixing method and law enforcement instrument

    CN113762110A