A voice wake-up method, device, equipment, medium and product
Through the improved combination method of CTC loss function and HMM model, the bispeech wake-up module structure is adopted to solve the problem of the balance between accuracy and calculation amount of speech wake-up method, and achieve efficient speech wake-up effect.
Patent Information
- Application Number
- CN202510147661.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing voice wake-up method is difficult to balance the wake-up effect and the calculation amount, and it is impossible to improve accuracy and reduce the consumption of computing resources at the same time.
Using improved fixed path decoding based on CTC loss function and improved target path recognition diagram based on HMM model, the identification and confirmation of potential targets is performed through the bispeaker wake-up module structure, reducing the computational amount while improving wake-up accuracy.
While ensuring the accuracy of voice wake-up, it reduces the consumption of computing resources, improves voice wake-up efficiency, supports custom wake-up words, and reduces memory and computing load.
Smart Images

Figure CN119993164B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of voice wake-up, and in particular to a voice wake-up method, device, equipment, medium and product. Background Art
[0002] Voice has always been a way of interaction between people. In the human-computer interaction of intelligent hardware such as smart speakers and smart set-top boxes, voice interaction is the most natural way of interaction, which is very convenient and easy to understand. In recent years, the development of artificial intelligence technology has advanced by leaps and bounds. At present, machines can already recognize and understand the internal meaning of voice and make corresponding responses in a timely manner, such as playing voice, calling relevant skills, etc. In this interaction process, the real-time performance and accuracy of voice wake-up greatly affect the user experience and are an important prerequisite for smooth voice interaction.
[0003] The current voice wake-up methods mostly adopt the following three: methods based on the Hidden Markov Model (HMM), methods based on the Cross-Entropy (CE) loss function, and methods based on the Connectionist Temporal Classification (CTC) loss function. However, in actual use, a single method is usually adopted to solve the voice wake-up problem, and it is impossible to balance the wake-up effect and the computational complexity. Summary of the Invention
[0004] The purpose of the present application is to provide a voice wake-up method, device, equipment, medium and product, which can improve the accuracy of voice wake-up while reducing the computational complexity and improving the voice wake-up efficiency.
[0005] To achieve the above purpose, the present application provides the following solutions:
[0006] In the first aspect, the present application provides a voice wake-up method, including:
[0007] Obtaining an original audio signal;
[0008] Performing preprocessing on the original audio signal to obtain a preprocessed audio signal;
[0009] Use the first voice wake-up method to identify potential targets in the preprocessed audio signal, and obtain a first recognition result. Among them, the first voice wake-up method is an improved voice wake-up method based on the CTC loss function. The first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on the CTC loss function to fixed-path decoding. The fixed-path decoding refers to decoding through a target path, and the target path includes the fixed path of the phonemes in the target wake-up word and the fixed path of the phonemes with similar pronunciations to the target wake-up word. The potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciations to the target wake-up word;
[0010] When the first recognition result meets the preset conditions, perform preliminary wake-up, and use the second voice wake-up method to identify whether the target audio frame contains the real potential target, and obtain a second recognition result. Among them, the target audio frame refers to the audio frame that may contain the potential target in the first recognition result. The second voice wake-up method is an improved voice wake-up method based on the HMM model. The second voice wake-up method includes improving the recognition graph in the traditional voice wake-up method based on the HMM model to a target-path recognition graph, and the target-path recognition graph is a recognition graph constructed according to the target path;
[0011] Confirm wake-up according to the second recognition result, and identify the time of each phoneme of the final target wake-up word in the second recognition result.
[0012] In a second aspect, the present application provides a voice wake-up method device, including:
[0013] An acquisition module, configured to acquire an original audio signal;
[0014] A voice energy detection module, configured to preprocess the original audio signal to obtain a preprocessed audio signal;
[0015] A first voice wake-up module, configured to use the first voice wake-up method to identify potential targets in the preprocessed audio signal, and obtain a first recognition result. Among them, the first voice wake-up method is an improved voice wake-up method based on the CTC loss function. The first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on the CTC loss function to fixed-path decoding. The fixed-path decoding refers to decoding through a preset target path, and the target path includes the fixed path of the phonemes in the target wake-up word and the fixed path of the phonemes with similar pronunciations to the target wake-up word. The potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciations to the target wake-up word;
[0016] A second voice wake-up module, configured to use a second voice wake-up method to identify whether a true potential target is included in a target audio frame, so as to obtain a second recognition result, where the target audio frame refers to an audio frame that may include the potential target in the first recognition result, the second voice wake-up method is an improved voice wake-up method based on an HMM model, and the second voice wake-up method includes improving an identification graph in a conventional voice wake-up method based on an HMM model into a target path identification graph, and the target path identification graph is an identification graph constructed according to the target path;
[0017] A final result obtaining module, configured to confirm wake-up according to the second recognition result and identify the time of each phoneme of the final target wake-up word in the second recognition result.
[0018] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the voice wake-up method described in the first aspect above.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the voice wake-up method described in the first aspect above is implemented.
[0020] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the voice wake-up method described in the first aspect above is implemented.
[0021] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0022] The present application provides a voice wake-up method, device, device, medium and product. The method includes: preprocessing an original audio signal; using a first voice wake-up method to perform potential target recognition on the preprocessed audio signal to obtain a first recognition result; using a second voice wake-up method to perform secondary potential target recognition on the first recognition result to obtain a second recognition result; and confirming wake-up when the second recognition result meets a second preset condition, and identifying the time of each phoneme of the final target wake-up word in the second recognition result. Wherein, the first voice wake-up method includes improving a decoding method in a conventional voice wake-up method based on a CTC loss function into a fixed path decoding, and the second voice wake-up method includes improving an identification graph in a conventional voice wake-up method based on an HMM model into a target path identification graph, which can improve the accuracy of voice wake-up while reducing the calculation amount and improving the voice wake-up efficiency. Description of the Drawings
[0023] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0024] Figure 1 Schematic diagram of the functional modules of a voice wake-up device provided in Embodiment 2 of the present application;
[0025] Figure 2 Schematic diagram of the structure of the first voice wake-up module in Embodiment 2 of the present application;
[0026] Figure 3 Schematic diagram of the structure of the second voice wake-up module in Embodiment 2 of the present application;
[0027] Figure 4 DFSMN structure diagram in Embodiment 2 of the present application;
[0028] Figure 5 CTC fixed path score diagram in Embodiment 2 of the present application;
[0029] Figure 6 Schematic diagram of the structure of a computer device provided in Embodiment 3 of the present application. Detailed implementation manners
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0031] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0032] Embodiment 1
[0033] It has been found through research that the voice wake-up method based on the HMM model needs to construct an identification network according to information such as wake-up words, approximate pronunciations of wake-up words, anti-wake-up words, garbage words, or garbage phonemes, and identify the user's voice according to this identification network. The loading and search results of this identification network consume a lot of memory resources and computing resources.
[0034] The voice wake-up method based on the CE loss function transforms the voice wake-up problem from a speech recognition task into an image classification task. It only needs to smooth the posterior probability results of the acoustic model, without the need to construct a recognition network and perform a search. However, its accuracy depends on the pre-annotated wake-word boundary information and does not support the function of custom wake words.
[0035] The voice wake-up method based on the CTC loss function solves the alignment problem by introducing the blank symbol "blank" and using the method of inferring the output probability with the time-step probability. Usually, the ctc prefix beam search method is used for decoding to obtain the output sequence. In the voice wake-up task, there is no need to construct a complex recognition network. Only a hot word graph of the wake word and the approximate pronunciation of the wake word needs to be constructed, and points are added during the beam search process in the form of shallow fusion (ShallowFusion), thereby affecting the path sorting. This method uses the speech recognition method to complete the voice wake-up task, still requiring prefix beam search and constructing a hot word graph, with a large memory consumption and computational complexity.
[0036] In view of the defect that using the above single method to solve the voice wake-up problem cannot balance the wake-up effect and the computational complexity, this embodiment provides a voice wake-up method, including:
[0037] S1: Obtain the original audio signal.
[0038] S2: Preprocess the original audio signal to obtain the preprocessed audio signal.
[0039] S21: Perform frame splitting on the original audio signal to obtain a number of audio frames.
[0040] S22: Calculate the sum of squares of all sampling points in each audio frame.
[0041] S23: Calculate the speech energy of each audio frame according to the sum of squares.
[0042] S24: Smooth each speech energy to obtain the smoothed speech energy.
[0043] S25: Select the audio frames with the smoothed speech energy greater than the energy threshold as the preprocessed audio signal.
[0044] S3: Identify potential targets in the preprocessed audio signal using a first voice wake-up method to obtain a first recognition result. The first voice wake-up method is an improved voice wake-up method based on the CTC loss function. The first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on the CTC loss function to fixed-path decoding. Fixed-path decoding refers to decoding through a target path, and the target path includes the fixed paths of the phonemes in the target wake-up word and the fixed paths of the phonemes with similar pronunciations to the target wake-up word. The potential targets refer to the phonemes of the target wake-up word and the phonemes with similar pronunciations to the target wake-up word.
[0045] S31: Extract features from each audio frame in the preprocessed audio signal to obtain the acoustic features of each audio frame.
[0046] S32: Use each of the acoustic features as input and calculate the posterior probability vector of each target phoneme at each moment using a pre-trained acoustic model. The acoustic model is a deep neural network model based on the CTC loss function, and the target phonemes include the blank item and all non-tonal pinyin items.
[0047] S33: Calculate the fixed-path score of the potential target according to the CTC fixed-path score map. The CTC fixed-path score map includes multiple time steps on the horizontal axis, multiple label groups on the vertical axis, and each time step-label intersection. Each label group includes the blank item and the potential target item, and each time step-label intersection represents the posterior probability vector of the corresponding label at each moment.
[0048] Replace each blank posterior probability vector in the CTC fixed-path score map with a non-potential target posterior probability vector. The blank posterior probability vector refers to the posterior probability vector of each time step corresponding to the blank item, and the non-potential target posterior probability vector refers to the sum of the posterior probability vectors of other target phonemes recognized using the pre-trained acoustic model except for the potential target item in the same label group.
[0049] S34: Determine the qualified fixed path according to the fixed-path score and the path score threshold.
[0050] S35: Determine whether to be preliminarily awakened according to the peak probability value of each phoneme and the peak probability threshold on the qualified fixed path and the distance between adjacent phoneme peaks and the distance threshold. The peak probability value of the phoneme is a variable calculated according to the posterior probability vector.
[0051] S4: When the first recognition result meets the first preset condition, perform preliminary wake-up, and use a second voice wake-up method to identify whether the target audio frame contains the real potential target, obtaining a second recognition result. Herein, the target audio frame refers to the audio frame that may contain the potential target in the first recognition result, the second voice wake-up method is an improved voice wake-up method based on the HMM model, and the second voice wake-up method includes improving the recognition graph in the traditional voice wake-up method based on the HMM model into a target path recognition graph, and the target path recognition graph is a recognition graph constructed according to the target path.
[0052] The determination process of the target audio frame specifically includes:
[0053] According to the peak time point of the first potential target in the preprocessed audio signal, retain the duration of one phoneme pronunciation forward to obtain an extended audio signal;
[0054] Remove the audio signal before the potential target in the extended audio signal, and retain the potential target and the subsequent audio signal to obtain the target audio frame.
[0055] S5: When the second recognition result meets the second preset condition, confirm wake-up, and identify the time of each phoneme of the final target wake-up word in the second recognition result.
[0056] The voice wake-up method provided in this embodiment uses the first voice wake-up method to perform potential target recognition on the preprocessed audio signal to obtain a first recognition result; uses the second voice wake-up method to perform re-potential target recognition on the first recognition result to obtain a second recognition result; and when the second recognition result meets the second preset condition, confirm wake-up and identify the time of each phoneme of the final target wake-up word in the second recognition result. Herein, the first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on the CTC loss function into fixed-path decoding, and the second voice wake-up method includes improving the recognition graph in the traditional voice wake-up method based on the HMM model into a target path recognition graph, which can improve the accuracy of voice wake-up while reducing the computational amount and improving the efficiency of voice wake-up.
[0057] Embodiment 2
[0058] As Figure 2 shown, the voice wake-up method device provided in this embodiment includes:
[0059] An acquisition module, configured to acquire an original audio signal.
[0060] A voice energy detection module, configured to preprocess the original audio signal to obtain a preprocessed audio signal.
[0061] The first voice wake-up module is used to identify potential targets in the preprocessed audio signal by using the first voice wake-up method to obtain a first recognition result. Among them, the first voice wake-up method is an improved voice wake-up method based on the CTC loss function. The first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on the CTC loss function to fixed-path decoding. The fixed-path decoding refers to decoding through a preset target path. The target path includes the fixed path of the phonemes in the target wake-up word and the fixed path of the phonemes with similar pronunciations to the target wake-up word. The potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciations to the target wake-up word.
[0062] The second voice wake-up module is used to identify whether the target audio frame contains the real potential target by using the second voice wake-up method to obtain a second recognition result. Among them, the target audio frame refers to the audio frame that may contain the potential target in the first recognition result. The second voice wake-up method is an improved voice wake-up method based on the HMM model. The second voice wake-up method includes improving the recognition graph in the traditional voice wake-up method based on the HMM model to a target-path recognition graph. The target-path recognition graph is a recognition graph constructed according to the target path.
[0063] The final result obtaining module is used to confirm the wake-up according to the second recognition result and identify the time of each phoneme of the final target wake-up word in the second recognition result.
[0064] To make the specific process of this embodiment clearer to those skilled in the art, the following is a specific explanation.
[0065] To improve the accuracy of the voice wake-up result, this embodiment adopts a dual voice wake-up module structure and adds a voice energy detection module in front of the dual voice wake-up module to reduce the overall computational amount of the voice wake-up module in a quiet scenario.
[0066] As Figure 1 shown, the device includes: the acquisition module acquires the user input audio, inputs the user audio into the voice energy detection module. If the voice energy reaches the preset threshold condition, the user audio is input into the first voice wake-up module. If it meets the preset condition, the user audio is input into the second voice wake-up module. If it meets the preset condition, the results of the first voice wake-up module and the second voice wake-up module are sent to the final result acquisition module to determine the final wake-up result.
[0067] In the voice energy module, the calculation method of voice energy is to frame the user input audio. For example, the frame length is 25 ms and the frame shift is 10 ms. Calculate the sum of the squares of the sampling points in each frame, take the logarithm and record it as the energy. Smooth the energy sequence. When the voice energy meets the preset threshold, send the user input audio to the first voice wake-up module.
[0068] The first voice wake-up module is an acoustic model based on the CTC loss function. This acoustic model is trained using a deep neural network and is used to identify whether there is a wake-up word in the audio. This model does not use the ctc prefix beam search method for decoding, nor does it adjust the path score in the way of a hot word graph. For the wake-up problem in this embodiment, it is only necessary to calculate the fixed path of the wake-up word and the fixed path score of the similar pronunciation of the wake-up word.
[0069] The first voice wake-up module in this embodiment overcomes the defect that "the voice wake-up method based on the CTC loss function solves the alignment problem by introducing the blank symbol blank and using the time step probability to infer the output probability. Usually, the ctcprefix beam search method is used for decoding to obtain the output sequence. In the voice wake-up task, there is no need to build a complex recognition network. It is only necessary to build a hot word graph of the wake-up word and the approximate pronunciation of the wake-up word, and add scores in the beam search process in the form of shallow fusion, so as to affect the path sorting. This method uses the speech recognition method to complete the voice wake-up task, and still needs to perform prefix beam search and build a hot word graph, with a large memory consumption and computational complexity".
[0070] When the fixed path score reaches the preset threshold, check the peak probability value of each phoneme and the distance between adjacent phoneme peaks of this fixed path. When it meets the preset threshold, the first voice wake-up module determines to wake up.
[0071] The first voice wake-up module consists of three parts: a feature extraction module, an encoding module, and a classification module. Such as Figure 3As shown. The feature extraction module extracts features from the user input audio. The feature extraction method can be the original sampling points, MFCC features and their differences, FBANK features and their differences, but is not limited to this. The encoding module can be a Conformer model, a lightweight model processed from the Conformer model, a Zipformer model, a lightweight model processed from the Zipformer model, but is not limited to this. The classification module can be a fully connected layer and an activation function. The number of layers of the fully connected layer is not limited. The activation function can be a Softmax function, a Sigmoid function, but is not limited to this. The classification module outputs a sequence of posterior probability vectors of the acoustic model. Usually, ctc prefixbeam search decoding is required here, but fixed-path decoding is used here, that is, assuming the maximum pronunciation duration of the wake word, such as 2000ms, the number of frames of the posterior probability vector required to calculate the fixed-path score can be determined. For example, if it is 40ms per frame, 50 frames are required. The fixed-path scores of the wake word and its approximate sounds are calculated respectively within the fixed number of frames. The single-path probability calculation formula is as follows:
[0072]
[0073] Among them, C is the path containing blanks and duplicates, X is the sequence of posterior probability vectors, where C = (c1,...c t ,...c T ), X = (x1,...x t ,...x T ), and here y(c t ,t) represents the probability of outputting label c t at time t.
[0074] The fixed-path probability is the sum of all possible single-path probabilities, and the calculation formula is as follows:
[0075] P(S|X) = ∑ c∈A(S) P(C|X) (2);
[0076] Among them, S is the pronunciation sequence of the wake word or approximate sound, and A(S) represents the set of all possible paths of S. The final fixed-path score is the negative logarithm of formula 2, and the calculation formula is as follows:
[0077] Score = -logP(S|X) (3).
[0078] The fixed paths that meet the conditions are screened according to the preset threshold, which is the result of the first voice wake-up module.
[0079] Usually, the way to calculate the fixed path is as Figure 5As shown in (CTC fixed path score graph), to reduce the computational complexity, the forward-backward algorithm is usually used to calculate the sum of all path probabilities. However, for the voice wake-up problem, other phoneme sequences can be inserted before the pronunciation sequences of the wake-up word and its approximate sounds, and it is not necessary to be the blank item. That is, in Figure 5 , for example, the blank item and the potential target item in the first tag group. The posterior probability sequence of the first row does not have to be the posterior probability sequence of the bank item and can be replaced by the sum of the posterior probabilities of other items except the S item in the second row. Thus, the wake-up rate can be improved without increasing the false wake-up rate. That is, other voices can be inserted before the wake-up word and its approximate sounds in the 50-frame posterior probability sequence. For example, if the wake-up word is Xiaoluo Xiaoluo, the voice "Hello, Xiaoluo Xiaoluo" can also wake it up.
[0080] When a certain fixed path score reaches the preset threshold, check the peak probability value of each phoneme in this fixed path and the distance between adjacent phoneme peaks. When it meets the preset threshold, the first voice wake-up module determines the wake-up, and according to the peak time point of the first phoneme of the wake-up word and its approximate sounds, retain the duration of one phoneme pronunciation forward, remove the pronunciation before the wake-up word and its approximate sounds in the user input audio, and send the remaining user input audio segment to the second voice wake-up module.
[0081] When the first voice wake-up module determines that the audio contains the wake-up word, the audio segment containing the wake-up word is input into the second voice wake-up module for secondary confirmation. The second voice wake-up module is constructed based on the HMM method, but there is no need to construct a complex recognition graph for garbage words, etc. It only needs to construct the recognition graph of the fixed path determined by the first voice wake-up module. Search and calculate the confidence in this graph. If it meets the preset threshold, the second voice wake-up module determines the wake-up.
[0082] The second voice wake-up module in this implementation overcomes the defect that "the voice wake-up method based on the HMM model needs to construct a recognition network according to information such as the wake-up word, the approximate pronunciation of the wake-up word, the anti-wake-up word, garbage words or garbage phonemes, and identify the user's speech according to this recognition network. The loading and search results of this recognition network consume a lot of memory resources and computing resources".
[0083] The second voice wake-up module consists of 4 parts: a feature extraction module, a convolutional module, a feed-forward sequence memory network module, and a module for constructing a fixed path recognition graph and searching. As Figure 4The DFSMN structure diagram shown above. The feature extraction module extracts features from the user input audio. The feature extraction method can be MFCC features and their differences, FBANK features and their differences, but is not limited to this. The convolutional module includes a convolutional module which is CNN (Convolutional Neural Networks) and a pooling layer. The number of layers of the CNN and the number of layers of the pooling layer are not limited, and the combination method is not limited. The pooling layer can be max pooling, average pooling, but is not limited to this.
[0084] The Feed-forward Sequence Memory Neural Network module is composed of DFSMN (Deep-FSMN). It adds a skip-connection operation to cFSMN and can stack deeper network structures. The structure is as Figure 4 shown. Its calculation formula is as follows:
[0085]
[0086] It consists of three parts. The first part is the output of the previous layer's sequence memory module, the second part is the output of the linear mapping layer, and the third part is the FSMN structure of this layer. FSMN refers to "Fully Connected State Memory Network", which models the historical and future temporal information and integrates the encoded information of a fixed dimension. After the three parts are added, they are input to the next hidden layer through a linear mapping. Here, l represents the number of layers of FSMN, represents the l-th layer linear mapping layer, represents the output of the l-th memory block, represents the length of the future time slice and represents the length of the historical time slice. H represents the skip-connection between the memory blocks and can be any linear or non-linear transformation.
[0087] The second voice wake-up module makes a second confirmation based on the result of the first voice wake-up module. Assuming a known pronunciation sequence, the audio segment of the wake-up word or its approximate sound determined by the first wake-up model is sent to the second voice wake-up module for feature extraction and the posterior probability is obtained through a neural network model. Only path search is performed in the recognition graph constructed by the fixed path of the known wake-up word or its approximate sound, and the path score is output. If the score meets the preset threshold, the wake-up is confirmed, and the path result is sent to the final result acquisition module, so that the time range of each pronunciation unit can be confirmed. While returning the confirmed wake-up, the time information of each pronunciation unit is also returned.
[0088] 1. Data Preparation
[0089] Select the speech data for training the neural network part of the first voice wake-up module and the neural network part of the second voice wake-up module. Require near-field recording data (within 1m) with clear and accurate pronunciation and a quiet room. Try to achieve an even distribution in terms of age, gender, and region. Data simulation can be appropriately performed, including adding noise, reverberation, amplitude adjustment, and speech rate adjustment, etc.
[0090] 2. Model Training
[0091] Select the model structure and number of network layers of the acoustic models of the first voice wake-up module and the second voice wake-up module according to the usage scenario and device computing power. Use the prepared audio data and its annotations to train the acoustic models of the first voice wake-up module and the second voice wake-up module. End when the loss functions of the two models converge.
[0092] 3. Engineering of the Wake-up Engine
[0093] Implement each sub-module in the first voice wake-up module and the second voice wake-up module according to the above solution. When the user continuously inputs audio, the wake-up result, as well as the audio of the wake-up word and its approximate pronunciation, can be continuously and real-time returned.
[0094] In the voice energy module, the frame segmentation parameters can be a frame length of 25ms and a frame shift of 10ms. Calculate the sum of the squares of the sampling points in each frame and take the logarithm, which is recorded as the energy. Smooth this energy sequence. The smoothing window length can be 4, and the energy threshold can be 11.
[0095] In the first voice wake-up module, the feature module can be selected as 80-dimensional FBANK features and their first-order and second-order differences. The encoding module can select a lightweight Conformer network structure. For the fixed-path score, if it is a four-character wake-up word, the threshold can be selected as 5, the probability of each spike is not less than 0.5, and the distance between adjacent phoneme spikes is not greater than 400ms.
[0096] In the second voice wake-up module, the feature module can be selected as 80-dimensional FBANK features and their first-order and second-order differences. Three CNN layers can be selected, and each CNN layer is connected to a max-pooling layer. After the output features are dimensionally transformed, they can be connected to five DFSMN networks. Perform path search in the recognition graph constructed for the fixed path of the known wake-up word or its approximate pronunciation. The output path score threshold is 2. When the path score meets this threshold, the wake-up is confirmed, and the specific time of each phoneme of the wake-up word is returned.
[0097] The voice wake-up method based on the HMM model needs to construct an identification network according to information such as wake-up words, approximate pronunciations of wake-up words, anti-wake-up words, garbage words, or garbage phonemes, and identify the user's speech according to this identification network. The loading and search results of this identification network consume a lot of memory resources and computing resources.
[0098] The voice wake-up method based on the CE loss function transforms the voice wake-up problem from a speech recognition task into an image classification task. It only needs to smooth the posterior probability results of the acoustic model, without the need to construct an identification network and search. However, its accuracy depends on the pre-annotated wake-up word boundary information and does not support the custom wake-up word function.
[0099] The voice wake-up method based on the CTC loss function solves the alignment problem by introducing the blank symbol "blank" and using the time-step probability to infer the output probability. Usually, the ctc prefix beam search method is used to decode to obtain the output sequence. In the voice wake-up task, there is no need to construct a complex identification network. Only a hot word graph of wake-up words and approximate pronunciations of wake-up words needs to be constructed, and points are added during the beam search process in the form of shallow fusion (ShallowFusion), thereby affecting the path sorting. This method uses the speech recognition method to complete the voice wake-up task, and still needs to perform prefix beam search and construct a hot word graph, with a large memory consumption and computational complexity.
[0100] In this embodiment, the modeling units in the first voice wake-up module and the second voice wake-up module are non-tonal pinyin (i.e., phonemes), so the neural network can support custom wake-up words. Just add the corresponding path in the post-processing to support new wake-up words at any time, overcoming the defect that "the voice wake-up method based on the CE loss function transforms the voice wake-up problem from a speech recognition task into an image classification task. It only needs to smooth the posterior probability results of the acoustic model, without the need to construct an identification network and search. However, its accuracy depends on the pre-annotated wake-up word boundary information and does not support the custom wake-up word function".
[0101] The voice wake-up device provided in this embodiment solves the following technical problems:
[0102] 1) When using the speech recognition method to solve the voice wake-up problem, there is a problem of constructing a weighted finite state transducer (WFST) recognition graph. The memory consumption and computational complexity of constructing this graph and searching this graph are both very large.
[0103] 2) When using a single method to solve the voice wake-up problem, it is impossible to balance the wake-up effect and the computational complexity. In this embodiment, a dual voice wake-up module structure is adopted. When the first voice wake-up module determines that the user input audio includes a wake-up word, the second voice wake-up module is used to determine whether to wake up again.
[0104] This embodiment realizes that under the conditions of ensuring the accuracy of voice wake-up and custom wake-up words, it is possible to avoid constructing a complex recognition graph and its search calculation, avoid performing ctc prefix beam search on all acoustic modeling units, and avoid constructing a hot word graph; in view of the particularity of the voice wake-up problem, it focuses on the fixed path of the wake-up word and its similar sounds; and uses the second voice wake-up module for secondary confirmation to further ensure the accuracy of voice wake-up.
[0105] Embodiment 3
[0106] This embodiment provides a computer device, which can be a server or a terminal, and its internal structure diagram can be as shown in Figure 6 The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data in the voice wake-up method. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes the voice wake-up method in Embodiment 1.
[0107] Those skilled in the art can understand that Figure 6 The structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor, and a computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are realized.
[0108] Embodiment 4
[0109] This embodiment provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the voice wake-up method in Embodiment 1.
[0110] Embodiment 5
[0111] This embodiment provides a computer program product including a computer program, which, when executed by a processor, implements the voice wake-up method in Embodiment 1.
[0112] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0113] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0114] In each of the embodiments provided in this application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., and is not limited thereto. In each of the embodiments provided in this application, the processor may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., and is not limited thereto.
[0115] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0116] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A voice wake-up method, characterized in that, The described voice wake-up method includes: Obtain the original audio signal; Perform preprocessing on the original audio signal to obtain the preprocessed audio signal; Use the first voice wake-up method to identify potential targets in the preprocessed audio signal, obtaining a first recognition result. Among them, the first voice wake-up method is an improved voice wake-up method based on the CTC loss function. The first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on the CTC loss function to fixed-path decoding. The fixed-path decoding refers to decoding through a target path, and the target path includes the fixed path of the phonemes in the target wake-up word and the fixed path of the phonemes with similar pronunciations to the target wake-up word. The potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciations to the target wake-up word; When the first recognition result meets the first preset condition, perform preliminary wake-up, and use the second voice wake-up method to identify whether the target audio frame contains the real potential target, obtaining a second recognition result. Among them, the target audio frame refers to the audio frame that may contain the potential target in the first recognition result. The second voice wake-up method is an improved voice wake-up method based on the HMM model. The second voice wake-up method includes improving the recognition graph in the traditional voice wake-up method based on the HMM model to a target-path recognition graph, and the target-path recognition graph is a recognition graph constructed according to the target path; When the second recognition result meets the second preset condition, confirm wake-up, and identify the time of each phoneme of the final target wake-up word in the second recognition result.
2. The voice wake-up method according to claim 1, wherein Performing preprocessing on the original audio signal to obtain the preprocessed audio signal specifically includes: Perform frame splitting on the original audio signal to obtain a number of audio frames; Calculate the sum of squares of all sampling points in each audio frame; Calculate the speech energy of each audio frame according to the sum of squares; Perform smoothing processing on each speech energy to obtain the smoothed speech energy; Select the audio frames with the smoothed speech energy greater than the energy threshold as the preprocessed audio signal.
3. The voice wake-up method according to claim 1, wherein Using the first voice wake-up method to identify potential targets in the preprocessed audio signal, obtaining a first recognition result specifically includes: Extract features from each audio frame in the preprocessed audio signal to obtain the acoustic features of each audio frame; Take each acoustic feature as input, and use a pre-trained acoustic model to calculate the posterior probability vector of each target phoneme at each moment. Among them, the acoustic model is a deep neural network model based on the CTC loss function, and the target phonemes include the blank item and all non-tonal pinyin items; Calculate the fixed-path score of the potential target according to the CTC fixed-path score map. Among them, the CTC fixed-path score map includes multiple time steps on the horizontal axis, multiple label groups on the vertical axis, and each time step-label intersection. Each label group includes the blank item and the potential target item, and each time step-label intersection represents the posterior probability vector of the corresponding label at each moment; Determine the eligible fixed paths according to the fixed path scores and the path score thresholds; Judge whether to perform a preliminary wake-up according to the phoneme peak probability values and peak probability thresholds on each eligible fixed path and the distances and distance thresholds between adjacent phoneme peaks, where the phoneme peak probability value is a variable calculated according to the posterior probability vector.
4. The voice wake-up method according to claim 3, wherein Replace each blank posterior probability vector in the CTC fixed path score map with a non-potential target posterior probability vector, where the blank posterior probability vector refers to the posterior probability vector at each time step corresponding to the blank item, and the non-potential target posterior probability vector refers to the sum of the posterior probability vectors of other target phonemes recognized by using a pre-trained acoustic model except for the potential target items in the same label group.
5. The voice wake-up method according to claim 3, wherein The calculation expression of the fixed path score of the potential target is: P(S|X) = ∑ c∈A(S) P(C|X); Score = -logP(S|X); Among them, P(C|X) represents the probability of a single path, C represents a path containing blank items and duplicate items, C = (c1,...c t ,...c T ); X represents a sequence of posterior probability vectors, which is calculated based on posterior probability vectors, X = (x1,...x t ,...x T ); y(c t ,t) represents the probability that the output label is c t at time t; P(S|X) is the fixed path probability; S represents the pronunciation sequence of the potential target; A(S) represents the set of all possible paths of S; Score represents the fixed path score.
6. The voice wake-up method according to claim 1, characterized in that The process of determining the target audio frame specifically includes: According to the peak time point of the first potential target in the preprocessed audio signal, retain the duration of one phoneme pronunciation forward to obtain an extended audio signal; Remove the audio signal before the potential target in the extended audio signal, and retain the potential target and its subsequent audio signal to obtain the target audio frame.
7. A voice wake-up method and device, characterized in that The speech wake-up method device includes: An acquisition module, configured to acquire an original audio signal; A speech energy detection module, configured to preprocess the original audio signal to obtain a preprocessed audio signal; A first speech wake-up module, configured to use a first speech wake-up method to identify potential targets in the preprocessed audio signal to obtain a first recognition result, where the first speech wake-up method is an improved speech wake-up method based on the CTC loss function, and the first speech wake-up method includes improving the decoding method in the traditional speech wake-up method based on the CTC loss function to fixed path decoding, and the fixed path decoding refers to decoding through a preset target path, and the target path includes the fixed paths of the phonemes in the target wake-up word and the fixed paths of the phonemes with similar pronunciations to the target wake-up word, and the potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciations to the target wake-up word; A second speech wake-up module, configured to use a second speech wake-up method to identify whether the target audio frame contains the real potential target to obtain a second recognition result, where the target audio frame refers to the audio frame that may contain the potential target in the first recognition result, and the second speech wake-up method is an improved speech wake-up method based on the HMM model, and the second speech wake-up method includes improving the recognition graph in the traditional speech wake-up method based on the HMM model to a target path recognition graph, and the target path recognition graph is a recognition graph constructed according to the target path; A final result obtaining module, configured to confirm wake-up according to the second recognition result and identify the time of each phoneme of the final target wake-up word in the second recognition result.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that the processor executes the computer program to implement the voice wake-up method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the voice wake-up method according to any one of claims 1-6.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the voice wake-up method according to any one of claims 1-6.
Citation Information
Patent Citations
Voice wake-up method, voice wake-up device and storage medium
CN115762480A
Audio recognition method, audio recognition device, vehicle, computer equipment and medium
CN117456999A