Voice wake-up method, device, equipment, medium and product
Through the improved CTC loss function and the voice wake-up method of the HMM model, the dual recognition technology of fixed path decoding and target path recognition diagram is adopted, and the problem of difficult to balance the voice wake-up effect and calculation amount in the prior art is solved, achieving more efficient and accurate voice wake-up.
Patent Information
- Application Number
- CN202510147661.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-11
AI Technical Summary
Existing voice wake-up methods are difficult to balance the wake-up effect and the amount of calculation, resulting in poor user experience.
Fixed path decoding is performed using an improved speech wake-up method based on CTC loss function, and a target path recognition diagram is used in combination with an improved speech wake-up method based on HMM model, and double recognition is performed to confirm wake-up.
It improves the accuracy of voice wake-up, while reducing the amount of calculation, and improving the efficiency of voice wake-up.
Smart Images

Figure CN119993164A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice wake-up technology, and in particular to a voice wake-up method, device, equipment, medium and product. Background Art
[0002] Voice has always been a way of interaction between people. In the human-computer interaction of smart hardware such as smart speakers and smart set-top boxes, voice interaction is the most natural way of interaction, which is very convenient and easy to understand. In recent years, the development of artificial intelligence technology has made great progress. At present, machines can recognize and understand the internal meaning of voice and respond accordingly in time, such as playing voice and calling related skills. In this interaction process, the real-time and accuracy of voice wake-up greatly affects the user experience and is an important prerequisite for smooth voice interaction.
[0003] The current voice wake-up methods mostly use the following three methods: based on the Hidden Markov Model (HMM) model, based on the Cross-Entropy (CE) loss function, and based on the Connectionist Temporal Classification (CTC) loss function. However, in actual use, a single method is usually used to solve the voice wake-up problem, which cannot balance the wake-up effect and the amount of calculation. Summary of the invention
[0004] The purpose of this application is to provide a voice wake-up method, device, equipment, medium and product, which can improve the accuracy of voice wake-up while reducing the amount of calculation and improving the efficiency of voice wake-up.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] In a first aspect, the present application provides a voice wake-up method, comprising:
[0007] Get the original audio signal;
[0008] Preprocessing the original audio signal to obtain a preprocessed audio signal;
[0009] A first voice wake-up method is used to identify potential targets in the preprocessed audio signal to obtain a first recognition result, wherein the first voice wake-up method is an improved voice wake-up method based on a CTC loss function, and the first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on a CTC loss function to fixed path decoding, wherein the fixed path decoding refers to decoding through a target path, wherein the target path includes a fixed path of phonemes in a target wake-up word and a fixed path of phonemes with similar pronunciation to the target wake-up word, and the potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciation to the target wake-up word;
[0010] When the first recognition result meets the preset conditions, a preliminary wake-up is performed, and a second voice wake-up method is used to identify whether the target audio frame contains the real potential target, so as to obtain a second recognition result, wherein the target audio frame refers to an audio frame that may contain the potential target in the first recognition result, and the second voice wake-up method is an improved voice wake-up method based on the HMM model, and the second voice wake-up method includes improving the recognition graph in the traditional voice wake-up method based on the HMM model into a target path recognition graph, and the target path recognition graph is a recognition graph constructed according to the target path;
[0011] The wake-up is confirmed according to the second recognition result, and the time of each phoneme of the final target wake-up word in the second recognition result is identified.
[0012] In a second aspect, the present application provides a voice wake-up method and device, comprising:
[0013] An acquisition module, used for acquiring an original audio signal;
[0014] A speech energy detection module, used for preprocessing the original audio signal to obtain a preprocessed audio signal;
[0015] A first voice wake-up module, used to identify potential targets in the preprocessed audio signal by using a first voice wake-up method to obtain a first recognition result, wherein the first voice wake-up method is an improved voice wake-up method based on a CTC loss function, and the first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on a CTC loss function to fixed path decoding, wherein the fixed path decoding refers to decoding through a preset target path, wherein the target path includes a fixed path of phonemes in a target wake-up word and a fixed path of phonemes with similar pronunciation to the target wake-up word, and the potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciation to the target wake-up word;
[0016] A second voice wake-up module, used for identifying whether the target audio frame contains the real potential target by using a second voice wake-up method, and obtaining a second recognition result, wherein the target audio frame refers to an audio frame that may contain the potential target in the first recognition result, and the second voice wake-up method is an improved voice wake-up method based on the HMM model, and the second voice wake-up method includes improving the recognition graph in the traditional voice wake-up method based on the HMM model into a target path recognition graph, and the target path recognition graph is a recognition graph constructed according to the target path;
[0017] The final result acquisition module is used to confirm the wake-up according to the second recognition result and identify the time of each phoneme of the final target wake-up word in the second recognition result.
[0018] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the voice wake-up method described in the first aspect above.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the voice wake-up method described in the first aspect above.
[0020] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the voice wake-up method described in the first aspect above.
[0021] According to the specific embodiments provided in this application, this application has the following technical effects:
[0022] The present application provides a voice wake-up method, device, equipment, medium and product, the method comprising: preprocessing an original audio signal; using a first voice wake-up method to identify a potential target on the preprocessed audio signal to obtain a first recognition result; using a second voice wake-up method to identify a potential target again on the first recognition result to obtain a second recognition result; and confirming the wake-up when the second recognition result meets a second preset condition, and identifying the time of each phoneme of the final target wake-up word in the second recognition result, wherein the first voice wake-up method comprises improving the decoding method in the traditional voice wake-up method based on the CTC loss function to fixed path decoding, and the second voice wake-up method comprises improving the recognition graph in the traditional voice wake-up method based on the HMM model to a target path recognition graph, which can improve the accuracy of voice wake-up while reducing the amount of calculation and improving the efficiency of voice wake-up. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0024] Figure 1 A schematic diagram of the functional modules of a voice wake-up device provided in Example 2 of the present application;
[0025] Figure 2 This is a schematic diagram of the structure of the first voice wake-up module in Example 2 of the present application;
[0026] Figure 3 This is a schematic diagram of the structure of the second voice wake-up module in Example 2 of the present application;
[0027] Figure 4 This is a structural diagram of the DFSMN in Example 2 of the present application;
[0028] Figure 5 It is the CTC fixed path score graph in Example 2 of the present application;
[0029] Figure 6 A schematic diagram of the structure of a computer device provided in Example 3 of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0031] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0032] Example 1
[0033] Research has found that the voice wake-up method based on the HMM model needs to build a recognition network based on information such as wake-up words, approximate pronunciations of wake-up words, anti-wake-up words, junk words or junk phonemes, and recognize the user's voice based on the recognition network. The loading and search results of the recognition network consume a lot of memory and computing resources.
[0034] The voice wake-up method based on the CE loss function transforms the voice wake-up problem from a speech recognition task to an image classification task. It only needs to smooth the posterior probability results of the acoustic model without building a recognition network and searching. However, its accuracy depends on the boundary information of the wake-up word annotated in advance, and it cannot support the custom wake-up word function.
[0035] The voice wake-up method based on the CTC loss function solves the alignment problem by introducing blank characters and using the time step probability to infer the output probability. The ctc prefix beam search method is usually used to decode and obtain the output sequence. In the voice wake-up task, there is no need to build a complex recognition network. It is only necessary to build a hot word graph of the wake-up word and the approximate pronunciation of the wake-up word, and add points in the beam search process in the form of shallow fusion, thereby affecting the path sorting. This method uses the speech recognition method to complete the voice wake-up task, and still needs to perform prefix beam search and build a hot word graph, which consumes a lot of memory and calculation.
[0036] In view of the defect that the above single method is used to solve the voice wake-up problem and cannot balance the wake-up effect and the amount of calculation, this embodiment provides a voice wake-up method, including:
[0037] S1: Get the original audio signal.
[0038] S2: Preprocess the original audio signal to obtain a preprocessed audio signal.
[0039] S21: performing frame processing on the original audio signal to obtain a plurality of audio frames.
[0040] S22: Calculate the sum of squares of all sampling points in each of the audio frames.
[0041] S23: Calculate the speech energy of each of the audio frames according to the square sum.
[0042] S24: performing smoothing processing on each of the speech energies to obtain smoothed speech energy.
[0043] S25: Select the audio frame whose smoothed speech energy is greater than the energy threshold as the preprocessed audio signal.
[0044] S3: Use a first voice wake-up method to identify potential targets in the preprocessed audio signal to obtain a first recognition result, wherein the first voice wake-up method is an improved voice wake-up method based on a CTC loss function, and the first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on a CTC loss function to fixed path decoding, and the fixed path decoding refers to decoding through a target path, and the target path includes a fixed path of phonemes in a target wake-up word and a fixed path of phonemes with similar pronunciation to the target wake-up word, and the potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciation to the target wake-up word.
[0045] S31: extracting features from each audio frame in the preprocessed audio signal to obtain acoustic features of each audio frame.
[0046] S32: Taking each of the acoustic features as input, using a pre-trained acoustic model to calculate the posterior probability vector of each target phoneme at each moment, wherein the acoustic model is a deep neural network model based on a CTC loss function, and the target phonemes include blank items and all toneless pinyin items.
[0047] S33: Calculate the fixed path score of the potential target according to the CTC fixed path score graph, wherein the CTC fixed path score graph includes multiple time steps on the horizontal axis, multiple label groups on the vertical axis and each time step-label intersection, each of the label groups includes a blank item and a potential target item, and each of the time step-label intersection represents the posterior probability vector of the corresponding label at each moment.
[0048] Each blank posterior probability vector in the CTC fixed path score graph is replaced with a non-potential target posterior probability vector, wherein the blank posterior probability vector refers to the posterior probability vector of each time step corresponding to the blank item, and the non-potential target posterior probability vector refers to the sum of the posterior probability vectors of other target phonemes recognized by the pre-trained acoustic model except the potential target items in the same label group.
[0049] S34: Determine a fixed path that meets the condition according to the fixed path score and the path score threshold.
[0050] S35: Determine whether to initially wake up based on each phoneme peak probability value and peak probability threshold on the qualified fixed path and the distance and distance threshold of adjacent phoneme peaks, wherein the phoneme peak probability value is a variable calculated based on the posterior probability vector.
[0051] S4: When the first recognition result satisfies the first preset condition, a preliminary wake-up is performed, and a second voice wake-up method is used to identify whether the target audio frame contains the actual potential target to obtain a second recognition result, wherein the target audio frame refers to the audio frame that may contain the potential target in the first recognition result, and the second voice wake-up method is an improved voice wake-up method based on the HMM model. The second voice wake-up method includes improving the recognition graph in the traditional HMM model-based voice wake-up method into a target path recognition graph, and the target path recognition graph is a recognition graph constructed according to the target path.
[0052] The process of determining the target audio frame specifically includes:
[0053] According to the peak time point of the first potential target in the preprocessed audio signal, retain the duration of a phoneme pronunciation forward to obtain an extended audio signal;
[0054] The audio signal before the potential target in the extended audio signal is removed, and the audio signal of the potential target and its subsequent audio signal are retained to obtain the target audio frame.
[0055] S5: When the second recognition result satisfies the second preset condition, confirm the wake-up, and identify the time of each phoneme of the final target wake-up word in the second recognition result.
[0056] The voice wake-up method provided in this embodiment adopts a first voice wake-up method to identify potential targets of a preprocessed audio signal to obtain a first recognition result; adopts a second voice wake-up method to identify potential targets again on the first recognition result to obtain a second recognition result; and confirms the wake-up when the second recognition result meets the second preset condition, and identifies the time of each phoneme of the final target wake-up word in the second recognition result, wherein the first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on the CTC loss function to fixed path decoding, and the second voice wake-up method includes improving the recognition graph in the traditional voice wake-up method based on the HMM model to a target path recognition graph, which can improve the accuracy of voice wake-up while reducing the amount of calculation and improving the efficiency of voice wake-up.
[0057] Example 2
[0058] like Figure 2 As shown, this embodiment provides a voice wake-up method device including:
[0059] The acquisition module is used to acquire the original audio signal.
[0060] The speech energy detection module is used to preprocess the original audio signal to obtain a preprocessed audio signal.
[0061] The first voice wake-up module is used to use a first voice wake-up method to identify potential targets in the preprocessed audio signal to obtain a first recognition result, wherein the first voice wake-up method is an improved voice wake-up method based on a CTC loss function, and the first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on a CTC loss function to fixed path decoding, and the fixed path decoding refers to decoding through a pre-set target path, and the target path includes a fixed path of phonemes in a target wake-up word and a fixed path of phonemes with similar pronunciation to the target wake-up word, and the potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciation to the target wake-up word.
[0062] The second voice wake-up module is used to use a second voice wake-up method to identify whether the target audio frame contains the real potential target, and obtain a second recognition result, wherein the target audio frame refers to the audio frame that may contain the potential target in the first recognition result, and the second voice wake-up method is an improved HMM model-based voice wake-up method. The second voice wake-up method includes improving the recognition graph in the traditional HMM model-based voice wake-up method into a target path recognition graph, and the target path recognition graph is a recognition graph constructed according to the target path.
[0063] The final result acquisition module is used to confirm the wake-up according to the second recognition result and identify the time of each phoneme of the final target wake-up word in the second recognition result.
[0064] In order to make the specific process of this embodiment more clear to those skilled in the art, a specific explanation is given below.
[0065] In order to improve the accuracy of the voice wake-up results, this embodiment adopts a dual voice wake-up module structure and adds a voice energy detection module before the dual voice wake-up module to reduce the overall calculation amount of the voice wake-up module in a quiet scene.
[0066] like Figure 1 As shown, the device includes: an acquisition module acquires user input audio, inputs the user audio into a speech energy detection module, if the speech energy reaches a preset threshold condition, inputs the user audio into a first voice wake-up module, if it meets the preset conditions, inputs the user audio into a second voice wake-up module, if it meets the preset conditions, sends the first voice wake-up module result and the second voice wake-up module result to a final result acquisition module to determine the final wake-up result.
[0067] In the speech energy module, the method for calculating speech energy is to divide the user input audio into frames, such as a frame length of 25ms and a frame shift of 10ms, calculate the square sum of the sampling points in each frame and take the logarithm as energy, smooth the energy sequence, and when the speech energy meets the preset threshold, send the user input audio to the first speech wake-up module.
[0068] The first voice wake-up module is an acoustic model based on the CTC loss function. The acoustic model is trained using a deep neural network and is used to identify whether there is a wake-up word in the audio. The model does not use the ctc prefix beam search method for decoding, nor does it use the hot word graph to adjust the path score. For the wake-up problem, this embodiment only needs to calculate the fixed path of the wake-up word and the fixed path score of the wake-up word with similar pronunciations.
[0069] The first speech wake-up module of this embodiment overcomes the defect that "the speech wake-up method based on the CTC loss function solves the alignment problem by introducing the blank character blank and using the time step probability to infer the output probability. The ctcprefix beam search method is usually used to decode and obtain the output sequence. In the speech wake-up task, there is no need to build a complex recognition network. It is only necessary to build a hot word graph of the wake-up word and the approximate pronunciation of the wake-up word, and add points in the beam search process in the form of shallow fusion, thereby affecting the path sorting. This method uses a speech recognition method to complete the speech wake-up task, and still needs to perform prefix beam search and build a hot word graph, which consumes a lot of memory and calculation."
[0070] When a fixed path score reaches a preset threshold, each phoneme peak probability value of the fixed path and the distance between adjacent phoneme peaks are checked, and when the preset threshold is met, the first voice wake-up module determines wake-up.
[0071] The first voice wake-up module consists of three parts: feature extraction module, encoding module, and classification module. Figure 3As shown. The feature extraction module extracts features from the user input audio. The feature extraction method can be the original sampling point, MFCC features and their differences, FBANK features and their differences, but is not limited to this. The encoding module can be a Conformer model, a lightweight model processed by the Conformer model, a Zipformer model, and a lightweight model processed by the Zipformer model, but is not limited to this. The classification module can be a fully connected layer and an activation function. The number of layers of the fully connected layer is not limited. The activation function can be a Softmax function or a Sigmoid function, but is not limited to this. The classification module outputs a sequence of posterior probability vectors of the acoustic model. Usually, CTC prefix beam search decoding is required here, but fixed path decoding is used here, that is, assuming that the maximum pronunciation duration of the wake-up word is 2000ms, the number of frames of the posterior probability vector required to calculate the fixed path score can be determined, such as 40ms per frame, which requires 50 frames. The scores of the fixed paths of the wake-up word and its approximate sound are calculated separately within the fixed number of frames. The probability calculation formula for a single path is as follows:
[0072]
[0073] Where C is the path containing blank and repeated items, X is the sequence of posterior probability vectors, where C = (c1, ...c t , ...c T ), X=(x1,...x t ,...x T ), where y(c t ,t) means the output label is c at time t t probability.
[0074] The fixed path probability is the sum of all possible single path probabilities, and the calculation formula is as follows:
[0075] P(S|X)=∑ c∈A(S) P(C|X) (2);
[0076] Where S is the pronunciation sequence of the wake-up word or similar sound, and A(S) represents the set of all possible paths of S. The final score of the fixed path is the negative logarithm of formula 2, and the calculation formula is as follows:
[0077] Score = -logP(S|X) (3).
[0078] The fixed paths that meet the conditions are screened out according to the preset threshold, which are the results of the first voice wake-up module.
[0079] Typically, fixed paths are calculated as Figure 5As shown in the figure (CTC fixed path score graph), in order to reduce the amount of calculation, a forward-backward algorithm is usually used to calculate the sum of all path probabilities. However, for the voice wake-up problem, other phoneme sequences can be inserted before the pronunciation sequence of the wake-up word and its similar sounds, without the need for blank items. Figure 5 For example, for the blank item and the potential target item in the first label group, the first row does not need to be the posterior probability sequence of the bank item, but can be replaced by the posterior probability sum of other items except the S item in the second row. This can improve the wake-up rate without increasing the false wake-up rate, that is, other voices can be inserted before the wake-up word and its approximate sound in the 50-frame posterior probability sequence. For example, if the wake-up word is Xiao Luo Xiao Luo, the voice "Hao Xiao Luo Xiao Luo" can also be used for wake-up.
[0080] When the score of a fixed path reaches a preset threshold, the peak probability value of each phoneme in the fixed path and the distance between adjacent phoneme peaks are checked. When the preset threshold is met, the first voice wake-up module determines the wake-up and reserves the duration of a phoneme pronunciation based on the peak time point of the first phoneme of the wake-up word and its approximate sound. The pronunciation before the wake-up word and its approximate sound is removed from the user input audio, and the remaining user input audio segment is sent to the second voice wake-up module.
[0081] When the first voice wake-up module determines that the audio contains the wake-up word, the audio segment containing the wake-up word is input into the second voice wake-up module for secondary confirmation. The second voice wake-up module is constructed based on the HMM method, but it does not need to construct complex recognition graphs such as junk words. It only needs to construct a recognition graph of the fixed path determined by the first voice wake-up module, search and calculate the confidence in the graph, and if it meets the preset threshold, the second voice wake-up module determines the wake-up.
[0082] The second voice wake-up module of this implementation overcomes the defect that "the voice wake-up method based on the HMM model needs to build a recognition network according to the information such as the wake-up word, the approximate pronunciation of the wake-up word, the anti-wake-up word, the junk word or the junk phoneme, and recognize the user's voice according to the recognition network. The loading and search results of the recognition network consume a lot of memory resources and computing resources."
[0083] The second voice wake-up module consists of four parts: feature extraction module, convolution module, feedforward sequence memory network module, fixed path recognition graph construction and search module. Figure 4In the DFSMN structure diagram shown, the feature extraction module extracts features from the user input audio, and the feature extraction method can be MFCC features and their differences, FBANK features and their differences, but is not limited thereto. The convolution module includes a convolution module that is a CNN (Convolutional Neural Networks) and a pooling layer, wherein the number of CNN layers and the number of pooling layers are not limited, and the combination method is not limited. The pooling layer can be maximum pooling or average pooling, but is not limited thereto.
[0084] The Feed-forward Sequence Memory Neural Network module is composed of DFSMN (Deep-FSMN). It adds skip-connection operations to cFSMN and can stack a deeper network structure. The structure is as follows: Figure 4 The calculation formula is as follows:
[0085]
[0086] It consists of three parts. The first part is the output of the sequence memory module of the previous layer, the second part is the output of the linear mapping layer, and the third part is the FSMN structure of this layer. FSMN refers to the "Fully Connected State Memory Network", which models historical and future time series information and integrates fixed-dimensional encoding. After the three parts are added, they are input to the next hidden layer through linear mapping. l represents the number of FSMN layers. represents the lth linear mapping layer, represents the output of the memory block at layer l, represents the length of future time slices and represents the length of the time slice of the history, and H represents the skip-connection between memory blocks, which can be any linear or nonlinear transformation.
[0087] The second voice wake-up module is a second confirmation based on the result of the first voice wake-up module. Assuming the pronunciation sequence is known, the audio segment of the wake-up word or its approximate sound determined by the first wake-up model is sent to the second voice wake-up module for feature extraction and the posterior probability is obtained through the neural network model. Only the path search is performed in the recognition graph constructed by the fixed path of the known wake-up word or its approximate sound, and the path score is output. If the score meets the preset threshold, the wake-up is confirmed, and the path result is sent to the final result acquisition module to confirm the time range of each pronunciation unit. When the wake-up confirmation is returned, the time information of each pronunciation unit is also returned.
[0088] 1. Data preparation
[0089] The speech data used to train the neural network part of the first voice wake-up module and the neural network part of the second voice wake-up module must be clear and accurate, and the near-field recording data (within 1m) in a quiet room should be evenly distributed in terms of age, gender, and region. Data simulation can be appropriately performed, including adding noise, reverberation, amplitude adjustment, speech speed adjustment, etc.
[0090] 2. Model training
[0091] The model structure and network layers of the acoustic model of the first voice wake-up module and the second voice wake-up module are selected according to the usage scenario and the computing power of the device, and the acoustic model training of the first voice wake-up module and the second voice wake-up module is performed using the prepared audio data and its annotations. The training ends when the loss functions of the two models converge.
[0092] 3. Engineering of wake-up engine
[0093] According to the above scheme, each submodule in the first voice wake-up module and the second voice wake-up module is implemented. When the user continues to input audio, the wake-up result and the audio of the wake-up word and its similar sound can be continuously returned in real time.
[0094] In the speech energy module, the frame parameters can be a frame length of 25ms and a frame shift of 10ms. The square sum of the sampling points in each frame is calculated and the logarithm is recorded as energy. The energy sequence is smoothed. The smoothing window length can be 4 and the energy threshold can be 11.
[0095] In the first voice wake-up module, the feature module can be selected as 80-dimensional FBANK features and their first-order and second-order differences. The encoding module can select a lightweight Conformer network structure. For fixed path scores, if it is a four-character wake-up word, the threshold can be selected as 5, the probability of each peak is not less than 0.5, and the distance between adjacent phoneme peaks is not more than 400ms.
[0096] In the second voice wake-up module, the feature module can be selected as 80-dimensional FBANK features and their first-order and second-order differences. Three CNN layers can be selected, and each CNN layer is connected to a maximum pooling layer. After the output features are transformed in dimension, they can be connected to a five-layer DFSMN network. Path search is performed in the recognition graph constructed by the fixed path of the known wake-up word or its approximate sound. The output path score threshold is 2. When the path score meets the threshold, the wake-up is confirmed and the specific time of each phoneme of the wake-up word is returned.
[0097] The voice wake-up method based on the HMM model needs to build a recognition network based on information such as the wake-up word, the approximate pronunciation of the wake-up word, the anti-wake-up word, junk words or junk phonemes, and recognize the user's voice based on the recognition network. The loading and search results of the recognition network consume a lot of memory resources and computing resources.
[0098] The voice wake-up method based on the CE loss function transforms the voice wake-up problem from a speech recognition task to an image classification task. It only needs to smooth the posterior probability results of the acoustic model without building a recognition network and searching. However, its accuracy depends on the boundary information of the wake-up word annotated in advance, and it cannot support the custom wake-up word function.
[0099] The voice wake-up method based on the CTC loss function solves the alignment problem by introducing blank characters and using the time step probability to infer the output probability. The ctc prefix beam search method is usually used to decode and obtain the output sequence. In the voice wake-up task, there is no need to build a complex recognition network. It is only necessary to build a hot word graph of the wake-up word and the approximate pronunciation of the wake-up word, and add points in the beam search process in the form of shallow fusion, thereby affecting the path sorting. This method uses the speech recognition method to complete the voice wake-up task, and still needs to perform prefix beam search and build a hot word graph, which consumes a lot of memory and calculation.
[0100] In this embodiment, the modeling units in the first voice wake-up module and the second voice wake-up module are toneless pinyin (i.e., phonemes), so the neural network can support custom wake-up words. Only the corresponding path needs to be added in the post-processing to support new wake-up words at any time, overcoming the defect that "the voice wake-up method based on the CE loss function transforms the voice wake-up problem from a speech recognition task to an image classification task. It only needs to smooth the posterior probability results of the acoustic model without building a recognition network and searching. However, its accuracy depends on the pre-annotated wake-up word boundary information, and it cannot support the custom wake-up word function."
[0101] The voice wake-up device provided in this embodiment solves the following technical problems:
[0102] 1) When using speech recognition methods to solve the voice wake-up problem, the problem of constructing a weighted finite state transducer (WFT) recognition graph is solved. The memory consumption and computational complexity of constructing and searching the graph are very large.
[0103] 2) Usually, when a single method is used to solve the voice wake-up problem, it is impossible to balance the wake-up effect and the amount of calculation. This embodiment adopts a dual voice wake-up module structure. When the first voice wake-up module determines that the user input audio includes a wake-up word, the second voice wake-up module is used to determine whether to wake up again.
[0104] This embodiment achieves the goal of avoiding the construction of complex recognition graphs and their search calculations, avoiding CTC prefix beam search searches for all acoustic modeling units and the construction of hot word graphs while ensuring the accuracy of voice wake-up and custom wake-up words. In view of the particularity of the voice wake-up problem, the fixed path of the wake-up word and its similar sounds is focused on; and a second voice wake-up module is used for secondary confirmation to further ensure the accuracy of voice wake-up.
[0105] Example 3
[0106] This embodiment provides a computer device, which may be a server or a terminal. Its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data in the voice wake-up method. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the voice wake-up method in Example 1 is implemented.
[0107] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0108] Example 4
[0109] This embodiment provides a computer-readable storage medium storing a computer program, which implements the voice wake-up method in Embodiment 1 when executed by a processor.
[0110] Example 5
[0111] This embodiment provides a computer program product, including a computer program, which implements the voice wake-up method in Embodiment 1 when executed by a processor.
[0112] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0113] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0114] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0115] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0116] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A voice wake-up method, characterized in that: The voice wake-up method comprises: Get the original audio signal; Preprocessing the original audio signal to obtain a preprocessed audio signal; A first voice wake-up method is used to identify potential targets in the preprocessed audio signal to obtain a first recognition result, wherein the first voice wake-up method is an improved voice wake-up method based on a CTC loss function, and the first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on a CTC loss function to fixed path decoding, wherein the fixed path decoding refers to decoding through a target path, wherein the target path includes a fixed path of phonemes in a target wake-up word and a fixed path of phonemes with similar pronunciation to the target wake-up word, and the potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciation to the target wake-up word; When the first recognition result satisfies the first preset condition, a preliminary wake-up is performed, and a second voice wake-up method is used to identify whether the target audio frame contains the real potential target, so as to obtain a second recognition result, wherein the target audio frame refers to an audio frame that may contain the potential target in the first recognition result, and the second voice wake-up method is an improved voice wake-up method based on the HMM model, and the second voice wake-up method includes improving the recognition graph in the traditional voice wake-up method based on the HMM model into a target path recognition graph, and the target path recognition graph is a recognition graph constructed according to the target path; When the second recognition result satisfies the second preset condition, the wake-up is confirmed, and the time of each phoneme of the final target wake-up word in the second recognition result is identified.
2. The voice wake-up method according to claim 1, characterized in that: Preprocessing the original audio signal to obtain a preprocessed audio signal specifically includes: Performing frame processing on the original audio signal to obtain a plurality of audio frames; Calculating the sum of squares of all sampling points in each of the audio frames; Calculate the speech energy of each of the audio frames according to the square sum; Smoothing each of the speech energies to obtain smoothed speech energy; The audio frame whose smoothed speech energy is greater than the energy threshold is selected as the preprocessed audio signal.
3. The voice wake-up method according to claim 1, characterized in that: The first voice wake-up method is used to identify potential targets in the preprocessed audio signal to obtain a first recognition result, specifically including: Extracting features from each audio frame in the preprocessed audio signal to obtain acoustic features of each audio frame; Taking each of the acoustic features as input, and using a pre-trained acoustic model to calculate a posterior probability vector of each target phoneme at each moment, wherein the acoustic model is a deep neural network model based on a CTC loss function, and the target phonemes include blank items and all toneless pinyin items; Calculate the fixed path score of the potential target according to the CTC fixed path score graph, wherein the CTC fixed path score graph includes multiple time steps on the horizontal axis, multiple label groups on the vertical axis, and each time step-label intersection, each of the label groups includes a blank item and a potential target item, and each of the time step-label intersections represents the posterior probability vector of the corresponding label at each moment; Determining a qualified fixed path according to the fixed path score and the path score threshold; Whether to initially wake up is determined based on each phoneme peak probability value and peak probability threshold on the qualified fixed path and the distance and distance threshold of adjacent phoneme peaks, wherein the phoneme peak probability value is a variable calculated based on the posterior probability vector.
4. The voice wake-up method according to claim 3, characterized in that: Each blank posterior probability vector in the CTC fixed path score graph is replaced with a non-potential target posterior probability vector, wherein the blank posterior probability vector refers to the posterior probability vector of each time step corresponding to the blank item, and the non-potential target posterior probability vector refers to the sum of the posterior probability vectors of other target phonemes recognized by the pre-trained acoustic model except the potential target items in the same label group.
5. The voice wake-up method according to claim 3, characterized in that: The calculation expression of the fixed path score of the potential target is: P(S|X)=∑ c∈A(S) P(C|X); Score = -logP(S|X); Where P(C|X) represents the probability of a single path, C represents the path containing blank items and duplicate items, and C = (c1, ...c t , ...c T ), ; X represents the posterior probability vector sequence, the posterior probability vector sequence is calculated according to the posterior probability vector, X = (x1, ...x t ,...x T );y(c t ,t) means the output label is c at time t t The probability of ; P(S|X) is the fixed path probability; S represents the pronunciation sequence of the potential target; A(S) represents the set of all possible paths of S; Score represents the fixed path score.
6. The voice wake-up method according to claim 1, characterized in that: The process of determining the target audio frame specifically includes: According to the peak time point of the first potential target in the preprocessed audio signal, retain the duration of a phoneme pronunciation forward to obtain an extended audio signal; The audio signal before the potential target in the extended audio signal is removed, and the audio signal of the potential target and its subsequent audio signal are retained to obtain the target audio frame.
7. A voice wake-up method and device, characterized in that: The voice wake-up method device comprises: An acquisition module, used for acquiring an original audio signal; A speech energy detection module, used for preprocessing the original audio signal to obtain a preprocessed audio signal; A first voice wake-up module, used to identify potential targets in the preprocessed audio signal by using a first voice wake-up method to obtain a first recognition result, wherein the first voice wake-up method is an improved voice wake-up method based on a CTC loss function, and the first voice wake-up method includes improving the decoding method in the traditional voice wake-up method based on a CTC loss function to fixed path decoding, wherein the fixed path decoding refers to decoding through a preset target path, wherein the target path includes a fixed path of phonemes in a target wake-up word and a fixed path of phonemes with similar pronunciation to the target wake-up word, and the potential target refers to the phonemes of the target wake-up word and the phonemes with similar pronunciation to the target wake-up word; A second voice wake-up module, used for identifying whether the target audio frame contains the real potential target by using a second voice wake-up method, and obtaining a second recognition result, wherein the target audio frame refers to an audio frame that may contain the potential target in the first recognition result, and the second voice wake-up method is an improved voice wake-up method based on the HMM model, and the second voice wake-up method includes improving the recognition graph in the traditional voice wake-up method based on the HMM model into a target path recognition graph, and the target path recognition graph is a recognition graph constructed according to the target path; The final result acquisition module is used to confirm the wake-up according to the second recognition result and identify the time of each phoneme of the final target wake-up word in the second recognition result.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the voice wake-up method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the voice wake-up method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the voice wake-up method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Voice wake-up method and device, storage medium and equipment
CN114220440A
Voice wake-up method, voice wake-up device and storage medium
CN115762480A
Audio recognition method, audio recognition device, vehicle, computer equipment and medium
CN117456999A
Display device and voice recognition method
CN118828083A
System and method that support voice recognition and identity verification based on multi-path CTC alignment
KR102579130B1