Model training method and device, voice wake-up method and device

Through a model training method combining phoneme information and semantic information, the problem of inaccurate speech judgment in the prior art is solved, and the accuracy and reliability of the speech wake-up function are improved.

CN113851113BActive Publication Date: 2025-06-20VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111137419.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-27
Publication Date
2025-06-20
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

In the prior art, inaccurate voice judgment leads to problems such as inaccurate awakening or false awakening.

Method used

A model training method is adopted to obtain the first feature information and the second feature information of the audio training data, combine the acoustic model and generate an adversarial network model, output phoneme information and semantic information, and train the model to improve the accuracy of audio judgment.

Benefits of technology

It improves the accuracy of the acoustic model for audio judgment, reduces the phenomenon of inaccurate or false awakening, and enhances the reliability of the voice wake-up function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113851113B_ABST
    Figure CN113851113B_ABST
Patent Text Reader

Abstract

The present application discloses a model training method and apparatus, a voice wake-up method and apparatus, an electronic device, and a readable storage medium, belonging to the technical field of data processing. Among them, the model training method includes: obtaining first feature information of audio training data, where the audio training data includes wake-up audio and non-wake-up audio; outputting phoneme information and semantic information of the audio training data through a to-be-trained acoustic model, a generative adversarial network model, and the first feature information; outputting second feature information of the audio training data through the to-be-trained generative adversarial network model, and the phoneme information and the semantic information; and training the acoustic model and the generative adversarial network model according to the first feature information and the second feature information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data processing, and particularly relates to a model training method and device, a voice wake-up method and device, an electronic device, and a readable storage medium. Background Art

[0002] Currently, voice interaction has become an important form of human-computer interaction. Among them, the voice wake-up function, as the entry of voice interaction, has been successfully applied to various types of electronic devices, such as smart speakers, smartphones, smart home devices, in-vehicle intelligent devices, and so on.

[0003] For example, the user can successfully wake up the smart speaker by specifying a wake-up word, and then can control the speaker to play audio through voice; for another example, the user can successfully wake up the mobile phone by specifying a wake-up word, and then can control the mobile phone to make a call through voice.

[0004] In the prior art, there often occur phenomena such as failure to wake up or false wake-up due to inaccurate voice judgment. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a model training method, which can solve the problem that in the prior art, there often occur phenomena such as failure to wake up or false wake-up due to inaccurate voice judgment.

[0006] In a first aspect, the embodiments of this application provide a model training method, which includes: obtaining first feature information of audio training data, where the audio training data includes wake-up audio and non-wake-up audio; outputting phoneme information and semantic information of the audio training data through a to-be-trained acoustic model, a generative adversarial network model, and the first feature information; outputting second feature information of the audio training data through the to-be-trained generative adversarial network model, and the phoneme information and the semantic information; and training the acoustic model and the generative adversarial network model according to the first feature information and the second feature information.

[0007] In a second aspect, the embodiments of this application provide a voice wake-up method, which includes: obtaining third feature information of a first audio; outputting first phoneme information of the first audio through the acoustic model and the third feature information; and outputting a wake-up instruction when the first phoneme information matches preset phoneme information of the wake-up audio; where the acoustic model is trained by the model training method described in the first aspect.

[0008] In a third aspect, an embodiment of the present application provides a model training device, which includes: a first acquisition module, configured to acquire first feature information of audio training data, where the audio training data includes wake-up audio and non-wake-up audio; a first output module, configured to output phoneme information and semantic information of the audio training data through a to-be-trained acoustic model, a generative adversarial network model, and the first feature information; a second output module, configured to output second feature information of the audio training data through the to-be-trained generative adversarial network model, and the phoneme information and the semantic information; and a training module, configured to train the acoustic model and the generative adversarial network model according to the first feature information and the second feature information.

[0009] In a fourth aspect, an embodiment of the present application provides a voice wake-up device, which includes: a second acquisition module, configured to acquire third feature information of a first audio; a third output module, configured to output first phoneme information of the first audio through the acoustic model and the third feature information; and a fourth output module, configured to output a wake-up instruction when the first phoneme information matches preset phoneme information of the wake-up audio; where the acoustic model is trained by the model training method described in the first aspect.

[0010] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect or the second aspect are implemented.

[0011] In a sixth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.

[0012] In a seventh aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect or the second aspect.

[0013] Thus, in the embodiments of the present application, in the voice wake-up function, it is necessary to train an acoustic model to ensure a relatively high accuracy rate of the acoustic model in audio judgment. First, a large number of audio including wake-up audio and non-wake-up audio are used as audio training data, and first feature information is extracted. The first feature information is input into the acoustic model, and the phoneme information of the audio training data is output. Secondly, the first feature information is input into the generative adversarial network model, and the semantic information of the audio training data is output. Then, the phoneme information is input into the generative adversarial network model, and the generative adversarial network model combines the semantic information and the phoneme information to output the second feature information of the audio training data. Further, based on the output second feature information and the first feature information, the acoustic model and the generative adversarial network model are trained to minimize the difference between the second feature information and the first feature information. It can be seen that in the embodiments of the present application, the method mainly combines two audio features of phoneme information and semantic information to enhance the representation of audio semantic feature information, so as to realize the model training in the whole function, thereby achieving the training purpose of the acoustic model, making the accuracy rate of the acoustic model in judging audio relatively high, and further improving the accuracy rate of judging wake-up audio, and avoiding the phenomena of failure to wake up or false wake-up. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a flowchart of the model training method according to an embodiment of the present application;

[0015] Figure 2 is an explanatory schematic diagram of the model training method according to an embodiment of the present application;

[0016] Figure 3 is a schematic diagram of the network structure of the model training method according to an embodiment of the present application;

[0017] Figure 4 is a flowchart of the voice wake-up method according to an embodiment of the present application;

[0018] Figure 5 is a block diagram of the model training device according to an embodiment of the present application;

[0019] Figure 6 is a block diagram of the voice wake-up device according to an embodiment of the present application;

[0020] Figure 7 is a schematic diagram of the hardware structure of the electronic device according to an embodiment of the present application;

[0021] Figure 8 is a schematic diagram of the hardware structure of the electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0023] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means that the related objects before and after are in an "or" relationship.

[0024] Next, in conjunction with the accompanying drawings, the model training method provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0025] See Figure 1 , which shows a flowchart of the model training method according to an embodiment of the present application. The method is applied to an electronic device and includes:

[0026] Step 110: Obtain first feature information of audio training data, where the audio training data includes wake-up audio and non-wake-up audio.

[0027] Among them, the wake-up audio is the audio used to output a wake-up instruction, and other audio except the wake-up audio is non-wake-up audio.

[0028] In this embodiment, based on the audio training data including wake-up audio and non-wake-up audio, the model in the wake-up function is trained to improve the accuracy of audio judgment in the wake-up function.

[0029] The first feature information is the sum of the feature information obtained from a large number of audios in the audio training data.

[0030] In this embodiment, the first feature information is used to represent the Fbank feature of the audio training data.

[0031] Optionally, based on the audio training data, Fbank feature extraction is performed on the training corpus. Generally, features of 80 dimensions can be extracted, and the sampling rate is 16KHz.

[0032] Step 120: Output the phoneme information and semantic information of the audio training data through the acoustic model to be trained, the generative adversarial network model, and the first feature information.

[0033] Among them, the phoneme information of the audio training data is output through the acoustic model to be trained.

[0034] It should be noted that the acoustic model is a model used to recognize sounds in speech recognition or speech wake-up.

[0035] In this step, the first feature information is used as the input, and through the acoustic model, the phoneme information of the audio training data is output.

[0036] Optionally, the phoneme information includes a phoneme probability matrix. Among them, for each audio in the audio training data, each frame corresponds to a group of phoneme probability sequences.

[0037] In addition, the semantic information of the audio training data is output through the generative adversarial network model to be trained.

[0038] Optionally, the generative adversarial network model in this embodiment is based on a conditional variational auto-encoder (C-VAE). It can be considered that the generative adversarial network model includes an encoder.

[0039] Therefore, in this step, the first feature information is used as the input, and through the encoder, the semantic information of the audio training data is output.

[0040] Among them, the semantic information is the sum of the semantic information obtained from a large number of audios in the audio training data.

[0041] Exemplarily, through the encoder, the semantic information corresponding to each frame of each audio in the audio training data can be obtained.

[0042] The second feature information in this embodiment is a semantic representation latent variable.

[0043] Step 130: Output the second feature information of the audio training data through the generative adversarial network model to be trained, and the phoneme information and semantic information.

[0044] In this embodiment, a variational adversarial (VAWGAN) network is used. The VAWGAN network is a deep learning model and one of the most promising methods for unsupervised learning on complex distributions in recent years. The model generates quite good outputs through the mutual game learning of at least two modules in the framework: the generative model and the discriminative model.

[0045] Optionally, the generation module includes a generator, and the discrimination module includes a discriminator.

[0046] In this step, the output of the acoustic model (i.e., the phoneme information of the audio training data) and the output of the encoder (i.e., the semantic information of the audio training data) are concatenated and input into the generator, and the second feature information of the audio training data is output.

[0047] In this embodiment, the second feature information is used to represent the fake feature of the audio training data.

[0048] Among them, the first feature information is the real feature obtained based on the audio training data, while the second feature information is the synthetic feature obtained based on the model output.

[0049] See Figure 2 , the output z of the encoder represents the semantic representation latent variable; the output A(x) of the acoustic model represents the phoneme posterior probability matrix, the horizontal axis is the time dimension, and the vertical axis is the phoneme posterior probability sequence. Since both of them can represent the semantic information representation of the audio, combining them can better enhance the representation of the audio semantic feature information, so that the acoustic model generated by modeling can better represent the phoneme probability of each frame of the audio.

[0050] Step 140: Train the acoustic model and the generative adversarial network model according to the first feature information and the second feature information.

[0051] In this step, the acoustic model and the VAWGAN network are trained to adjust the parameters in each model, and finally the trained acoustic model and the VAWGAN network are obtained.

[0052] During the training process, the parameters of the acoustic model A, the encoder parameter φ, the generator parameter θ, and the discriminator parameter ψ are optimized respectively.

[0053] Among them, in this step, the purpose of training is to minimize the difference between the synthesized second feature information and the real first feature information. In this way, the audio recognized by the acoustic model is the closest to the real audio, thereby improving the judgment accuracy of the audio.

[0054] Thus, in the embodiments of the present application, in the voice wake-up function, it is necessary to train the acoustic model to ensure a relatively high accuracy rate of the acoustic model in audio judgment. First, a large number of audio including wake-up audio and non-wake-up audio are used as audio training data, and first feature information is extracted. The first feature information is input into the acoustic model, and the phoneme information of the audio training data is output. Secondly, the first feature information is input into the generative adversarial network model, and the semantic information of the audio training data is output. Then, the phoneme information is input into the generative adversarial network model, and the generative adversarial network model combines the semantic information and the phoneme information to output the second feature information of the audio training data. Further, based on the output second feature information and the first feature information, the acoustic model and the generative adversarial network model are trained to minimize the difference between the second feature information and the first feature information. It can be seen that in the embodiments of the present application, the method mainly combines two audio features of phoneme information and semantic information to enhance the representation of audio semantic feature information, so as to realize the model training in the whole function, thereby achieving the training purpose of the acoustic model, making the accuracy rate of the acoustic model in judging audio relatively high, and further improving the accuracy rate of judging wake-up audio, and avoiding the phenomena of not being able to wake up or mis-wake-up.

[0055] In the process of the model training method in another embodiment of the present application, step 140 includes:

[0056] Sub-step A1: Train the acoustic model and the generative adversarial network model until the first error rate between the first feature information and the second feature information meets the first preset condition.

[0057] In this step, the first feature information and the second feature information are input into the discriminator, and the difference between the first feature information and the second feature information is output.

[0058] Optionally, the difference between the first feature information and the second feature information is represented by the first error rate.

[0059] For this embodiment, one explanation is that the training purpose ultimately to be achieved in the present application is to make the first error rate between the first feature information and the second feature information less than a certain threshold. Therefore, the first preset condition is: the first error rate is less than the threshold.

[0060] For this embodiment, another explanation is that the training purpose ultimately to be achieved in the present application is to reach the preset number of iterations so that the first error rate between the first feature information and the second feature information no longer changes and reaches the minimum value. Therefore, the first preset condition is: the error rate under the preset number of iterations.

[0061] Exemplarily, in an experiment, the number of iterations is selected to be 200,000 times.

[0062] In this embodiment, based on the first preset condition, the final training effect is achieved, minimizing the difference between the first feature information and the second feature information, thereby improving the judgment accuracy of the acoustic model for audio.

[0063] In the process of the model training method according to another embodiment of the present application, before step 120, the method further includes:

[0064] Step B1: Train the acoustic model until the matching rate between the phoneme information of the audio training data and the preset phoneme information meets the fourth preset condition.

[0065] Optionally, the audio training data further includes text annotations corresponding to each audio.

[0066] Optionally, using the trained speech recognition network with high accuracy, combined with the text annotations corresponding to each audio, align the audio training data to obtain the phoneme label corresponding to each frame of each audio in the audio training data. Further, all the phoneme labels form the preset phoneme information of this embodiment.

[0067] Among them, the preset phoneme information includes the phoneme label corresponding to each frame of each audio.

[0068] The matching rate in this embodiment is the total matching rate obtained based on the matching between the phoneme probability sequence of each frame in the audio training data and the phoneme label of the corresponding frame.

[0069] Optionally, the training process of this embodiment is:

[0070] Establish a mapping relationship between the Fbank feature x of the audio training data and the preset phoneme information.

[0071] First step, through the acoustic model and the first feature information, output the phoneme probability sequence of each frame; through the speech recognition network, obtain the phoneme label of each frame.

[0072] Second step, use the cross-entropy loss function to measure the error loss function between the inferred phoneme probability sequence z p and the phoneme label:

[0073]

[0074] z p =[p i1 ,p i2 ,...,p ic (2)

[0075] Among them, M is the sum of all phoneme labels; y ic is the sign function (0 or 1) of the phoneme label. If the phoneme label of the i-th frame is equal to c, take 1, otherwise take 0; pic is the predicted probability that the i-th frame belongs to c; z p is the phoneme probability sequence.

[0076] Among them, z p is inferred from the input Fbank features by the acoustic model:

[0077] z p = A(x) (3)

[0078] Among them, A are the parameters of the acoustic model. During the training process of the acoustic model, through continuous iteration, the cross-entropy loss in (1) is minimized, so that the acoustic model converges continuously.

[0079] Among them, the matching rate satisfies the fourth preset condition, corresponding to: the error L between the phoneme probability sequence z p and the phoneme label is minimized.

[0080] In this embodiment, before using the phoneme information of the output audio training data as the input of the encoder, the acoustic model can be preliminarily trained according to the above training method, so that the difference between the phoneme information of the audio training data obtained by the acoustic model and the preset phoneme information is minimized. It can be seen that based on the training method provided in this embodiment, combined with the training method provided in the previous embodiment, the purpose of more refined training of the acoustic model can be achieved to ensure that the accuracy of the acoustic model for audio judgment is as high as possible.

[0081] In the process of the model training method in another embodiment of the present application, the generative adversarial network model includes a discriminator module and a generator module. Step 140 includes:

[0082] Sub-step C1: Train the generator module and the acoustic model until the second feature information output by the generator module meets the second preset condition.

[0083] Sub-step C2: Train the discriminator module until the second error rate between the first feature information and the second feature information output by the generator module meets the third preset condition.

[0084] As can be seen from the foregoing embodiments, in the present application, based on the variational autoencoder, the VAWGAN network is incorporated into the decoder to improve the VAE effect. Among them, VAWGAN includes two parts, one part is a generator for generating synthetic spectra, and the other part is a discriminator for judging whether the synthetic spectra are real spectra. It can be understood that: the decoder includes a generator and a discriminator.

[0085] In the VAWGAN network, the objective function is:

[0086] J vawgan = L(x; φ,θ) + αJwgan (4)

[0087] Among them, \(L(x;\varphi,\theta)\) is the objective function of the encoder part:

[0088]

[0089] Among them, \(D\) KL (q φ (z|x)||p θ (z)) represents the relative entropy (Kullback-Leibler Divergence, abbreviated as KL divergence) between the discriminative module \(q\) φ (z|x) and the true posterior probability \(p(z|x)\). The prior probability \(p\) θ (z) is a standard multi-dimensional Gaussian distribution. \(q\) φ (z|x) and \(p\) θ (x|z) are the encoder and decoder respectively, following a multi-dimensional Gaussian distribution, and their mean vectors and covariance matrices are \((\mu\) φ (z),\(\sigma\) φ (z)) and \((\mu\) θ (x),\(\sigma\) θ (x)) respectively. Therefore, the two terms on the right can be simplified to:

[0090]

[0091]

[0092] Among them, \(K\) is the dimension of the intermediate variable \(z\), and \(L\) is the number of times of sampling \(q\) φ (z|x). Since the sampling process is a discontinuous operation and cannot be differentiated, the network parameters of the encoder and decoder cannot be updated through backpropagation. Therefore, another random variable \(\varepsilon\) is introduced to reparameterize the hidden variable \(z\), making \(z\) (l) =\(\mu\) θ (x)+\(\varepsilon\) (l) *\(\sigma\) θ (x), \(\varepsilon(l)\sim N(0,I)\), then:

[0093]

[0094] Among them, \(D\) is the number of samples of \(x\).

[0095] So far, the objective loss function for optimizing the VAWGAN network can be obtained.

[0096] Among them, the parameters of the acoustic model \(A(x)\) change dynamically according to the loss function during the training process, making the model converge continuously; the output \(z\) of the encoder changes dynamically according to the output of the encoder.

[0097] Based on the above content, continue to explain how the acoustic model A(x) and VAWGAN are trained simultaneously to make the acoustic model A(x) achieve better results.

[0098] J wgan represents the objective function of the VAWGAN part:

[0099]

[0100] where α is the loss coefficient of the VAWGAN, and D Ψ is the discrimination output of the discriminator on the authenticity of the features. Combine A(x) with z and send it into the generator, and then let the discriminator judge. The second half of the above formula is the loss function of the generator's two-dimensional convolutional neural network:

[0101]

[0102] Since the acoustic model A(x) needs to continuously optimize its parameters during this process, the objective function for optimizing the generator becomes:

[0103]

[0104] where min represents minimizing the losses of the generator and the acoustic model, and solving for the optimal parameters of the generator and the acoustic model A; the second half of the above formula is the loss function of the acoustic model, which needs to be combined with the generator's loss function to make the overall loss reach the optimal value.

[0105] Since the loss function optimization of the acoustic model is added to the generator, the loss function of the discriminator's two-dimensional convolutional neural network becomes:

[0106]

[0107] The objective function for optimizing the discriminator is:

[0108]

[0109] where max represents maximizing the discriminator's loss function, that is, the discriminator's goal is to maximize the gap between distinguishing real features and fake features, so as to continuously optimize the discriminator's model parameters.

[0110] In this embodiment, the decoder consists of a generator and a discriminator. During the training process, first fix the discriminator parameters, and train the generator and the acoustic model to make the overall loss function L of the generator G as small as possible, that is, the second feature information meets the second preset condition, and obtain the generated Fbank feature x′ (i.e., the second feature information); then fix the generator and acoustic model parameters, and train the discriminator to make the discriminator's loss function L D as large as possible, that is, -LD Minimize, that is, the second error rate between the first feature information and the second feature information satisfies a third preset condition.

[0111] It should be noted that in one interpretation, the two steps in this embodiment are alternately repeated. For example, for the first time: based on the first parameters of the discriminator, perform step C1 to train the generator and the acoustic model, so that the overall loss function L of the generator G is as small as possible; further, based on the parameters of the generator and the acoustic model after this training, perform step C2 to train the discriminator; for the second time, because in step C2 of the first time, the parameters of the discriminator are adjusted, therefore, based on the adjusted second parameters, train the generator and the acoustic model, so that the overall loss function L of the generator G is as small as possible; further, based on the parameters of the generator and the acoustic model after this training, perform step C2 to train the discriminator. And so on until the number of iterations is completed.

[0112] Correspondingly, the second error rate is used to represent the error rate obtained in one repetition of the steps, and the first error rate is used to represent the error rate finally obtained through training.

[0113] In another interpretation, the two steps in this embodiment respectively represent two types of steps. For example, step C1 represents all the steps of training the generator and the acoustic model, and can be an overview of multiple repeated steps; step C2 represents all the steps of training the discriminator and can be an overview of multiple repeated steps.

[0114] Correspondingly, both the first error rate and the second error rate are used to represent the error rate finally obtained through training.

[0115] Under this interpretation, there is no limitation on the order between step C1 and step C2.

[0116] Among them, the generator adopts a two-dimensional convolutional neural network, including 4 convolutional layers. The filter sizes of the 4 convolutional layers are 9*1, 7*1, 7*1, 1025*1 respectively, the strides are 3, 3, 3, 1 respectively, the filter depths are 32, 16, 8, 1 respectively, and the activation function adopts the LReLU function.

[0117] The discriminator adopts a two-dimensional convolutional neural network, including 3 convolutional layers and 1 fully connected layer. The filter sizes of the 3 convolutional layers are 7*1, 7*1, 115*1 respectively, the strides are all 3, the filter depths are 16, 32, 64 respectively, and the activation function adopts the LReLU function.

[0118] The acoustic model and the encoder have the same structure. A two-dimensional convolutional neural network is adopted, including 5 convolutional layers and 1 fully-connected layer. The filter sizes of the 5 convolutional layers are all 7*1, the strides are all 3, and the filter depths are 16, 32, 64, 128, and 256 respectively. The activation function uses the LReLU function. The network structure is as shown in Figure 3 shown. During the training process, the Stochastic Gradient Descent (SGD) method is used to update the network model parameters.

[0119] In this embodiment, a method for training an acoustic model for voice wake-up based on a generative adversarial network is provided to improve the modeling effect of the acoustic model. Among them, a generative adversarial network based on a variational autoencoder and an acoustic model are used in combination to implement the training of the voice wake-up system. Combining the VAWGAN network in the acoustic model can better improve the modeling quality of the acoustic model and achieve high-quality voice wake-up.

[0120] In the process of the model training method in another embodiment of this application, step 120 includes:

[0121] Sub-step D1: Through the acoustic model to be trained and the first feature information, output the phoneme information corresponding to the target frames of each audio in the audio training data.

[0122] Sub-step D2: Through the generative adversarial network model to be trained and the first feature information, output the semantic information corresponding to the target frames of each audio in the audio training data.

[0123] Optionally, the target frames include each frame of each audio in the audio training data.

[0124] Optionally, the target frames include partial frames of each audio in the audio training data. Among them, the partial frames can be collected at a certain frequency to ensure that the target frames are evenly distributed in the audio training data.

[0125] Correspondingly, based on the obtained phoneme information of the target frames, the semantic information of the corresponding frames is further obtained, so as to combine the phoneme information and semantic information of any frame to generate the fake feature corresponding to this frame for comparison with the Fbank feature corresponding to this frame.

[0126] Furthermore, each frame in the target frames is sequentially subjected to feature comparison, so as to complete the feature comparison of the entire audio training data.

[0127] In this embodiment, a method for obtaining the phoneme information and semantic information of audio training data is provided to describe this embodiment in more detail. Among them, for the audio training data in this embodiment, the corresponding phoneme information and semantic information are regularly obtained for the target frames therein to generate the composite features of the frames, so as to be used for comparison with the true features of the frames. It can be seen that based on the feature comparison of the target frames, this embodiment can infer the overall situation of the audio training data for model training in this application.

[0128] See Figure 4 , which shows the flowchart of the voice wake-up method according to another embodiment of the present application, applied to an electronic device. The method includes:

[0129] Step 150: Obtain the third feature information of the first audio.

[0130] Among them, the model training method and the voice wake-up method provided in this application are respectively applied to two stages. The first stage is the training stage in the foregoing embodiment, and the other stage is the wake-up stage in this embodiment.

[0131] In this step, the third feature information is used to represent the Fbank feature of the first audio.

[0132] In this embodiment, the first audio can be an audio stream. Therefore, an audio stream is sent into a storage buffer (buffer). The general frame length is 10 ms. In order to reduce the calculation amount, a frame skipping method can be used (such as sending 1 frame every 3 frames), and then its features are extracted.

[0133] Step 160: Output the first phoneme information of the first audio through an acoustic model and the third feature information.

[0134] Among them, the acoustic model is trained by the model training method in any of the foregoing embodiments.

[0135] The extracted Fbank features are input into the acoustic model trained in the foregoing embodiment for inference to obtain the first phoneme information of the corresponding first audio.

[0136] Among them, the first phoneme information includes a phoneme probability matrix.

[0137] Step 170: Output a wake-up instruction when the first phoneme information matches the preset phoneme information of the wake-up audio.

[0138] Among them, the wake-up instruction is used to wake up the terminal device and is applied to the voice wake-up function.

[0139] This step corresponds to the Viterbi decoding step.

[0140] In this step, the phoneme probability matrix obtained in step 160 is sent into the decoding graph of the wake-up audio, and decoded using the Viterbi algorithm to obtain a score. Determine whether the score is greater than a certain threshold. If "yes", wake up; if "no", continue to send the next frame of data.

[0141] Among them, the score here can be understood as: the degree of association between the first phoneme information and the preset phoneme information of the wake-up audio. If the degree of association is greater than a certain threshold, the first phoneme information matches the preset phoneme information of the wake-up audio.

[0142] Exemplarily, based on the first phoneme information, the phoneme label with the highest probability corresponding to each frame in the input audio stream can be obtained through decoding, and compared with the preset phoneme labels corresponding to each frame in the wake-up audio. If the similarity is greater than a certain set value, the degree of association between the first phoneme information and the preset phoneme information of the wake-up audio is greater than a certain threshold.

[0143] In this way, based on the foregoing embodiments, by combining two audio features of phoneme information and semantic information, the representation of audio semantic feature information is enhanced to implement model training in the entire function. Therefore, in this embodiment, in the wake-up stage, through the inference calculation of the trained acoustic model for the received first audio, more accurate phoneme information of the first audio can be obtained, so that after comparing the first audio with the wake-up audio, terminal devices such as mobile phones can be accurately and timely woken up.

[0144] In summary, the modeling process of voice wake-up is usually to train an acoustic model to establish the mapping between voice features and phonemes, and then use the optimal path algorithm for decoding. However, due to the requirements of low power consumption and fast response, the resources of the acoustic model are limited, and it is often easy for the acoustic model to make inaccurate judgments, resulting in the situation of not being able to wake up or being woken up by mistake. Based on this, in this application, in the training stage of the voice wake-up acoustic model, a generative adversarial network based on a variational autoencoder is used, combined with the acoustic model, which can better improve the modeling quality of the acoustic model, make phoneme inference more accurate, reduce false wake-up, and improve the wake-up rate.

[0145] It should be noted that for the model training method provided in the embodiments of this application, the execution subject can be a model training device, or a control module in the model training device for executing the model training method. In the embodiments of this application, the model training method is executed by the model training device as an example to illustrate the model training device provided in the embodiments of this application.

[0146] Figure 5 The block diagram of the model training device according to another embodiment of this application is shown, and the device includes:

[0147] The first acquisition module 10 is configured to acquire first feature information of audio training data, where the audio training data includes wake-up audio and non-wake-up audio;

[0148] The first output module 20 is configured to output phoneme information and semantic information of the audio training data through a to-be-trained acoustic model, a generative adversarial network model, and the first feature information;

[0149] The second output module 30 is configured to output second feature information of the audio training data through the to-be-trained generative adversarial network model, as well as the phoneme information and the semantic information;

[0150] The training module 40 is configured to train the acoustic model and the generative adversarial network model according to the first feature information and the second feature information.

[0151] In this way, in the embodiments of the present application, in the voice wake-up function, it is necessary to train the acoustic model to ensure a high accuracy rate of the acoustic model in audio judgment. First, a large amount of audio including wake-up audio and non-wake-up audio is used as audio training data, and the first feature information is extracted. The first feature information is input into the acoustic model to output the phoneme information of the audio training data. Secondly, the first feature information is input into the generative adversarial network model to output the semantic information of the audio training data. Then, the phoneme information is input into the generative adversarial network model, and the generative adversarial network model combines the semantic information and the phoneme information to output the second feature information of the audio training data. Further, based on the output second feature information and the first feature information, the acoustic model and the generative adversarial network model are trained to minimize the difference between the second feature information and the first feature information. It can be seen that in the embodiments of the present application, the method mainly combines two audio features of phoneme information and semantic information to enhance the representation of audio semantic feature information, so as to realize the model training in the whole function, thereby achieving the training purpose of the acoustic model, making the accuracy rate of the acoustic model in judging audio higher, and further improving the accuracy rate of judging wake-up audio, and avoiding the phenomena of not being able to wake up or being misawakened.

[0152] Optionally, the training module 40 includes:

[0153] The first training unit is configured to train the acoustic model and the generative adversarial network model until a first error rate between the first feature information and the second feature information meets a first preset condition.

[0154] Optionally, the generative adversarial network model includes a discrimination module and a generation module; the training module 40 includes:

[0155] The second training unit is configured to train the generation module and the acoustic model until the second feature information output through the generation module meets a second preset condition;

[0156] A third training unit for training the discrimination module until a second error rate between the first feature information and the second feature information output by the generation module meets a third preset condition.

[0157] Optionally, the first output module 20 includes:

[0158] A first output unit for outputting phoneme information corresponding to target frames of each audio in the audio training data through the acoustic model to be trained and the first feature information;

[0159] A second output unit for outputting semantic information corresponding to target frames of each audio in the audio training data through the generative adversarial network model to be trained and the first feature information.

[0160] It should be noted that for the voice wake-up method provided in the embodiments of the present application, the execution subject may be a voice wake-up device, or a control module in the voice wake-up device for executing the voice wake-up method. In the embodiments of the present application, taking the voice wake-up device executing the voice wake-up method as an example, the voice wake-up device provided in the embodiments of the present application is described.

[0161] Figure 6 The block diagram of a model training device according to another embodiment of the present application is shown. The device includes:

[0162] A second acquisition module 50 for acquiring third feature information of a first audio;

[0163] A third output module 60 for outputting first phoneme information of the first audio through the acoustic model and the third feature information;

[0164] A fourth output module 70 for outputting a wake-up instruction when the first phoneme information matches the preset phoneme information of the wake-up audio;

[0165] Wherein, the acoustic model is trained by the model training method in any of the foregoing embodiments.

[0166] In this way, based on the foregoing embodiments, by combining two audio features of phoneme information and semantic information, the representation of audio semantic feature information is enhanced to implement model training in the entire function. Therefore, in this embodiment, during the wake-up phase, through the trained acoustic model to perform inference calculation on the received first audio, more accurate phoneme information of the first audio can be obtained, so that after comparing the first audio with the wake-up audio, the terminal device such as a mobile phone can be accurately and timely woken up.

[0167] The model training device / voice wake-up device in the embodiments of the present application can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0168] The model training device / voice wake-up device in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0169] The model training device / voice wake-up device provided in the embodiments of the present application can implement each process implemented in the corresponding method embodiments above. To avoid repetition, it will not be elaborated here.

[0170] Optionally, as Figure 7 shown, the embodiments of the present application further provide an electronic device 100, including a processor 101, a memory 102, a program or instruction stored on the memory 102 and executable on the processor 101. When the program or instruction is executed by the processor 101, it implements each process of any of the above model training methods / voice wake-up method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0171] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0172] Figure 8 Schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0173] The electronic device 1000 includes, but is not limited to: a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010, etc.

[0174] Those skilled in the art can understand that the electronic device 1000 may further include a power source (such as a battery) for powering each component. The power source can be logically connected to the processor 1010 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system. Figure 8 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0175] Among them, in one scenario, the processor 1010 is configured to obtain first feature information of audio training data, where the audio training data includes wake-up audio and non-wake-up audio; output phoneme information and semantic information of the audio training data through a to-be-trained acoustic model, a generative adversarial network model, and the first feature information; output second feature information of the audio training data through the to-be-trained generative adversarial network model, and the phoneme information and the semantic information; and train the acoustic model and the generative adversarial network model according to the first feature information and the second feature information.

[0176] In this way, in the embodiments of the present application, in the voice wake-up function, it is necessary to train the acoustic model to ensure a high accuracy rate of the acoustic model in audio judgment. First, a large amount of audio including wake-up audio and non-wake-up audio is used as audio training data, and first feature information is extracted. The first feature information is input into the acoustic model to output phoneme information of the audio training data. Secondly, the first feature information is input into the generative adversarial network model to output semantic information of the audio training data. Then, the phoneme information is input into the generative adversarial network model. The generative adversarial network model combines the semantic information and the phoneme information to output second feature information of the audio training data. Further, based on the output second feature information and the first feature information, the acoustic model and the generative adversarial network model are trained to minimize the difference between the second feature information and the first feature information. It can be seen that in the embodiments of the present application, the method mainly combines two audio features of phoneme information and semantic information to enhance the representation of audio semantic feature information, so as to implement model training in the entire function, thereby achieving the training purpose of the acoustic model, making the accuracy rate of the acoustic model in audio judgment higher, and further improving the accuracy rate of judging wake-up audio, and avoiding the phenomenon of failure to wake up or false wake-up.

[0177] Optionally, the processor 1010 is further configured to train the acoustic model and the generative adversarial network model until a first error rate between the first feature information and the second feature information meets a first preset condition.

[0178] Optionally, the generative adversarial network model includes a discriminative module and a generative module; the processor 1010 is further configured to train the generative module and the acoustic model until the second feature information output by the generative module meets a second preset condition; and train the discriminative module until a second error rate between the first feature information and the second feature information output by the generative module meets a third preset condition.

[0179] Optionally, the processor 1010 is further configured to output phoneme information corresponding to target frames of each audio in the audio training data through the acoustic model to be trained and the first feature information; and output semantic information corresponding to the target frames of each audio in the audio training data through the generative adversarial network model to be trained and the first feature information.

[0180] In another scenario, the processor 1010 is configured to obtain third feature information of a first audio; output first phoneme information of the first audio through the acoustic model and the third feature information; and output a wake-up instruction when the first phoneme information matches preset phoneme information of the wake-up audio, where the acoustic model is trained by the foregoing scenario.

[0181] In this way, based on the foregoing embodiments, by combining two audio features of phoneme information and semantic information, the representation of audio semantic feature information is enhanced to implement model training in the entire function. Therefore, in this embodiment, in the wake-up stage, through the trained acoustic model to perform inference calculation on the received first audio, more accurate phoneme information of the first audio can be obtained, so that after comparing the first audio with the wake-up audio, the terminal device such as a mobile phone can be accurately and timely woken up.

[0182] In summary, the modeling process of voice wake-up usually trains an acoustic model to establish a mapping between voice features and phonemes, and then uses the optimal path algorithm for decoding. However, due to the requirements of low power consumption and fast response, the resources of the acoustic model are limited, and it is often easy for the acoustic model to make inaccurate judgments, resulting in the situation of not being able to wake up or being woken up by mistake. Based on this, in this application, in the voice wake-up acoustic model training stage, a generative adversarial network based on a variational autoencoder is used, combined with the acoustic model, which can better improve the modeling quality of the acoustic model, make phoneme inference more accurate, reduce false wake-up, and improve the wake-up rate.

[0183] It should be understood that in the embodiments of the present application, the input unit 1004 may include a Graphics Processing Unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes the image data of static pictures or video processed by an image capture device (such as a camera) in the video processing capture mode or the image capture mode. The display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1007 includes a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. The other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an action bar, which will not be elaborated here. The memory 1009 may be used to store software programs and various data, including but not limited to application programs and operating systems. The processor 1010 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1010.

[0184] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned model training method / voice wake-up method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0185] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer-readable storage medium, such as a computer Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc.

[0186] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned model training method / voice wake-up method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0187] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, a system chip, a chip system, or a system-on-chip, etc.

[0188] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0189] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A model training method, characterized in that, The method includes: Obtaining first feature information of audio training data, where the audio training data includes wake-up audio and non-wake-up audio; Outputting phoneme information of the audio training data through the acoustic model to be trained and the first feature information; Outputting semantic information of the audio training data through the encoder of the generative adversarial network model to be trained and the first feature information; Outputting second feature information of the audio training data through the generator of the generative adversarial network model to be trained, the phoneme information, and the semantic information; Training the acoustic model and the generative adversarial network model according to the first feature information and the second feature information.

2. The method according to claim 1, characterized in that, The training of the acoustic model and the generative adversarial network model according to the first feature information and the second feature information includes: Training the acoustic model and the generative adversarial network model until a first error rate between the first feature information and the second feature information meets a first preset condition.

3. The method according to claim 1, characterized in that, The generative adversarial network model includes a discriminative module and a generative module; the training of the acoustic model and the generative adversarial network model according to the first feature information and the second feature information includes: Training the generative module and the acoustic model until the second feature information output by the generative module meets a second preset condition; Training the discriminative module until a second error rate between the first feature information and the second feature information output by the generative module meets a third preset condition.

4. The method according to claim 1, characterized in that, The outputting of the phoneme information of the audio training data through the acoustic model to be trained and the first feature information includes: Outputting phoneme information corresponding to a target frame of each audio in the audio training data through the acoustic model to be trained and the first feature information.

5. The method according to claim 4, characterized in that, The outputting of the semantic information of the audio training data through the encoder of the generative adversarial network model to be trained and the first feature information includes: Outputting semantic information corresponding to the target frame of each audio in the audio training data through the generative adversarial network model to be trained and the first feature information.

6. A voice wake-up method, characterized in that, The method includes: Obtaining third feature information of a first audio; Outputting first phoneme information of the first audio through the acoustic model and the third feature information; Outputting a wake-up instruction when the first phoneme information matches preset phoneme information of the wake-up audio; wherein, the acoustic model is trained by the model training method according to any one of claims 1-5.

7. A model training device, characterized in that, The device includes: A first acquisition module, configured to obtain first feature information of audio training data, where the audio training data includes wake-up audio and non-wake-up audio; A first output module, configured to output phoneme information of the audio training data through the acoustic model to be trained and the first feature information; and output semantic information of the audio training data through the encoder of the generative adversarial network model to be trained and the first feature information; A second output module, configured to output second feature information of the audio training data through a generator of the to-be-trained generative adversarial network model, the phoneme information, and the semantic information. A training module, configured to train the acoustic model and the generative adversarial network model according to the first feature information and the second feature information.

8. The device according to claim 7, characterized in that, The training module includes: A first training unit, configured to train the acoustic model and the generative adversarial network model until a first error rate between the first feature information and the second feature information meets a first preset condition.

9. The device according to claim 7, characterized in that, The generative adversarial network model includes a discrimination module and a generation module. The training module includes: A second training unit, configured to train the generation module and the acoustic model until the second feature information output through the generation module meets a second preset condition. A third training unit, configured to train the discrimination module until a second error rate between the first feature information and the second feature information output through the generation module meets a third preset condition.

10. The device according to claim 7, characterized in that,The first output module includes: A first output unit, configured to output phoneme information corresponding to a target frame of each audio in the audio training data through the to-be-trained acoustic model and the first feature information.

11. The device according to claim 10, wherein The first output module includes: A second output unit, configured to output semantic information corresponding to the target frame of each audio in the audio training data through the to-be-trained generative adversarial network model and the first feature information.

12. A voice wake-up device, wherein The apparatus includes: A second acquisition module, configured to acquire third feature information of a first audio. A third output module, configured to output first phoneme information of the first audio through the acoustic model and the third feature information. A fourth output module, configured to output a wake-up instruction when the first phoneme information matches preset phoneme information of the wake-up audio. Wherein, the acoustic model is trained by the model training method according to any one of claims 1-5.

13. An electronic device, wherein It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the model training method according to any one of claims 1-5 or the voice wake-up method according to claim 6 are implemented.

14. A readable storage medium, wherein A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the model training method according to any one of claims 1-5 or the voice wake-up method according to claim 6 are implemented.