A speech recognition method, apparatus, device, and storage medium

Through the adversarial multi-task training speech recognition model, the impact of echo phenomenon on speech recognition accuracy is solved, and highly robust speech recognition in echo scenarios is achieved, adapting to different users and device statuses, and protecting user privacy.

CN114664288BActive Publication Date: 2025-07-08HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011524726.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-22
Publication Date
2025-07-08
Estimated Expiration
2040-12-22

AI Technical Summary

Technical Problem

The echo phenomenon affects the accuracy of the speech recognition device in the closed space, resulting in the user's voice being unable to be accurately recognized.

Method used

Adversarial multi-task training is adopted to obtain a robust speech recognition model in echo scenarios through joint training of feature extraction network, speech recognition network and domain classification network, and train it using the speech sample data of the source domain and the target domain to form an adversarial relationship to improve the robustness of speech recognition.

Benefits of technology

In echo scenes, the accuracy of speech recognition is improved, the impact of echo phenomena is reduced, the status of different users and devices is adapted to different users and devices, the privacy of users is protected, and the robustness of speech recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114664288B_ABST
    Figure CN114664288B_ABST
Patent Text Reader

Abstract

The present application provides a speech recognition method, apparatus, device, and storable medium, relating to the field of artificial intelligence technology, and particularly to the field of speech recognition. The method includes: training an acoustic model by means of adversarial multi-task training, where the network structure includes a feature extraction network, a speech recognition network, and a domain classification network. First, speech data in a non-echo scenario is collected as the speech sample data of the source domain, and speech data in an echo scenario is collected as the speech sample data of the target domain. The speech recognition network is trained using the speech sample data of the source domain, and at the same time, the feature extraction network and the domain classification network with an adversarial relationship are trained using the speech sample data of the source domain and the speech sample data of the target domain, so that the features extracted by the feature extraction network are domain-invariant and recognizable by the speech recognition network. Finally, a speech recognition model with robustness in an echo scenario is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a speech recognition method, apparatus, device, and storage medium. Background Art

[0002] With the continuous development of computer science and technology, especially artificial intelligence (AI) technology, speech recognition technology has begun to move from the laboratory to the market and is being applied in more and more fields, such as industrial control, smart home, smart toys, voice control of terminal devices, etc. Speech recognition technology makes the acquisition and processing of information more convenient, improves the work efficiency of users, and brings convenience to people's lives.

[0003] However, due to the existence of the echo phenomenon, the accuracy of speech recognition decreases. As shown in the echo phenomenon Figure 1 In the enclosed space 1, such as a room, the speaker 21 of the speech recognition device 20 itself plays audio, and the audio signal is reflected by entities such as walls in the room and is received again by the microphone 22 of the speech recognition terminal through multiple paths. The echo signal received by the speech recognition module 23 drowns out the voice of the target user to be picked up by the microphone, resulting in the user's voice not being accurately recognized by the speech recognition terminal.

[0004] Therefore, it is urgent to develop a speech recognition device that is not affected by the echo phenomenon.

[0005] Application Content

[0006] Embodiments of this application provide a speech recognition method, apparatus, device, and storage medium, which enhance the robustness of the speech recognition model in an echo scenario, so that the speech recognition device with the speech recognition model is not affected by the echo phenomenon.

[0007] In a first aspect, the present application provides a speech recognition method, which includes first separately obtaining speech sample data of a source domain including source domain speech samples, text labels of the source domain speech samples, and domain labels of the source domain, and speech sample data of a target domain including target domain speech samples and domain labels of the target domain; wherein, the source domain speech samples do not include echo data, and the target domain speech samples include echo data; then extracting features of the source domain speech samples based on a feature extraction network to obtain first speech features, and extracting features of the target domain speech samples to obtain second speech features; inputting the first speech features as sample features into a speech recognition network to obtain a speech recognition result; inputting the first speech features and the second speech features as sample features into a domain classification network to obtain a domain classification result; jointly training the feature extraction network, the speech recognition network, and the domain classification network according to the speech recognition result and the text labels, and according to the domain classification result and the domain labels, to obtain a trained feature extraction network and a speech recognition network as a speech recognition model; and finally inputting the speech data to be recognized into the trained speech recognition model to obtain a speech recognition result.

[0008] The speech recognition method of the present application trains an acoustic model in an adversarial multi-task training manner. The network structure includes a feature extraction network, a speech recognition network, and a domain classification network. First, speech data in a non-echo scenario is collected as speech sample data of the source domain, and speech data in an echo scenario is collected as speech sample data of the target domain. The speech recognition network is trained using the speech sample data of the source domain, and at the same time, the feature extraction network and the domain classification network with an adversarial relationship are trained using the speech sample data of the source domain and the speech sample data of the target domain, so that the features extracted by the feature extraction network are domain-invariant and recognizable by the speech recognition network. Finally, a speech recognition model with robustness in an echo scenario is obtained.

[0009] In a possible implementation, the domain classification network includes a gradient reversal layer and a domain classification layer; the gradient reversal layer makes the feature extraction network and the domain classification layer form an adversarial relationship, and inputs the source domain speech samples and the domain labels of the source domain, and the target domain speech samples and the domain labels of the target domain into the feature extraction network respectively to train the feature extraction network and the domain classification layer.

[0010] In a possible implementation, the training of the feature extraction network and the domain classification layer includes: a forward propagation training process, where the speech features extracted by the feature extraction network pass through the gradient reversal layer and are input into the domain classification layer, and the domain classification layer updates the parameters of the domain classification layer according to the domain classification result and the domain labels; a backward propagation training process, where the gradient reversal layer takes the gradient of the domain classification layer and passes it to the feature extraction network to update the parameters of the feature extraction network.

[0011] In another possible implementation, the joint training of the feature extraction network, the speech recognition network, and the domain classification network includes: performing a weighted sum of the loss function of the speech recognition network and the loss function of the domain classification network to obtain a total loss function; and jointly training the feature extraction network, the speech recognition network, and the domain classification network by minimizing the total loss function.

[0012] In another possible implementation, the loss function of the domain classification network is a cross-entropy loss function or a KL divergence loss function; or the domain label includes a soft label, and the loss function of the domain classification network is a uniform distribution function obtained based on the soft label.

[0013] In another possible implementation, the obtaining of the speech sample data of the source domain and the speech sample data of the target domain includes: collecting speech data and determining whether the speech data includes far-end speech data; if so, using the speech data as the speech sample of the target domain and annotating the domain label of the target domain; if not, using the speech data as the speech sample of the source domain and annotating the text label and the domain label of the source domain.

[0014] In another possible implementation, the determining whether the speech data includes far-end speech data includes: determining whether the far-end speech energy in the speech data is greater than a first preset threshold; if so, it includes far-end speech data; if not, it does not include far-end speech data.

[0015] In another possible implementation, before extracting the features of the source domain speech sample based on the feature extraction network to obtain the first speech feature, it further includes: determining whether the speech sample data of the source domain and the speech sample data of the target domain are greater than or equal to a preset quantity, and the difference between the confidence of the speech recognition result output by the speech recognition network for the source domain speech sample and the confidence of the speech recognition result output by the speech recognition network for the target domain speech sample is greater than or equal to a second preset threshold; if so, then performing the extraction of the features of the source domain speech sample based on the feature extraction network to obtain the first speech feature.

[0016] In another possible implementation, the target domain includes multiple target domains determined based on the signal echo energy ratio.

[0017] In another possible implementation, the echo data at least includes one or more of the following: audio data processed by an echo cancellation module, far-end speech data, echo estimation data of the echo cancellation module, residual echo data of an echo suppression module, and echo path data estimated by the echo cancellation module.

[0018] In another possible implementation, the training process of the speech recognition model is completed in a cloud server, and the text label of the source domain speech sample is obtained based on manual annotation.

[0019] In another possible implementation, the training process of the speech recognition model is completed on the terminal device; the obtaining of the speech sample data of the source domain and the speech sample data of the target domain includes: the terminal device collects speech data, determines whether the speech data includes remote speech data, if so, uses the speech data as the speech sample of the target domain and labels the domain label of the target domain, if not, uses the speech data as the speech sample of the source domain and labels the domain label of the source domain; the text label of the source domain speech sample is obtained based on the acoustic model preset in the terminal device.

[0020] To achieve self-learning on the device side, on the one hand, there is no need to collect data offline to update the model and then update it to the user device. Instead, the device side collects data and can complete the iterative optimization of the model automatically for different levels of problems without the need for users to label data, without involving the offline red-labeling and cloud uploading of user data, ensuring that user personal data does not leave their own devices and protecting user privacy. On the other hand, for different device states and usage scenarios of different users, personalized optimization is carried out to improve the speech recognition performance of the user device in the echo scenario.

[0021] In another possible implementation, the speech recognition result is the speech content of the speech data to be recognized, or whether the speech data to be recognized contains a wake-up word, or whether the voiceprint feature of the speech data to be recognized matches a preset voiceprint feature.

[0022] In a second aspect, the present application also provides a speech recognition device, including: an acquisition module, configured to acquire speech sample data of a source domain and speech sample data of a target domain, where the speech sample data of the source domain includes a source domain speech sample, a text label of the source domain speech sample, and a domain label of the source domain, and the speech sample data of the target domain includes a target domain speech sample and a domain label of the target domain; wherein, a partial distribution feature of the source domain can be transferred and learned to the target domain, the source domain speech sample does not include echo data, and the target domain speech sample includes echo data; an extraction module, configured to extract the feature of the source domain speech sample based on a feature extraction network to obtain a first speech feature, and extract the feature of the target domain speech sample to obtain a second speech feature; a training module, configured to use the first speech feature as a sample feature to input into a speech recognition network to obtain a speech recognition result; use the first speech feature and the second speech feature as sample features to input into a domain classification network to obtain a domain classification result; and jointly train the feature extraction network, the speech recognition network, and the domain classification network according to the speech recognition result and the text label, and according to the domain classification result and the domain label, to obtain a trained feature extraction network and a speech recognition network as a speech recognition model; an identification module, configured to input the speech data to be recognized into the trained speech recognition model to obtain a speech recognition result.

[0023] In a possible implementation, the domain classification network includes a gradient reversal layer and a domain classification layer; the gradient reversal layer makes the feature extraction network and the domain classification layer form an adversarial relationship, and inputs the source domain speech samples and the domain labels of the source domain, and the target domain speech samples and the domain labels of the target domain into the feature extraction network respectively to train the feature extraction network and the domain classification layer.

[0024] In another possible implementation, training the feature extraction network and the domain classification layer includes: a forward propagation training process, where the speech features extracted by the feature extraction network are input into the domain classification layer through the gradient reversal layer, and the domain classification layer updates the parameters of the domain classification layer according to the domain classification result and the domain label; a backpropagation training process, where the gradient reversal layer reverses the gradient of the domain classification layer and passes it to the feature extraction network to update the parameters of the feature extraction network.

[0025] In another possible implementation, the training module is specifically configured to: perform weighted summation on the loss function of the speech recognition network and the loss function of the domain classification network to obtain a total loss function; and jointly train the feature extraction network, the speech recognition network, and the domain classification network by minimizing the total loss function.

[0026] In another possible implementation, the loss function of the domain classification network is a cross-entropy loss function or a KL divergence loss function; or the domain label includes a soft label, and the loss function of the domain classification network is a uniform distribution function obtained based on the soft label.

[0027] In another possible implementation, the acquisition module is specifically configured to: collect speech data, and determine whether the speech data includes remote speech data; if so, use the speech data as target domain speech samples and label the domain labels of the target domain; if not, use the speech data as source domain speech samples and label the text label and the domain labels of the source domain.

[0028] In another possible implementation, determining whether the speech data includes remote speech data includes: determining whether the remote speech energy in the speech data is greater than a first preset threshold; if so, it includes remote speech data; if not, it does not include remote speech data.

[0029] In another possible implementation, the device further includes: a judgment module, configured to judge whether the speech sample data of the source domain and the speech sample data of the target domain are greater than or equal to a preset quantity, and whether the difference between the confidence of the speech recognition result output by the speech recognition network for the source domain speech sample and the confidence of the speech recognition result output by the speech recognition network for the target domain speech sample is greater than or equal to a second preset threshold; if so, then execute extracting the features of the source domain speech sample by the feature extraction network to obtain the first speech features.

[0030] In another possible implementation, the target domain includes multiple target domains determined based on the signal echo energy ratio.

[0031] In another possible implementation, the echo data at least includes one or more of the following: audio data processed by an echo cancellation module, far-end speech data, echo estimation data of the echo cancellation module, residual echo data of the echo suppression module, and echo path data estimated by the echo cancellation module.

[0032] In another possible implementation, the training process of the speech recognition model is completed in a cloud server, and the text labels of the source domain speech samples are obtained based on manual annotation.

[0033] In another possible implementation, the training process of the speech recognition model is completed on a terminal device; obtaining the speech sample data of the source domain and the speech sample data of the target domain includes: the terminal device collecting speech data, judging whether the speech data includes far-end speech data, if so, then using the speech data as the target domain speech sample and labeling the domain label of the target domain, if not, then using the speech data as the source domain speech sample and labeling the domain label of the source domain; the text labels of the source domain speech samples are obtained based on the recognition by an acoustic model preset in the terminal device.

[0034] In another possible implementation, the speech recognition result is the speech content of the speech data to be recognized, or whether the speech data to be recognized contains a wake-up word, or whether the voiceprint feature of the speech data to be recognized matches a preset voiceprint feature.

[0035] In a third aspect, the present application further provides a speech recognition device, including a memory and a processor, where an executable code is stored in the memory, and the processor executes the executable code to implement the method described in the first aspect or any possible implementation manner of the first aspect.

[0036] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in the first aspect or any possible implementation manner of the first aspect. Description of the Drawings

[0037] Figure 1 It is a schematic diagram of the echo phenomenon;

[0038] Figure 2a It is a histogram of the word error rate of speech recognition varying with the echo energy ratio in an echo scenario;

[0039] Figure 2b It is a schematic structural diagram of the speech recognition device in the first solution;

[0040] Figure 2c It is a schematic diagram of the principle of echo cancellation in the second solution;

[0041] Figure 2d It is a schematic diagram of the principle of residual echo estimation in the third solution;

[0042] Figure 2e It is a schematic block diagram of the speech recognition device in the third solution;

[0043] Figure 3 It is a schematic structural diagram of the speech recognition model during the training process provided by the embodiments of the present application;

[0044] Figure 4 It is a flowchart of the speech recognition method provided by the embodiments of the present application;

[0045] Figure 5 It is a schematic structural diagram of the domain classification network in the embodiments of the present application;

[0046] Figure 6 It is a schematic diagram of the training process of the speech recognition model in the cloud server during the training process provided by the embodiments of the present application;

[0047] Figure 7 It is a schematic diagram of the training process of the speech recognition model in the speech recognition device during the training process provided by the embodiments of the present application;

[0048] Figure 8 It is an architecture diagram of the speech recognition system provided by the embodiments of the present application;

[0049] Figure 9 It is an architecture diagram of the voice wake-up system provided by the embodiments of the present application;

[0050] Figure 10 It is a schematic structural diagram of the speech recognition device provided by the embodiments of the present application;

[0051] Figure 11 It is a schematic structural diagram of the speech recognition device provided by the embodiments of the present application. Detailed Description of the Invention

[0052] The technical solution of the present application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0053] The voice recognition interaction scenarios in the echo scenario include: wake-up-free interruption, wake-up interruption, specific person wake-up, multi-language voice recognition, etc.

[0054] For example, when a voice robot or smart speaker is playing music or telling a story, the user hopes that the device can recognize voice commands such as "stop talking" and "next song" to control the device, achieving wake-up-free interruption.

[0055] For example, when a voice robot or smart speaker is playing music or telling a story, the user hopes to interrupt the playback by using a wake-up word, such as "wake-up word one", achieving wake-up interruption.

[0056] For example, when a voice robot or smart speaker is playing music or telling a story, only a specific person can wake up and interrupt, achieving specific person wake-up.

[0057] For example, for a voice robot or smart speaker playing background music, whether the user speaks Chinese or other languages (such as English), the speech content can be accurately recognized, achieving multi-language voice recognition.

[0058] However, the echo phenomenon seriously affects the speech recognition accuracy of speech recognition devices. As Figure 2a shown, as the echo energy ratio of the speech data decreases (the smaller the echo energy ratio, the stronger the echo), the word error rate (WER) of speech recognition increases significantly.

[0059] Therefore, to achieve the speech recognition function in the above echo scenario, it is necessary to solve the problem that the echo phenomenon affects the speech recognition accuracy.

[0060] The first solution is to add an echo cancellation module. As Figure 2b shown, an echo cancellation module 24 is added to the speech recognition device 2, so that the audio data picked up by the microphone first passes through the echo cancellation module 24 for processing, and then enters the speech recognition module 23 for speech recognition.

[0061] Although this solution solves the echo problem in most speech recognition interaction scenarios through the echo cancellation module, for the intelligent voice interaction scenario, due to challenges such as the non-linearity of low-cost speakers, strong echo levels caused by far-field voice interaction, and inter-channel correlation in multi-channels, problems such as proximal speech distortion or significant residuals in the distal speech will significantly reduce the usability of the speech recognition module. In addition, as the usage time of the user device increases, different user device speakers have different degrees of non-linearity problems, and it is difficult to have a unified offline algorithm to optimize the performance of relevant speech algorithms for different non-linearity degrees.

[0062] The second solution is an echo cancellation solution based on signal processing. By linearly modeling the echo path and using adaptive filtering to estimate the echo path, the user's voice after echo control can be obtained. For the residual echo, the Wiener filtering method is used to further suppress the echo residue.

[0063] Specifically, as Figure 2c shown, on the basis of the echo cancellation module 23, a residual echo suppression module 25 is added. First, the microphone signal containing the echo signal is modeled as:

[0064] d = s + y.

[0065] The microphone signal after the echo cancellation algorithm is processed is expressed as:

[0066]

[0067] The residual echo suppression module using Wiener filtering is designed as:

[0068]

[0069] The residual echo is estimated as:

[0070]

[0071] Although this solution effectively suppresses the echo in some scenarios, its processing effect is poor for non-linear distortion and strong echo scenarios. For voice interaction products such as smart speakers, robots, and mobile phones, there is significant non-linear distortion in the speaker, and the strong echo caused by the close distance between the speaker and the microphone is a high-frequency scenario. Eventually, this processing solution is difficult to meet the interaction requirements in voice interaction products.

[0072] In addition, this solution will also bring serious non-linear distortion, and voice interaction functions, including speech recognition, wake-up, speaker recognition, and language recognition, are very sensitive to non-linear distortion, affecting the performance of voice interaction.

[0073] The third solution starts from the perspective of deep learning. As Figure 2d shown, the residual echo is learned through a neural network. The input of the network is the output of the echo cancellation module, and the output of the network is the estimation of the residual echo.

[0074] The method using a neural network can utilize the powerful learning ability of the neural network to effectively suppress the residual echo.

[0075] However, on the one hand, neural network learning requires a large amount of paired labeled data, and the acquisition cost is high. And since the residual echo output by the echo cancellation module is weaker than the proximal user's voice under normal circumstances, it is difficult to effectively estimate the residual echo. On the other hand, the learning objective of this solution is not related to speech recognition and cannot guarantee the accuracy of speech recognition.

[0076] As Figure 2e shown, the fourth solution is to directly provide information such as the original microphone signal, residual echo signal, echo estimation signal, and reference signal to the speech recognition model. The speech recognition model is trained with labeled training data to improve the speech recognition effect in an echo scenario.

[0077] On the one hand, this solution can improve the processing effect of the echo scenario by using a neural network. On the other hand, training the speech recognition model with multi-condition data for speech recognition can avoid the problem that the speech recognition effect cannot be improved in the first three solutions. However, since a large amount of labeled data is required for training the speech recognition model, a relatively large training cost is needed. In addition, the multi-condition training method may not guarantee the robustness of the speech recognition model to residual echo, and the improvement of the speech recognition effect in the echo scenario is limited.

[0078] Figure 3 is a schematic structural diagram of the speech recognition model during the training process provided by an embodiment of the present application. As Figure 3 described, it includes a first branch composed of a feature extraction network 30 and a speech recognition network 31, and a second branch composed of a feature extraction network 30 and a domain discriminant network 32.

[0079] First, obtain the speech data in a non-echo scenario as the speech sample data of the source domain, and obtain the speech data in an echo scenario as the speech sample data of the target domain.

[0080] Then, input the speech sample data of the source domain into the first branch to train the speech recognition network and the feature extraction network, and adjust the parameters of the feature extraction network and the speech recognition network. Input the speech sample data of the source domain and the speech sample data of the target domain into the second branch to train the domain discriminant network and the feature extraction network, and adjust the parameters of the discriminant network and the feature extraction network to make the features extracted by the feature extraction network be domain-invariant and recognizable by the speech recognition network.

[0081] Finally, use the trained feature extraction network and speech recognition network as the speech recognition model to perform speech recognition on the speech data to be recognized, and obtain a speech recognition model with better robustness in a non-echo scenario.

[0082] It can be understood that the "model" mentioned in this article can learn the corresponding association between the input and the output from the training data, so that after training, for a given input, the corresponding output can be generated. The "model" can also be referred to as a "neural network", a "learning model", or a "learning network", etc.

[0083] Figure 4 is a flowchart of the speech recognition method provided by an embodiment of the present application. As Figure 4 shown, the speech recognition method includes the following steps:

[0084] S401. Obtain the speech sample data of the source domain and the speech sample data of the target domain.

[0085] Among them, the speech sample data of the source domain includes the source domain speech sample, the text label of the source domain speech sample, and the domain label of the source domain. The speech sample data of the target domain includes the target domain speech sample and the domain label of the target domain. The source domain speech sample does not include echo data, and the target domain speech sample includes echo data. For example, the speech data collected in a non-echo scenario can be used as the source domain speech sample, and the speech data collected in a non-echo scenario can be used as the target domain speech sample. Two types of modeling methods for the source domain and the target domain are proposed in this application, which is convenient for obtaining training data.

[0086] It should be explained that the echo scenario here is a scenario in a closed or partially closed space (such as a room), where the audio signal played by the speaker of the speech recognition device hits an obstacle and then is received by the pickup component of the speech recognition device through various paths, resulting in an echo phenomenon. The non-echo scenario is a scenario where the speaker of the speech recognition device does not play or plays in a non-closed space without an echo phenomenon.

[0087] Obtaining the speech sample of the source domain and the speech sample of the target domain can be achieved by collecting through the pickup component configured in the speech recognition device. For example, the pickup component is a microphone or an array microphone, that is, this application does not limit whether the speech sample is obtained through a single channel or multiple channels. It can also be obtained through data transmission.

[0088] The speech recognition device can be a device configured with a pickup component such as a smart speaker, a smart phone, a tablet computer, a notebook computer, a handheld computer, a personal digital assistant, a smart wearable device, etc.

[0089] Since the speech sample of the source domain does not have echo data and there is no problem that the echo affects the recognition accuracy of the speech recognition model, it can be automatically labeled after being recognized by an existing speech recognition model. Therefore, the text label of the speech sample of the source domain is relatively easy to obtain, and the first branch can be trained in a supervised training manner. However, the speech sample of the target domain includes echo data, and there is a problem that the echo affects the recognition accuracy of the speech recognition model and cannot be automatically labeled by an existing speech recognition model. All need to be manually labeled, and the workload is huge. Therefore, the second branch is trained in an unsupervised training manner.

[0090] For example, the annotation of the text label of the source domain speech sample can be automatically annotated after being recognized by the acoustic model preset in the speech recognition device, or can be automatically annotated after being recognized by any device with speech recognition function, or the training data that has been manually annotated by other acoustic models. This application does not make a limitation.

[0091] In one example, the voice samples of the source domain and the voice samples of the target domain are input into the feature extraction network in batches, and the domain labels of the source domain and the domain labels of the target domain are automatically labeled during input. Alternatively, the domain labels are automatically labeled by determining whether there is echo data during input, that is, the domain labels of the target domain are automatically labeled for those with echo data, and the domain labels of the source domain are automatically labeled for those without echo data.

[0092] S402. Extract the features of the source domain voice samples based on the feature extraction network to obtain the first voice feature, and extract the features of the target domain voice samples to obtain the second voice feature.

[0093] The first voice feature and the second voice feature are acoustic features, generally the spectral features of voice data. For example, Mel Frequency Cepstrum Coefficient (MFCC) features, FilterBank (FBank) features, etc.

[0094] Taking the first voice feature and the second voice feature extracted by the feature extraction network as FBank features as an example for illustration. The FBank feature is a front-end processing algorithm that processes audio in a way similar to the human ear and can improve the performance of speech recognition. The extraction process mainly includes Fourier transform, energy spectrum calculation, Mel filtering, and taking the Log value. For the sake of simplicity, it will not be elaborated here in detail. For specific details, please refer to the FBank feature extraction process in the prior art.

[0095] Generally speaking, the purpose of extracting FBank features is to reduce the dimension of audio data. For example, for an audio file with a length of one second, if the sampling rate is 16,000 sampling points per second, the number of bits after converting this audio file into an array will be very long. By extracting FBank features, the length of one frame of the audio file can be reduced to 80 bits.

[0096] The feature extraction network can specifically be a deep neural network such as a Convolutional Neural Networks (CNN) or a Recurrent Neural Networks (RNN).

[0097] S403. Input the first voice feature as the sample feature into the speech recognition network to obtain the speech recognition result; input the first voice feature and the second voice feature as the sample feature into the domain classification network to obtain the domain classification result.

[0098] In one example, the speech recognition network and the domain classification network can be neural networks such as Long Short-Term Memory (LSTM), Deep Neural Networks (DNN), and Convolutional Neural Networks (CNN).

[0099] For example, the speech recognition network is a DNN with 6 hidden layers, 1936 output nodes, and 1024 hidden layer nodes, and the domain classification network is a fully connected network with one hidden layer.

[0100] The speech recognition network can be an existing speech recognition network trained in a non-echo scenario or an untrained speech recognition network, which is not limited in this application. The speech recognition network outputs a speech recognition result based on the first speech feature, and the domain classification network outputs a domain classification result based on the first speech feature and the second speech feature.

[0101] S404. According to the speech recognition result and the text label, and according to the domain classification result and the domain label, jointly train the feature extraction network, the speech recognition network, and the domain classification network to obtain the trained feature extraction network and speech recognition network as the speech recognition model.

[0102] In one example, as Figure 5 shown, the domain classification network 32 includes a gradient reversal layer 321 and a domain classification layer 322. During the training process, it includes forward propagation, that is, the speech features extracted by the feature extraction network are input into the domain classification layer through the gradient reversal layer, and the domain classification layer outputs the domain classification result. The signal propagation direction is from the feature extraction network to the domain classification layer; backward propagation, that is, the error between the domain classification result output by the domain classification layer and the domain label is returned to the feature extraction network. The signal propagation direction is from the domain classification layer to the feature extraction network.

[0103] During the forward propagation process, the gradient reversal layer does not process the first speech feature and / or the second speech feature and directly passes them to the domain classification layer. During the backward propagation process, the gradient reversal layer takes the opposite of the gradient of the classification layer and passes it to the feature extraction network. This makes the update direction of the feature extraction network opposite to that of the domain classification network. That is, the training objective of the domain classification network is to accurately identify the domain to which the speech sample belongs as much as possible, while the training objective of the feature extraction network is that the extracted speech features make the domain classification network unable to identify the domain to which the speech sample belongs as much as possible. Therefore, through the gradient reversal layer, an adversarial relationship is formed between the feature extraction network and the domain classification network. Finally, the feature extraction network is trained to extract domain-invariant features, enabling the speech recognition model to still have good robustness in the echo scenario.

[0104] Meanwhile, the speech recognition network in the first branch trains the feature extraction network to extract speech features recognizable by the speech recognition network based on the speech recognition result and the text label of the speech sample in the source domain.

[0105] Of course, it also includes aligning the speech samples in the source domain with the text labels of the speech samples in the source domain to facilitate the training in the first branch. Based on the speech recognition result and the text label of the corresponding speech sample in the source domain, the feature extraction network is trained to extract speech features recognizable by the speech recognition network.

[0106] Generally speaking, according to the speech recognition result and the text label, and according to the domain classification result and the domain label, the feature extraction network, the speech recognition network, and the domain classification network are jointly trained, and finally the trained feature extraction network and the speech recognition network are obtained as the speech recognition model.

[0107] Specifically, the loss functions of the speech recognition network and the domain classification network are weighted and summed to obtain the total loss function. By minimizing the total loss function, the feature extraction network, the speech recognition network, and the domain classification network are jointly trained. The total cost function is obtained by summing the total loss functions of multiple trainings.

[0108] The cost function is:

[0109]

[0110] Among them, E(θ f , θ y , θ d ) represents the cost function, θ f represents the parameters of the feature extraction network, θ y represents the parameters of the speech recognition network, θ d represents the parameters of the domain classification network, N represents the number of samples in each training data block, L y represents the loss function of the speech recognition network, L d represents the loss function of the domain classification network. The overall goal is to minimize the total cost function, that is, to minimize the loss of the speech recognition network while maximizing the loss of the domain classification network, so as to maximize the speech recognition of the features of the feature extraction network while making the domain classification network unable to distinguish. Optimizing the three parameters separately can be decomposed into

[0111]

[0112]

[0113]

[0114] Further calculate the gradient of the network and backpropagate, and the obtained calculation formula is

[0115]

[0116]

[0117]

[0118] where α is the learning rate, and the negative sign of the gradient in the first sub-formula represents the operation of the gradient reversal layer.

[0119] In one example, the loss function of the domain classification network and the loss function of the speech recognition network can be the cross-entropy loss function, or the KL divergence loss function, or the domain label is a soft label, and the loss function of the domain classification network is a uniform distribution function obtained based on the soft label, etc. The present application does not limit the loss functions of the domain classification network and the speech recognition network, as long as the purpose of training the above-mentioned speech recognition network and domain classification network can be achieved.

[0120] Finally, in step S405, the speech data to be recognized is input into the trained speech recognition model to obtain a speech recognition result.

[0121] In some embodiments, the training process of the above speech recognition model is performed on a cloud server. As Figure 6 shown, the uploaded speech data is received, and it is determined whether there is a far-end speech in the speech data (the far-end speech here is the speech played by the speaker of the speech recognition device). If so, it is used as the speech sample of the source domain, the domain label of the source domain is marked, and the text label of the domain label of the source domain is retrieved to form the speech sample data of the source domain. If not, the speech data, and / or the speech data obtained by subjecting the speech data to echo cancellation processing, and / or the speech data obtained by subjecting the speech data obtained by echo cancellation processing to residual echo suppression processing is used as the speech sample of the target domain, and the domain label of the target domain is marked to form the speech sample data of the target domain. Then, the speech sample data of the source domain and the speech sample data of the target domain are input into the training network in batches to train the feature extraction network, the domain classification network, and the speech recognition network. The trained feature extraction network and speech recognition network are used as the speech recognition model to recognize speech data in an echo scenario or a non-echo scenario. The present application can obtain far-end speech and algorithm-estimated echo, and the characteristics of these data can be additionally utilized during the training of the speech recognition model to further improve the recognition accuracy under residual distortion.

[0122] In addition, the speech recognition method of the present application only models the residual distortion of echo cancellation and does not limit to a specific algorithm, and the echo cancellation algorithm is a linear process. Therefore, the speech recognition model trained by the present application has good generalization ability for echo cancellation algorithms.

[0123] In some other embodiments, the training process of the above speech recognition model can be carried out on a speech recognition device. As Figure 7 shown, the speech recognition device collects speech data through a sound pickup component, and determines whether there is distal speech in the speech data (the distal speech here is the speech played by the speaker of the speech recognition device). If so, it is used as the speech sample in the source domain. If not, the speech data, and / or the speech data obtained by subjecting the speech data to echo cancellation processing, and / or the speech data obtained by subjecting the speech data obtained by echo cancellation processing to residual echo suppression processing is used as the speech sample in the target domain. It is determined whether the speech recognition device needs end-side self-learning (that is, whether the speech recognition model needs to be trained). If so, the preset acoustic model is used to recognize the speech sample in the source domain to obtain its corresponding text label and automatically annotated domain label, which together with the speech sample in the source domain form the speech sample data in the source domain. The automatically annotated domain label of the speech sample in the target domain is combined with the speech sample in the target domain to form the speech sample data in the target domain. Then, the speech sample data in the source domain and the speech sample data in the target domain are input into the training network in batches to train the feature extraction network, domain classification network, and speech recognition network. The trained feature extraction network and speech recognition network are used as the speech recognition model to recognize speech data in an echo scenario or a non-echo scenario.

[0124] End-side self-learning is realized. On the one hand, there is no need to collect data offline to update the model and then update it to the speech recognition device. Instead, according to the design principle, data is collected on the end side, and there is no need for users to annotate data. It can automatically complete algorithm iteration optimization for different degrees of problems. For different device states of different users, personalized optimization can be carried out to improve the robustness of speech recognition of the device in an echo scenario. On the other hand, since it does not involve offline labeling and uploading of user data to the cloud, the training process can be automatically completed regularly on the user device, ensuring that the user's personal data does not leave their own device and ensuring the user's privacy requirements.

[0125] It can be understood that there are various ways to determine whether there is distal speech in the speech data. For example, in the energy decision method, the distal speech is a pure echo signal, and as long as there is a signal, it needs to be processed. Therefore, it only needs to distinguish between silence and non-silence. The optional criterion is Pr>delta, and it can be considered that there is a distal speech signal. Pr is the short-time estimate of the energy of the distal speech. A typical choice is Pr (t) =(1 - alpha)*Pr (t-1) +alpha*|x (t) | ^2 , Pr (t) is the energy estimate at time t, Pr (t-1) is the estimate at time t - 1, alpha is a constant, and a typical value is 0.95. x (t)is the instantaneous amplitude of the remote voice at time t, and delta is the threshold value (i.e., the first preset threshold), for example, it can take a value of 1e -5 。

[0126] In one example, the method for determining whether the voice recognition device needs end-side self-learning is as follows: Determine whether the voice sample data in the source domain and the voice sample data in the target domain are greater than or equal to a preset quantity, and the difference between the confidence of the voice recognition result output by the voice recognition network for the source domain voice sample and the confidence of the voice recognition result output by the voice recognition network for the target domain voice sample is greater than or equal to the second preset threshold. If so, it is determined that the raincoat recognition device needs to perform end-side self-learning.

[0127] For example, every fixed time, such as one week, and collect sufficient training samples, such as collecting 1000 pieces (i.e., the preset quantity) respectively for the echo scenario and the non-echo scenario, about 1 hour of data, and there is a significant difference in the performance of the source domain and target domain data on the existing recognition system. The index performance is the confidence of the voice recognition system. When there is a significant difference in the confidence of the source domain and target domain data, |C s -C t |≥delta, where C s is the average confidence of the source domain data, C t is the average confidence of the target domain data, delta is the threshold value (i.e., the second preset threshold), for example, delta≥0.3, and the number of samples of C s and C t is at least 100. The confidence is the standard statistic of the existing voice recognition system and is an index used to measure the voice recognition quality.

[0128] In some examples, the above-mentioned target domain is multiple target domains determined based on the signal echo energy ratio. For example, target domain one, echo energy ratio -20dB; target domain two, echo energy ratio -15dB; target domain three, echo energy ratio -10dB; target domain four, echo energy ratio -5dB; target domain five, echo energy 0; target domain six, echo energy ratio 5dB; target domain seven, echo energy ratio 10dB, etc. Then the corresponding domain classification network is a multi-domain classification task. On the one hand, the target domain performs multi-domain modeling according to the echo residual level. Compared with the binary classification situation of two-domain modeling, it further improves the invariance of the features extracted by the feature extraction network, further improves the voice recognition performance, and expands the applicable range of this method; on the other hand, since this application does not distinguish between modeling the echo cancellation residual and the echo suppression module, the target domain can be designed as the output of the echo suppression module. Therefore, the echo control module can be flexibly designed.

[0129] In addition, since the present application does not utilize channel information, the speech recognition method provided by the present application can be applied not only to single-channel echo cancellation algorithms, but also to stereo and multi-channel echo cancellation algorithms, etc., for the additional echo cancellation distortion problems caused by channel correlation.

[0130] The echo data can further include one or more of the audio data processed by the acoustic cancellation module, the echo estimation data of the echo cancellation module, the residual echo data of the echo suppression module, and the echo path data estimated by the echo cancellation module, further improving the performance of the algorithm, adapting to more scenarios, and increasing the robustness of the trained speech recognition model.

[0131] In some embodiments, the above-trained speech recognition model can be applied to a speech recognition system, and the speech recognition result is output as the speech content (i.e., the corresponding text content) of the speech data to be recognized. For example Figure 8 As shown, the speech recognition system includes a front-end processing module, a decoder (including a speech recognition model, a pronunciation dictionary, and a language model), and a text output module. The speech to be recognized (the speech to be recognized in the present application can be either far-field speech or near-field speech) is input, preprocessed by the front-end processing module such as noise reduction and feature extraction, and the feature data is transmitted to the decoder. The decoder calculates the text sequence with the best feature match according to resources such as the trained speech recognition model, pronunciation dictionary, and language model, and outputs it in the text output module.

[0132] In other embodiments, the above-trained speech recognition model can be applied to speech recognition-related systems such as a voice wake-up system and a voiceprint recognition system. The following takes the application of the speech recognition model to a voice wake-up system as an example to illustrate this application solution.

[0133] Figure 9 This is the architecture diagram of the voice wake-up system provided in this embodiment. For example Figure 9 As shown, the voice wake-up system includes a front-end processing module, a speech recognition model, a post-processing module, and a wake-up result module. The voice input module converts the acoustic signal of the input voice into an electrical signal, and then transmits the electrical signal to the front-end processing module. The front-end processing module preprocesses the electrical signal, such as echo cancellation, noise reduction, etc., and transmits the preprocessed data to the speech recognition model. The speech recognition model can adapt to complex acoustic environments (such as echo environments) and various data processed by different front-end processing modules (such as echo cancellation), and converts the preprocessed data into a modeling unit module defined by wake-up labels, and then transmits it to the post-processing module to convert the output probability of the speech recognition model into a wake-up confidence score. Finally, the wake-up result module outputs the result of whether to wake up according to the confidence score, and controls whether to wake up the device according to the wake-up result.

[0134] The embodiments of this application also provide a voice recognition device, such as Figure 10 shown. The voice recognition device 100 at least includes:

[0135] An acquisition module 101, configured to acquire voice sample data of a source domain and voice sample data of a target domain. The voice sample data of the source domain includes a source domain voice sample, a text label of the source domain voice sample, and a domain label of the source domain. The voice sample data of the target domain includes a target domain voice sample and a domain label of the target domain. Wherein, part of the distribution characteristics of the source domain can be transferred and learned to the target domain. The source domain voice sample does not include echo data, and the target domain voice sample includes echo data;

[0136] An extraction module 103, configured to extract features of the source domain voice sample based on a feature extraction network to obtain a first voice feature, and extract features of the target domain voice sample to obtain a second voice feature;

[0137] A training module 104, configured to use the first voice feature as sample features to input into a voice recognition network to obtain a voice recognition result; use the first voice feature and the second voice feature as sample features to input into a domain classification network to obtain a domain classification result; according to the voice recognition result and the text label, according to the domain classification result and the domain label, jointly train the feature extraction network, the voice recognition network, and the domain classification network to obtain a trained feature extraction network and a voice recognition network as a voice recognition model;

[0138] A recognition module 105, configured to input the voice data to be recognized into the trained voice recognition model to obtain a voice recognition result.

[0139] In a possible implementation, the domain classification network includes a gradient reversal layer and a domain classification layer; the gradient reversal layer makes the feature extraction network and the domain classification layer form an adversarial relationship, and inputs the source domain voice sample and the domain label of the source domain, and the target domain voice sample and the domain label of the target domain into the feature extraction network respectively to train the feature extraction network and the domain classification layer.

[0140] In another possible implementation, the training of the feature extraction network and the domain classification layer includes: a forward propagation training process, where the voice features extracted by the feature extraction network pass through the gradient reversal layer and are input into the domain classification layer, and the domain classification layer updates the parameters of the domain classification layer according to the domain classification result and the domain label; a backward propagation training process, where the gradient reversal layer takes the gradient of the domain classification layer and reverses it and then passes it to the feature extraction network to update the parameters of the feature extraction network.

[0141] In another possible implementation, the training module 104 is specifically configured to: perform weighted summation on the loss function of the speech recognition network and the loss function of the domain classification network to obtain a total loss function; and jointly train the feature extraction network, the speech recognition network, and the domain classification network by minimizing the total loss function.

[0142] In another possible implementation, the loss function of the domain classification network is a cross-entropy loss function or a KL divergence loss function; or the domain label includes a soft label, and the loss function of the domain classification network is a uniform distribution function obtained based on the soft label.

[0143] In another possible implementation, the obtaining module 101 is specifically configured to: collect speech data, and determine whether the speech data includes remote speech data; if so, use the speech data as a target domain speech sample and label the domain label of the target domain; if not, use the speech data as a source domain speech sample and label the text label and the domain label of the source domain.

[0144] In another possible implementation, determining whether the speech data includes remote speech data includes: determining whether the remote speech energy in the speech data is greater than a first preset threshold; if so, it includes remote speech data; if not, it does not include remote speech data.

[0145] In another possible implementation, the apparatus further includes: a judgment module 102, configured to judge whether the speech sample data of the source domain and the speech sample data of the target domain are greater than or equal to a preset quantity, and whether the difference between the confidence of the speech recognition result output by the speech recognition network for the source domain speech sample and the confidence of the speech recognition result output by the speech recognition network for the target domain speech sample is greater than or equal to a second preset threshold; if so, execute extracting the features of the source domain speech sample based on the feature extraction network to obtain first speech features.

[0146] In another possible implementation, the target domain includes multiple target domains determined based on the signal echo energy ratio.

[0147] In another possible implementation, the echo data at least includes one or more of: audio data processed by an echo cancellation module, remote speech data, echo estimation data of the echo cancellation module, residual echo data of the echo suppression module, and echo path data estimated by the echo cancellation module.

[0148] In another possible implementation, the training process of the speech recognition model is completed in a cloud server, and the text label of the source domain speech sample is obtained based on manual annotation.

[0149] In another possible implementation, the training process of the speech recognition model is completed on a terminal device; the obtaining of the speech sample data of the source domain and the speech sample data of the target domain includes: the terminal device collects speech data, determines whether the speech data includes remote speech data, if so, the speech data is used as the speech sample of the target domain and the domain label of the target domain is labeled, if not, the speech data is used as the speech sample of the source domain and the domain label of the source domain is labeled; the text label of the speech sample of the source domain is obtained based on the acoustic model preset in the terminal device.

[0150] In another possible implementation, the speech recognition result is the speech content of the speech data to be recognized, or whether the speech data to be recognized contains a wake-up word, or whether the voiceprint feature of the speech data to be recognized matches a preset voiceprint feature.

[0151] The speech recognition device 100 according to the embodiment of the present application can correspondingly execute the method described in the embodiment of the present application, and the above and other operations and / or functions of each module in the speech recognition device 100 are respectively for implementing Figures 3 - 9 the corresponding processes of each method in, for the sake of brevity, will not be described in detail here.

[0152] In addition, it should be noted that the above-described embodiments are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiment provided by the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0153] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute any one of the above methods.

[0154] The present application also provides a computer program or a computer program product. The computer program or the computer program product includes instructions. When the instructions are executed, the computer is made to execute any one of the above methods.

[0155] The present application also provides a speech recognition device, including a memory and a processor. An executable code is stored in the memory, and the processor executes the executable code to implement any one of the above methods.

[0156] Figure 11 is a schematic structural diagram of the speech recognition device provided by the present application.

[0157] As shown in FIG. 11, the speech recognition device 1100 includes a processor 1101, a memory 1102, a bus 1103, a microphone 1104, a speaker 1105, and a communication interface 1106. Among them, the processor 1101, the memory 1102, the microphone 1104, the speaker 1105, and the communication interface 1106 communicate through the bus 1103, and can also communicate through other means such as wireless transmission. The microphone 1104 can receive and pick up voice data, the speaker 1105 can play voice data, the communication interface is used to communicate and connect with other communication devices, the memory 1102 stores executable program codes, and the processor 1101 can call the program codes stored in the memory 1102 to execute the speech recognition method in the foregoing method embodiments.

[0158] It should be understood that in the embodiments of the present application, the processor 1101 may be a central processing unit CPU, and the processor 1101 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0159] The memory 1102 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1101. The memory 1102 may also include a non-volatile random access memory. For example, the memory 1102 may also store a training data set.

[0160] The memory 1102 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0161] In addition to including a data bus, the bus 1103 can also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clear illustration, all kinds of buses are labeled as bus 1103 in the figure.

[0162] It should be understood that the speech recognition device 100 according to the embodiments of the present application can correspond to the speech recognition apparatus in the embodiments of the present application, and can correspond to the corresponding main body that executes the method shown in the embodiments of the present application Figures 3 - 9 and the above and other operations and / or functions of each device in the speech recognition device 1100 respectively implement the corresponding processes of the respective methods shown, and for the sake of brevity, they will not be described herein again. Figures 3 - 9

[0163] ​Those of ordinary skill in the art should further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0164] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the technical field.

[0165] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application and are not used to limit the protection scope of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included in the protection scope of this application.

Claims

1. A voice recognition method, characterized in that, Including: Obtain the speech sample data of the source domain and the speech sample data of the target domain. The speech sample data of the source domain includes the source domain speech sample, the text label of the source domain speech sample, and the domain label of the source domain. The speech sample data of the target domain includes the target domain speech sample and the domain label of the target domain. Among them, part of the distribution characteristics of the source domain can be transferred and learned to the target domain. The source domain speech sample does not include echo data, and the target domain speech sample includes echo data. The echo data includes at least one or more of the following: audio data processed by an echo cancellation module, far-end speech data, echo estimation data of the echo cancellation module, residual echo data of the echo suppression module, and echo path data estimated by the echo cancellation module. Extract the features of the source domain speech sample based on the feature extraction network to obtain the first speech feature, and extract the features of the target domain speech sample based on the feature extraction network to obtain the second speech feature. Input the first speech feature as the sample feature into the speech recognition network to obtain the speech recognition result; input the first speech feature and the second speech feature as the sample feature into the domain classification network to obtain the domain classification result. According to the speech recognition result and the text label, and according to the domain classification result and the domain label, jointly train the feature extraction network, the speech recognition network, and the domain classification network to obtain the trained feature extraction network and speech recognition network as the speech recognition model. Input the speech data to be recognized into the trained speech recognition model to obtain the speech recognition result.

2. The method according to claim 1, wherein The domain classification network includes a gradient reversal layer and a domain classification layer. The gradient reversal layer makes the feature extraction network and the domain classification layer form an adversarial relationship. Input the source domain speech sample and the domain label of the source domain, and the target domain speech sample and the domain label of the target domain into the feature extraction network respectively to train the feature extraction network and the domain classification layer.

3. The method according to claim 2, wherein The training of the feature extraction network and the domain classification layer includes: In the forward propagation training process, the speech features extracted by the feature extraction network are input into the domain classification layer through the gradient reversal layer, and the domain classification layer updates the parameters of the domain classification layer according to the domain classification result and the domain label. In the backward propagation training process, the gradient reversal layer takes the gradient of the domain classification layer and passes it to the feature extraction network to update the parameters of the feature extraction network.

4. The method according to claim 1, characterized in that, The joint training of the feature extraction network, the speech recognition network, and the domain classification network includes: Perform a weighted sum of the loss function of the speech recognition network and the loss function of the domain classification network to obtain the total loss function. By minimizing the total loss function, jointly train the feature extraction network, the speech recognition network, and the domain classification network.

5. The method according to claim 4, wherein The loss function of the domain classification network is the cross-entropy loss function or the KL divergence loss function. Or the domain label includes soft labels, and the loss function of the domain classification network is a uniform distribution function obtained based on the soft labels.

6. The method according to claim 1, characterized in that, The obtaining of the speech sample data of the source domain and the speech sample data of the target domain includes: Collect speech data and determine whether the speech data includes far-end speech data. If so, use the voice data as a target domain voice sample and label the domain label of the target domain; If not, use the voice data as a source domain voice sample and label the text label and the domain label of the source domain.

7. The method according to claim 6, wherein The determination of whether the voice data includes remote voice data includes: Determine whether the remote voice energy in the voice data is greater than a first preset threshold; If so, it includes remote voice data; If not, it does not include remote voice data.

8. The method according to claim 1, wherein Before the step of extracting the features of the source domain voice sample based on the feature extraction network to obtain the first voice feature, it further includes: Determine whether the voice sample data of the source domain and the voice sample data of the target domain are greater than or equal to a preset quantity, and whether the difference between the confidence of the voice recognition result output by the voice recognition network for the source domain voice sample and the confidence of the voice recognition result output by the voice recognition network for the target domain voice sample is greater than or equal to a second preset threshold; If so, execute the step of extracting the features of the source domain voice sample based on the feature extraction network to obtain the first voice feature.

9. The method according to claim 1, characterized in that, The target domain includes multiple target domains determined based on the signal echo energy ratio.

10. The method according to claim 1, wherein The training process of the voice recognition model is completed in the cloud server, and the text label of the source domain voice sample is obtained based on manual annotation.

11. The method according to claim 1, wherein The training process of the voice recognition model is completed on the terminal device; The step of obtaining the voice sample data of the source domain and the voice sample data of the target domain includes: The terminal device collects voice data, determines whether the voice data includes remote voice data. If so, use the voice data as a target domain voice sample and label the domain label of the target domain. If not, use the voice data as a source domain voice sample and label the domain label of the source domain; The text label of the source domain voice sample is obtained by recognition based on the acoustic model preset in the terminal device.

12. The method according to any one of claims 1-11, characterized in that, The voice recognition result is the voice content of the voice data to be recognized, or whether the voice data to be recognized contains a wake-up word, or whether the voiceprint feature of the voice data to be recognized matches a preset voiceprint feature.

13. A voice recognition device, characterized in that, It includes: An acquisition module, configured to acquire voice sample data of the source domain and voice sample data of the target domain. The voice sample data of the source domain includes a source domain voice sample, the text label of the source domain voice sample, and the domain label of the source domain. The voice sample data of the target domain includes a target domain voice sample and the domain label of the target domain; wherein, part of the distribution features of the source domain can be transferred and learned to the target domain. The source domain voice sample does not include echo data, and the target domain voice sample includes echo data. The echo data includes at least one or more of the following: audio data processed by an echo cancellation module, remote voice data, echo estimation data of the echo cancellation module, residual echo data of the echo suppression module, and echo path data estimated by the echo cancellation module; An extraction module, configured to extract the features of the source domain voice sample based on a feature extraction network to obtain a first voice feature, and extract the features of the target domain voice sample based on the feature extraction network to obtain a second voice feature; A training module, configured to input the first voice feature as a sample feature into a speech recognition network to obtain a speech recognition result; input the first voice feature and the second voice feature as sample features into a domain classification network to obtain a domain classification result; and jointly train the feature extraction network, the speech recognition network, and the domain classification network according to the speech recognition result and the text label, and according to the domain classification result and the domain label, to obtain the trained feature extraction network and speech recognition network as a speech recognition model; A recognition module, configured to input the voice data to be recognized into the trained speech recognition model to obtain a speech recognition result.

14. The device according to claim 13, characterized in that, The domain classification network includes a gradient reversal layer and a domain classification layer; The gradient reversal layer makes the feature extraction network and the domain classification layer form an adversarial relationship, and inputs a source domain voice sample and a domain label of the source domain, and a target domain voice sample and a domain label of the target domain into the feature extraction network respectively to train the feature extraction network and the domain classification layer.

15. The device according to claim 14, characterized in that, Training the feature extraction network and the domain classification layer includes: A forward propagation training process, in which the voice feature extracted by the feature extraction network is input into the domain classification layer through the gradient reversal layer, and the domain classification layer updates the parameters of the domain classification layer according to the domain classification result and the domain label; A backward propagation training process, in which the gradient reversal layer reverses the gradient of the domain classification layer and passes it to the feature extraction network to update the parameters of the feature extraction network.

16. The device according to claim 13, characterized in that, The training module is specifically configured to: Perform weighted summation on the loss function of the speech recognition network and the loss function of the domain classification network to obtain a total loss function; Jointly train the feature extraction network, the speech recognition network, and the domain classification network by minimizing the total loss function.

17. The device according to claim 13, characterized in that The acquisition module is specifically configured to: Collect voice data and determine whether the voice data includes remote voice data; If so, use the voice data as a target domain voice sample and label the domain label of the target domain; If not, use the voice data as a source domain voice sample and label the text label and the domain label of the source domain.

18. The device according to any one of claims 13-17, characterized in that The device further includes: A judgment module, configured to judge whether the voice sample data of the source domain and the voice sample data of the target domain are greater than or equal to a preset quantity, and whether the difference between the confidence of the speech recognition result output by the speech recognition network for the source domain voice sample and the confidence of the speech recognition result output by the speech recognition network for the target domain voice sample is greater than or equal to a second preset threshold; If so, execute extracting the feature of the source domain voice sample based on the feature extraction network to obtain the first voice feature.

19. A voice recognition device, characterized in that, It includes a memory and a processor, characterized in that executable code is stored in the memory, and the processor executes the executable code to implement the method according to any one of claims 1-12.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed on a computer, the computer is made to execute the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Cheating recording detecting neural network model optimization method and system

    CN110223676A

  • Speech recognition method based on domain-invariant feature

    CN110570845A