Speech recognition and model training methods and devices

By constructing a Gaussian hybrid model based on the audio acquisition device and the environment, determining the source probability distribution of the speech signal, and combining these probability distributions for wake-up word recognition, the problem of difficulty in speech recognition under different devices and environments is solved, and the accuracy of recognition is improved.

CN114582323BActive Publication Date: 2025-05-27LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210277532.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-21
Publication Date
2025-05-27
Estimated Expiration
2042-03-21

AI Technical Summary

Technical Problem

The audio acquisition devices and the acquisition environment on different electronic devices vary greatly, resulting in high difficulty in identifying voice wake-up words, prone to recognition errors, and low recognition performance.

Method used

By constructing a Gaussian hybrid model based on the audio distribution characteristics of multiple audio acquisition devices and environments, a Gaussian hybrid model is determined to determine the source probability distribution of the speech signal, and combining these probability distributions to recognize the speech signal wake-up word.

Benefits of technology

It improves the accuracy and reliability of speech wake-up word recognition, and is suitable for speech signal recognition in different audio acquisition devices and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114582323B_ABST
    Figure CN114582323B_ABST
Patent Text Reader

Abstract

The present application discloses a speech recognition and model training method and apparatus. The method includes: obtaining a speech signal to be recognized; determining a first probability distribution of the speech signal from various audio acquisition devices respectively based on the first audio distribution characteristics of the various audio acquisition devices, where the first audio distribution characteristics of the audio acquisition device characterize the audio distribution characteristics of the audio signal collected by the audio acquisition device; determining a second probability distribution of the speech signal from various audio acquisition environments respectively based on the second audio distribution characteristics of the various audio acquisition environments, where the second audio distribution characteristics of the audio acquisition environment characterize the audio distribution characteristics of the audio signal collected in the audio acquisition environment; combining at least one of the first probability distribution and the second probability distribution of the speech signal to perform wake-word recognition on the speech signal, and obtaining a wake-word recognition result. The solution of the present application can improve the accuracy of wake-word recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and more specifically, to a method and device for speech recognition and model training. Background Art

[0002] Speech wake-up word recognition, also known as wake-up word detection, refers to detecting whether a predefined wake-up word appears in a speech signal.

[0003] With the continuous development of electronic devices, there are significant differences in audio acquisition devices on different electronic devices, resulting in significant differences in the speech signals collected by the audio acquisition devices of different electronic devices; moreover, even for the same electronic device, the speech signals collected in different environments will also have significant differences. Due to the significant differences in speech signal acquisition devices and acquisition environments, the difficulty of speech wake-up word recognition is relatively high, and it is easy to have wake-up word recognition errors, resulting in low recognition performance of wake-up word recognition. Summary of the Invention

[0004] This application provides a method and device for speech recognition and model training.

[0005] Among them, a speech recognition method includes:

[0006] Obtain a speech signal to be recognized;

[0007] Based on the first audio distribution characteristics of various audio acquisition devices, determine the first probability distribution of the speech signal from each of the various audio acquisition devices, where the first audio distribution characteristic of the audio acquisition device represents the audio distribution characteristic of the audio signal collected by the audio acquisition device;

[0008] Based on the second audio distribution characteristics of various audio acquisition environments, determine the second probability distribution of the speech signal from each of the various audio acquisition environments, where the second audio distribution characteristic of the audio acquisition environment represents the audio distribution characteristic of the audio signal collected in the audio acquisition environment;

[0009] Combine at least one of the first probability distribution and the second probability distribution of the speech signal to perform wake-up word recognition on the speech signal to obtain a wake-up word recognition result.

[0010] In a possible implementation, the determining the first probability distribution of the speech signal from each of the various audio acquisition devices based on the first audio distribution characteristics of the various audio acquisition devices includes:

[0011] Determine the first probability distribution of the speech signal from each of the multiple audio acquisition devices based on the respective first Gaussian mixture models of the multiple audio acquisition devices, where the first Gaussian mixture model of the audio acquisition device is a Gaussian mixture model constructed based on the audio features of the audio signals collected by the audio acquisition device;

[0012] The determining the second probability distribution of the speech signal from each of the multiple audio acquisition environments based on the respective second audio distribution features of the multiple audio acquisition environments includes:

[0013] Determine the second probability distribution of the speech signal from each of the multiple audio acquisition environments based on the respective second Gaussian mixture models of the multiple audio acquisition environments, where the second Gaussian mixture model of the audio acquisition environment is a Gaussian mixture model constructed based on the audio features of the audio signals collected in the audio acquisition environment.

[0014] In a possible implementation, the performing wake word recognition on the speech signal by combining at least one of the first probability distribution and the second probability distribution of the speech signal to obtain a wake word recognition result includes:

[0015] Input at least one of the first probability distribution and the second probability distribution of the speech signal, and the speech signal into a wake word recognition model to obtain the wake word recognition result output by the wake word recognition model;

[0016] Wherein, the wake word recognition model is obtained through supervised training based on multiple audio training samples and at least one of the first probability distribution and the second probability distribution of the audio training samples, the first probability distribution of the audio training sample is the probability distribution of the audio training sample from each of the multiple audio acquisition devices, and the second probability distribution of the audio training sample is the probability distribution of the audio training sample from each of the multiple audio acquisition environments.

[0017] In yet another possible implementation, the first probability distribution of the speech signal includes the first probability that the speech signal comes from each of the multiple audio acquisition devices;

[0018] The second probability distribution of the speech signal includes: the second probability that the speech signal comes from each of the multiple audio acquisition environments;

[0019] The inputting at least one of the first probability distribution and the second probability distribution of the speech signal, and the speech signal into a wake word recognition model includes:

[0020] Construct a first vector from each first probability in the first probability distribution of the speech signal, and construct a second vector from each second probability in the second probability distribution of the speech signal;

[0021] Input at least one of the first vector and the second vector corresponding to the speech signal, and the speech signal into the wake word recognition model.

[0022] In another possible implementation, the wake word recognition model is trained as follows:

[0023] Obtain a first audio training set collected by each of multiple audio collection devices, where the first audio training set collected by the audio collection device includes at least one audio training sample;

[0024] Obtain a second audio training set for each of multiple audio collection environments, where the second audio training set for the audio collection environment includes at least one audio training sample collected in the audio collection environment;

[0025] Based on the first audio distribution characteristics of each of the multiple audio collection devices, determine the first probability distribution from which the audio training samples respectively originate from the multiple audio collection devices;

[0026] Based on the second audio distribution characteristics of each of the multiple audio collection environments, determine the second probability distribution from which the audio training samples respectively originate from the multiple audio collection environments;

[0027] Based on at least one of the first probability distribution and the second probability distribution of the audio training sample, and the audio training sample, train the wake word recognition model to be trained until the wake word recognition model meets the training end condition.

[0028] In another possible implementation, after obtaining the first audio training set of the audio collection device, it further includes:

[0029] Based on the audio characteristics of each audio training sample in the first audio training set of the audio collection device, construct a first Gaussian mixture model corresponding to the audio collection device, and the first Gaussian mixture model corresponding to the audio collection device is used to represent the audio distribution characteristics of the audio signal collected by the audio collection device;

[0030] After obtaining the second audio training set in the audio collection environment, it further includes:

[0031] Based on the audio characteristics of each audio training sample in the second audio training set of the audio collection environment, construct a second Gaussian mixture model corresponding to the audio collection environment, and the second Gaussian mixture model corresponding to the audio collection environment is used to represent the audio distribution characteristics of the audio signal collected in the audio collection environment;

[0032] Determining the first probability distribution from which the audio training samples respectively originate from the multiple audio acquisition devices based on the respective first audio distribution characteristics of the multiple audio acquisition devices includes:

[0033] Determining the first probability distribution from which the audio training samples respectively originate from the multiple audio acquisition devices based on the respective first Gaussian mixture models of the multiple audio acquisition devices;

[0034] Determining the second probability distribution from which the audio training samples respectively originate from the multiple audio acquisition environments based on the respective second audio distribution characteristics of the multiple audio acquisition environments includes:

[0035] Determining the second probability distribution from which the audio training samples respectively originate from the multiple audio acquisition environments based on the respective second Gaussian mixture models of the multiple audio acquisition environments.

[0036] In yet another possible implementation manner, combining at least one of the first probability distribution and the second probability distribution of the speech signal to perform wake word recognition on the speech signal includes:

[0037] Determining the audio feature of the speech signal;

[0038] Determining a set of model parameters required to process the speech signal based on at least one of the first probability distribution and the second probability distribution of the speech signal;

[0039] Combining the set of model parameters to determine the wake word recognition result corresponding to the audio feature.

[0040] Wherein, a model training method includes:

[0041] Obtaining a first audio training set collected by each of multiple audio acquisition devices, where the first audio training set collected by the audio acquisition device includes at least one audio training sample;

[0042] Obtaining a second audio training set of each of multiple audio acquisition environments, where the second audio training set of the audio acquisition environment includes at least one audio training sample collected in the audio acquisition environment;

[0043] Determining the first probability distribution from which the audio training samples respectively originate from the multiple audio acquisition devices based on the respective first audio distribution characteristics of the multiple audio acquisition devices;

[0044] Determining the second probability distribution from which the audio training samples respectively originate from the multiple audio acquisition environments based on the respective second audio distribution characteristics of the multiple audio acquisition environments;

[0045] Train the wake word recognition model to be trained based on at least one of the first probability distribution and the second probability distribution of the audio training samples, and the audio training samples, until the wake word recognition model meets the training end condition.

[0046] A voice recognition device includes:

[0047] A signal acquisition unit for acquiring a voice signal to be recognized;

[0048] A first probability determination unit for determining a first probability distribution of the voice signal from each of the multiple audio acquisition devices based on the first audio distribution characteristics of each of the multiple audio acquisition devices, where the first audio distribution characteristics of the audio acquisition device characterize the audio distribution characteristics of the audio signal acquired by the audio acquisition device;

[0049] A second probability determination unit for determining a second probability distribution of the voice signal from each of the multiple audio acquisition environments based on the second audio distribution characteristics of each of the multiple audio acquisition environments, where the second audio distribution characteristics of the audio acquisition environment characterize the audio distribution characteristics of the audio signal acquired in the audio acquisition environment;

[0050] A wake recognition unit for performing wake word recognition on the voice signal by combining at least one of the first probability distribution and the second probability distribution of the voice signal to obtain a wake word recognition result.

[0051] A model training device includes:

[0052] A first set acquisition unit for acquiring a first audio training set collected by each of the multiple audio acquisition devices, where the first audio training set collected by the audio acquisition device includes at least one audio training sample;

[0053] A second set acquisition unit for acquiring a second audio training set of each of the multiple audio acquisition environments, where the second audio training set of the audio acquisition environment includes at least one audio training sample collected in the audio acquisition environment;

[0054] A first distribution determination unit for determining a first probability distribution of the audio training sample from each of the multiple audio acquisition devices based on the first audio distribution characteristics of each of the multiple audio acquisition devices;

[0055] A second distribution determination unit for determining a second probability distribution of the audio training sample from each of the multiple audio acquisition environments based on the second audio distribution characteristics of each of the multiple audio acquisition environments;

[0056] A training control unit is configured to train a wake word recognition model to be trained based on at least one of a first probability distribution and a second probability distribution of the audio training samples and the audio training samples until the wake word recognition model meets the training end condition.

[0057] As can be seen from the above solution, after obtaining the speech signal to be recognized, the present application combines the first audio distribution features of various audio acquisition devices to determine the first probability distribution of the speech signal from various audio acquisition devices respectively; at the same time, it also combines the second audio distribution features of various audio acquisition environments to determine the second probability distribution of the speech signal from the various audio acquisition environments respectively. Since the first probability distribution can characterize the audio acquisition device used to acquire the speech signal to be recognized, and the second probability distribution can characterize the audio acquisition environment for acquiring the speech signal to be recognized, therefore, by performing wake word recognition in combination with at least one of the first probability distribution and the second probability distribution, the influence of at least one of the speech signal acquisition environment and the acquisition device can be considered for wake word recognition, thereby improving the accuracy of wake word recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0059] Figure 1 It is a schematic flowchart of a speech recognition method provided by an embodiment of the present application;

[0060] Figure 2 It is another schematic flowchart of a speech recognition method provided by an embodiment of the present application;

[0061] Figure 3 It is a schematic flowchart of a model training method provided by an embodiment of the present application;

[0062] Figure 4 It is another schematic flowchart of a speech recognition method provided by an embodiment of the present application;

[0063] Figure 5 It is another schematic flowchart of a model training method provided by an embodiment of the present application;

[0064] Figure 6 It is a schematic diagram of the composition structure of a speech recognition device provided by an embodiment of the present application;

[0065] Figure 7A schematic structural diagram of a component of the model training device provided in an embodiment of the present application;

[0066] Figure 8 A schematic structural diagram of a component of the electronic device provided in an embodiment of the present application.

[0067] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above-mentioned drawings are used to distinguish similar parts, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than that illustrated here. Specific embodiments

[0068] The speech recognition method of the present application can be applied to recognize whether there is a wake word in the speech signal to improve the accuracy and reliability of recognizing the wake word from the speech signal.

[0069] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0070] As Figure 1 shown, it shows a schematic flow diagram of the speech recognition method provided in an embodiment of the present application. The method of this embodiment can be applied to an electronic device. The method of this embodiment can be applied to any electronic device with a speech wake word recognition function, such as a mobile phone, a laptop computer, or a smart speaker, etc., without limitation.

[0071] The method of this embodiment may include:

[0072] S101, obtain the speech signal to be recognized.

[0073] Among them, the speech signal to be recognized is an audio signal collected by the audio collection device of the electronic device and needs to be detected (or recognized) for the wake word.

[0074] S102, based on the first audio distribution characteristics of each of the multiple audio collection devices, determine the first probability distribution of the speech signal from each of the multiple audio collection devices.

[0075] Among them, the first audio distribution characteristic of the audio collection device characterizes the audio distribution characteristic of the audio signal collected by the audio collection device. For the convenience of distinction, the audio distribution characteristic corresponding to the audio collection device is called the first audio distribution characteristic.

[0076] It can be understood that the audio distribution characteristics of different audio acquisition devices can be determined in advance and configured into the electronic device.

[0077] Among them, the first probability distribution may include the first probability that the voice signal respectively comes from each of the multiple audio acquisition devices. The first probability that the voice signal comes from a certain audio acquisition device can indicate the possibility that the voice signal is collected by this type of audio acquisition device. For example, the first probability that the audio signal comes from a certain audio acquisition device may be the similarity between the voice characteristics of the voice signal and the audio signal collected by the audio acquisition device.

[0078] It can be seen that the type of audio acquisition device that collects the voice signal can be reflected through the first probability distribution.

[0079] In a possible implementation manner, for each of the multiple audio acquisition devices, the first probability that the voice signal comes from the audio acquisition device can be determined based on the possibility that the audio characteristics of the voice signal conform to the first audio distribution characteristic of the audio acquisition device, or the similarity between the audio characteristics of the voice signal and the first audio distribution characteristic of the audio acquisition device.

[0080] It can be understood that the multiple audio acquisition devices mentioned here can be relatively scenario-based selected audio acquisition devices. For example, the audio acquisition devices configured in multiple different types of electronic devices can be selected, and the audio distribution characteristics of each of the multiple audio acquisition devices can be determined.

[0081] S103. Based on the second audio distribution characteristics of multiple audio acquisition environments respectively, determine the second probability distribution that the voice signal respectively comes from the multiple audio acquisition environments.

[0082] Among them, the audio acquisition environment reflects the environmental characteristics of collecting the audio signal. For example, the audio acquisition environment may include indoor quiet environment, indoor noisy environment, square, airport and other environments, without limitation.

[0083] Among them, the second audio distribution characteristic of the audio acquisition environment characterizes the audio distribution characteristic of the audio signal collected in the audio acquisition environment. In order to distinguish it from the audio distribution characteristic of the audio acquisition device, the audio distribution characteristic of the audio signal collected in the audio acquisition environment is called the second audio distribution characteristic.

[0084] The second probability distribution may include the second probability that the speech signal is from each of a variety of audio acquisition environments. The second probability that the speech signal is from a certain audio acquisition environment can indicate the likelihood that the speech signal is acquired in that audio acquisition environment. For example, the second probability can be the similarity between the audio features of the speech signal and the audio features of the audio signal acquired in that audio acquisition environment.

[0085] It can be seen that the characteristics of the audio acquisition environment for acquiring the speech signal can be reflected by this second probability distribution.

[0086] In yet another possible implementation, for each of a variety of audio acquisition environments, based on the likelihood that the audio features of the speech signal conform to the second audio distribution characteristics of that audio acquisition environment, or the similarity between the audio features of the speech signal and the second audio distribution characteristics of that audio acquisition environment, the second probability that the speech signal is from that audio acquisition environment is determined.

[0087] S104, Combining at least one of the first probability distribution and the second probability distribution of the speech signal, perform wake word recognition on the speech signal to obtain a wake word recognition result.

[0088] Among them, performing wake word recognition on the speech signal can be to recognize whether a wake word is included in the speech signal. Correspondingly, the wake word recognition result can be whether the speech signal includes a wake word. For example, the probability that a wake word is included in the speech signal can be recognized. If the probability is greater than a set threshold, it is determined that the speech signal includes a wake word.

[0089] Of course, in practical applications, the wake word recognition result can also be the probability that the speech signal includes a wake word, and there is no limitation on this.

[0090] It can be understood that the first probability distribution and the second probability distribution respectively reflect the characteristics of the audio acquisition device for acquiring the speech signal and the audio acquisition environment. And both the audio acquisition device for acquiring the speech signal and the audio acquisition environment of the speech signal may affect the wake word recognition of the speech signal. Based on this, in order to improve the accuracy of wake word recognition, the present application combines one or both of the characteristics reflecting that the speech signal is from a variety of different audio acquisition devices and the characteristics from a variety of different audio acquisition environments, that is, combines one or both of the first probability distribution and the second probability distribution, and performs wake word recognition on the speech signal, so that the influence of at least one of the audio acquisition environment and the audio acquisition device of the speech signal is considered in the wake word recognition process.

[0091] In an optional manner, the present application can combine the first probability distribution and the second probability distribution of the speech signal at the same time to perform wake word recognition on the speech signal.

[0092] As can be seen from the above, after obtaining the speech signal to be recognized, the present application combines the first audio distribution characteristics of various audio acquisition devices to determine the first probability distribution of the speech signal from various audio acquisition devices; at the same time, it also combines the second audio distribution characteristics of various audio acquisition environments to determine the second probability distribution of the speech signal from the various audio acquisition environments. Since the first probability distribution can characterize the audio acquisition device used to acquire the speech signal to be recognized, and the second probability distribution can characterize the audio acquisition environment for acquiring the speech signal to be recognized, therefore, by combining at least one of the first probability distribution and the second probability distribution for wake word recognition, the influence of at least one of the speech signal acquisition environment and the acquisition device can be considered for wake word recognition, thereby improving the accuracy of wake word recognition.

[0093] It can be understood that in the present application, there are various possible specific implementations for performing speech recognition on the speech signal by combining one of the first probability distribution and the second probability distribution of the speech signal.

[0094] For example, in one possible case, after determining the audio features of the speech signal, the present application can determine the model parameter set required to process the speech signal based on at least one of the first probability distribution and the second probability distribution of the speech signal. Correspondingly, the wake word recognition result corresponding to the audio features of the speech signal can be determined by combining the determined model parameter set.

[0095] Among them, when the first probability distribution of the speech signal from various different audio acquisition devices and the second probability distribution from various different audio acquisition environments change, the model parameter set will also change accordingly, thereby affecting the wake word recognition result of the speech signal.

[0096] For example, a model parameter library suitable for different audio acquisition devices and audio acquisition environments can be pre-built, and then, according to the differences in the audio acquisition devices and audio acquisition environments, different parameters can be selected from the model parameter library for combination to obtain a model parameter set suitable for wake word recognition.

[0097] In another possible case, the present application can use the trained wake word recognition model to combine at least one of the first probability distribution and the second probability distribution to determine the wake word recognition result of the speech signal. The following will illustrate this possible case with a flowchart. As Figure 2 shown, it shows another schematic flowchart of the speech recognition method provided by the embodiment of the present application. The method of this embodiment may include:

[0098] S201, obtain the speech signal to be recognized.

[0099] S202. Based on the first audio distribution characteristics of respective multiple audio collection devices, determine the first probability distribution of the voice signal originating from the multiple audio collection devices respectively.

[0100] For example, the first probability distribution may include the first probability of the voice signal originating from each of the multiple audio collection devices respectively.

[0101] S203. Based on the second audio distribution characteristics of respective multiple audio collection environments, determine the second probability distribution of the voice signal originating from the multiple audio collection environments respectively.

[0102] For example, the second probability distribution may include the first probability of the voice signal originating from each of the multiple audio collection environments respectively.

[0103] S204. Input at least one of the first probability distribution and the second probability distribution of the voice signal, and the voice signal into the wake word recognition model, to obtain the wake word recognition result output by the wake word recognition model.

[0104] It can be understood that inputting the voice signal into the wake word recognition model may be to first extract the audio features of the voice signal, and then input the audio features of the voice signal into the wake word recognition model.

[0105] Among them, the wake word recognition model is obtained through supervised training based on multiple audio training samples, and at least one of the first probability distribution and the second probability distribution of the audio training samples.

[0106] Among them, the audio training sample is the audio signal used to train the wake word recognition model. The first probability distribution of the audio training sample is the probability distribution of the audio training sample originating from multiple audio collection devices respectively. The second probability distribution of the audio training sample is the probability distribution of the audio training sample originating from multiple audio collection environments respectively.

[0107] It can be understood that the audio training samples for training the wake word recognition model in this application may include the audio signals of multiple different audio collection devices, or include the audio signals collected in multiple different audio collection environments. Of course, it may also include the audio signals collected by multiple different audio collection devices and the audio signals collected in multiple different audio collection environments at the same time.

[0108] Correspondingly, in the process of training the wake word recognition model of the present application, at least one of the first probability distribution of audio training samples from multiple audio acquisition devices and the second probability distribution of audio training samples from multiple different audio acquisition environments is combined, so that the trained wake word recognition model can comprehensively consider at least one of the audio acquisition device and the audio acquisition environment of the audio signal to perform wake word recognition. That is, the wake word recognition model can be applied to the wake word recognition of voice signals collected in different audio acquisition environments and by different audio acquisition devices, and can also reduce the situation where the wake word cannot be accurately recognized due to differences in the audio recording environment or audio acquisition device, thereby improving the accuracy of wake word recognition.

[0109] It can be understood that by training the wake word recognition model with audio training samples from different audio acquisition devices and different audio recording environments, the wake word recognition model can finally determine the model parameters suitable for recognizing audio signals from different audio acquisition devices and different audio acquisition environments. Based on this, in the process of using the wake word recognition model to perform wake word recognition on a voice signal, the wake word recognition model will also combine at least one of the first probability distribution of the voice signal from different audio acquisition devices and the second probability distribution from multiple audio acquisition environments to determine some model parameter sets required for the wake word recognition model to perform wake word recognition, so as to accurately recognize the wake word recognition result of the voice signal.

[0110] It can be understood that there are various possibilities for the process of the wake word recognition model. For the sake of easy understanding, an implementation manner of training the wake word recognition model will be described below as an example.

[0111] As Figure 3 shown, it shows a schematic flowchart of a model training method provided by an embodiment of the present application. The method of this embodiment may include:

[0112] S301, obtain the first audio training sets collected by various audio acquisition devices respectively.

[0113] Among them, the first audio training set collected by the audio acquisition device includes at least one audio training sample.

[0114] S302, obtain the second audio training sets of various audio acquisition environments respectively.

[0115] Among them, the second audio training set of the audio acquisition environment includes at least one audio training sample collected in this audio acquisition environment.

[0116] Among them, the order of steps S302 and S302 is not limited to Figure 3 shown. The order of these two steps can be interchanged or they can be executed synchronously.

[0117] S303. Determine the first probability distribution from which the audio training samples are respectively from multiple audio acquisition devices based on the respective first audio distribution characteristics of the multiple audio acquisition devices.

[0118] Among them, the first audio distribution characteristic of the audio signal collected by each audio acquisition device is determined based on the audio characteristics of the audio signal collected by the audio acquisition device.

[0119] Among them, the method for determining the first probability distribution from which the audio training samples are from different audio acquisition devices is similar to the process of determining the first probability distribution from which the speech signals are from each audio acquisition device before, and will not be elaborated here.

[0120] S304. Determine the second probability distribution from which the audio training samples are respectively from multiple audio acquisition environments based on the respective second audio distribution characteristics of the multiple audio acquisition environments.

[0121] Among them, the second audio distribution characteristic of each audio acquisition environment is determined based on the audio characteristics of the audio signal collected in the audio acquisition environment.

[0122] Among them, the process of determining the second probability distribution from which the audio training samples are from multiple audio acquisition environments is similar to the process of determining the second probability distribution of the speech signals before, and will not be elaborated here.

[0123] It can be understood that the order of steps S303 and S304 is not limited to Figure 3 as shown. In practical applications, these two steps can be interchanged or executed synchronously.

[0124] S305. Train the wake-up word recognition model to be trained based on at least one of the first probability distribution and the second probability distribution of the audio training samples, and the audio training samples, until the wake-up word recognition model meets the training end condition.

[0125] For example, the audio training samples can be labeled, and the label can represent whether the audio training sample is an audio containing a wake-up word. On this basis, one or both of the first probability distribution and the second probability distribution of the audio training samples, and the audio training samples can be used to train the wake-up word recognition model in a supervised training manner.

[0126] Among them, the training end condition can be: based on the label marked on the audio training sample, it is determined that the accuracy of the wake-up word recognition result of the wake-up word recognition model meets the requirements; or, based on the wake-up word recognition result recognized by the wake-up word recognition model, it is determined that the loss function value reaches convergence; or, the number of iterations of model training reaches a set number, etc., and there is no limit to this.

[0127] As introduced above, by training the wake word recognition model with audio training samples from different audio acquisition devices and different audio recording environments, the wake word recognition model can finally determine the model parameters suitable for recognizing audio signals from different audio acquisition devices and different audio acquisition environments. Thus, the trained wake word recognition model can be applied to audio signals collected in different audio acquisition environments and audio signals collected by different audio acquisition devices, which can improve the recognition accuracy of the wake word recognition model.

[0128] It can be understood that in this application, the first audio distribution feature of the audio acquisition device and the second audio distribution feature corresponding to the audio acquisition environment can have various forms. For the sake of easy understanding, a possible form of the audio distribution feature will be described below.

[0129] In a possible implementation, the audio distribution feature can be presented by a Gaussian mixture model. The Gaussian mixture model can more accurately represent the audio distribution feature of the audio signal.

[0130] Specifically, the first audio distribution feature of the audio acquisition device can be the first Gaussian mixture model of the audio acquisition device, and the first Gaussian mixture model of the audio acquisition device is a Gaussian mixture model constructed based on the audio features of the audio signals collected by the audio acquisition device. Correspondingly, when determining the first probability distribution of the speech signal, it can be: based on the first Gaussian mixture models of the respective multiple audio acquisition devices, determine the first probability distribution of the speech signal from the respective multiple audio acquisition devices.

[0131] Similarly, the first audio distribution feature of the audio acquisition environment can be the second Gaussian mixture model in this audio acquisition environment, and the second Gaussian mixture model is a Gaussian mixture model constructed based on the audio features of the audio signals collected in the audio acquisition environment. Correspondingly, the second probability distribution of the speech signal from multiple audio acquisition environments can be determined by combining the second Gaussian mixture models of the respective multiple audio acquisition environments.

[0132] For the sake of easy understanding, below, taking the first audio distribution feature as the first Gaussian mixture model, the second audio distribution feature as the second Gaussian mixture model, and combining the implementation of using the wake word recognition model to determine the wake word recognition result of the speech signal as an example for description.

[0133] As Figure 4 shown, it shows a schematic flowchart of another embodiment of a speech recognition method of this application. The method of this embodiment can include:

[0134] S401, obtain the speech signal to be recognized.

[0135] S402. Based on the first Gaussian mixture models of various audio acquisition devices, determine the first probability distributions of the speech signals from the various audio acquisition devices respectively.

[0136] Among them, the first Gaussian mixture model of an audio acquisition device is a Gaussian mixture model constructed based on the audio features of the audio signals collected by the audio acquisition device. The Gaussian mixture model can be regarded as a model composed of multiple single Gaussian models. Among them, the first Gaussian mixture model of the audio acquisition device can characterize the audio distribution characteristics of the audio signals collected by the audio acquisition device.

[0137] For example, for each type of audio acquisition device, multiple audio training samples collected by the audio acquisition device can be obtained, and the audio features of each audio training sample can be determined respectively. On this basis, a Gaussian mixture model can be constructed by combining the feature distributions of the audio features of the multiple audio training samples.

[0138] Of course, when the audio features of the audio training samples are known, there can be various processes for constructing the Gaussian mixture model, and there is no restriction on this.

[0139] Similar to the previous one, the first probability distribution can include the first probabilities of the speech signal from each of the various audio acquisition devices. Among them, the first probability of the speech signal from each audio acquisition device can be the possibility that the audio features of the speech signal belong to the first Gaussian mixture model of the audio acquisition device.

[0140] S403. Based on the second Gaussian mixture models of various audio acquisition environments, determine the second probability distributions of the speech signals from the various audio acquisition environments respectively.

[0141] Among them, the second Gaussian mixture model of an audio acquisition environment is a Gaussian mixture model constructed based on the audio features of the audio signals collected in the audio acquisition environment. The second Gaussian mixture model of the audio acquisition environment can characterize the audio distribution characteristics of the audio signals collected in the audio acquisition environment.

[0142] For example, for each type of audio acquisition environment, multiple audio training samples collected in the audio acquisition environment can be obtained, and the audio features of each audio training sample can be determined respectively. Then, a Gaussian mixture model can be constructed by combining the feature distributions of the audio features of the multiple audio training samples.

[0143] Similar to the previous one, the second probability distribution can include the second probabilities of the speech signal from each of the various audio acquisition environments. Among them, the second probability of the speech signal from each audio acquisition environment can be the possibility that the audio features of the speech signal belong to the second Gaussian mixture model of the audio acquisition environment.

[0144] S404. Input the voice signal, the first probability distribution, and the second probability distribution of the voice signal into the wake word recognition model to obtain the wake word recognition result output by the wake word recognition model.

[0145] Among them, the wake word recognition model is obtained through supervised training based on multiple audio training samples, the first probability distribution, and the second probability distribution of the audio training samples. The first probability distribution of the audio training samples is the probability distribution that the audio training samples are respectively from the multiple audio collection devices, and the second probability distribution of the audio training samples is the probability distribution that the audio training samples are respectively from multiple audio collection environments.

[0146] In a possible implementation, considering that the input of the wake word recognition model is generally in vector form, therefore, this application can form a first vector from each first probability in the first probability distribution of the voice signal, and form a second vector from each second probability in the second probability distribution of the voice signal. On this basis, the voice signal, the first vector, and the second vector of the voice signal can be input into the wake word recognition model.

[0147] Furthermore, before this step S404, this application can also extract the audio features of the voice signal. Correspondingly, the audio features of the voice signal, the first vector, and the second vector of the voice signal can be input into the wake word recognition model to obtain the wake word recognition result.

[0148] In this embodiment, taking the example of performing wake word recognition by combining the first probability distribution and the second probability distribution of the voice signal at the same time, of course, the present embodiment is also applicable to wake word recognition based on only any one of the first probability distribution and the second probability distribution of the voice signal. For example, based on at least one of the first vector and the second vector of the voice signal, the wake word recognition model can be used to recognize the wake word recognition result of the voice signal.

[0149] It can be understood that in the case of characterizing the audio distribution characteristics of the audio signal through a Gaussian mixture model, in the process of training the wake word recognition model in this application, the first probability distribution of the audio training samples can also be determined by combining the first Gaussian mixture model of each audio collection device, and the second probability distribution of the audio training samples can be determined according to the second Gaussian mixture model of each audio collection environment.

[0150] Furthermore, before training the wake word recognition model in this application, the first Gaussian mixture model and the second Gaussian mixture model can also be constructed by combining the audio features of the audio training samples used for training. Next, in combination with this implementation, the model training method provided by the embodiments of this application will be introduced.

[0151] Such as Figure 5As shown, it shows another schematic flowchart of the model training method provided by the embodiments of the present application. The method of this embodiment may include:

[0152] S501, obtain the first audio training sets collected by various audio acquisition devices respectively.

[0153] Among them, the first audio training set collected by the audio acquisition device includes at least one audio training sample.

[0154] Among them, for the convenience of distinction, the set composed of audio training samples collected by the audio acquisition device is called the first audio training set, and the set composed of audio training samples collected for the audio acquisition environment subsequently is called the second audio training set.

[0155] The various audio acquisition devices can be various different types of audio acquisition devices selected in advance, such as audio acquisition devices on various different types of electronic devices.

[0156] S502, based on the audio features of each audio training sample in the first audio training set of the audio acquisition device, construct a first Gaussian mixture model corresponding to the audio acquisition device.

[0157] The first Gaussian mixture model corresponding to the audio acquisition device is used to characterize the audio distribution characteristics of the audio signals collected by the audio acquisition device.

[0158] It can be understood that after constructing the first Gaussian mixture models of various audio acquisition devices respectively during the model training process, the first Gaussian mixture models of the various audio acquisition devices can be configured into the electronic devices that need to be involved in wake word recognition, so that the electronic devices can determine the first probability distribution of the voice signals collected by the electronic devices belonging to each audio acquisition device based on the first Gaussian mixture models of each audio acquisition device.

[0159] S503, obtain the second audio training sets of various audio acquisition environments respectively.

[0160] Among them, the second audio training set of each audio acquisition environment includes at least one audio training sample collected in this audio acquisition environment.

[0161] S504, based on the audio features of each audio training sample in the second audio training set of the audio acquisition environment, construct a second Gaussian mixture model corresponding to the audio acquisition environment.

[0162] Among them, the second Gaussian mixture model corresponding to the audio acquisition environment is used to characterize the audio distribution characteristics of the audio signals collected in the audio acquisition environment.

[0163] It can be understood that after constructing the second Gaussian mixture models corresponding to the audio features of various audio acquisition environments during the model training process, the second Gaussian mixture models corresponding to various audio acquisition environments can also be configured into the electronic device that needs to involve wake-up word recognition, so as to determine the second probability distribution of the voice signals collected by the electronic device belonging to each audio acquisition environment.

[0164] Among them, the process of constructing the first Gaussian mixture model and the second Gaussian mixture model can be referred to the previous introduction and will not be elaborated here.

[0165] S505. For each audio training sample, based on the first Gaussian mixture models of various audio acquisition devices, determine the first probability distribution of the audio training sample originating from various audio acquisition devices respectively.

[0166] S506. For each audio training sample, based on the second Gaussian mixture models of various audio acquisition environments, determine the second probability distribution of the audio training sample originating from various audio acquisition environments respectively.

[0167] It should be noted that this step S505 and S506 need to be executed for each audio training sample in the first audio training set and the second audio training set. Of course, the execution order of steps S505 and S506 is not limited to Figure 5 as shown. In practical applications, these two steps can be interchanged or executed synchronously.

[0168] S507. Based on the audio training sample and the first probability distribution and the second probability distribution of the audio training sample, train the wake-up word recognition model to be trained until the wake-up word recognition model meets the training end condition.

[0169] For example, the audio features of the audio training sample, the first vector converted from the first probability distribution of the audio training sample, and the second vector converted from the second probability distribution can be input into the wake-up word recognition model, and combined with the wake-up word recognition result predicted by the wake-up word recognition model and the wake-up word recognition result label annotated by the audio training sample, to determine whether the training end condition is met.

[0170] Among them, the training end condition can be referred to the relevant introduction in the previous embodiments and will not be elaborated here.

[0171] Corresponding to a voice recognition method provided in an embodiment of the present application, the present application also provides a voice recognition device.

[0172] Such as Figure 6 shown, which shows a schematic structural diagram of a composition of a voice recognition device provided in an embodiment of the present application. The device in this embodiment may include:

[0173] A signal acquisition unit 601, configured to acquire a speech signal to be recognized;

[0174] A first probability determination unit 602, configured to determine a first probability distribution of the speech signal from each of the multiple audio acquisition devices based on the respective first audio distribution characteristics of the multiple audio acquisition devices, where the first audio distribution characteristic of the audio acquisition device characterizes the audio distribution characteristic of the audio signal acquired by the audio acquisition device;

[0175] A second probability determination unit 603, configured to determine a second probability distribution of the speech signal from each of the multiple audio acquisition environments based on the respective second audio distribution characteristics of the multiple audio acquisition environments, where the second audio distribution characteristic of the audio acquisition environment characterizes the audio distribution characteristic of the audio signal acquired in the audio acquisition environment;

[0176] A wake-up recognition unit 604, configured to perform wake-up word recognition on the speech signal by combining at least one of the first probability distribution and the second probability distribution of the speech signal, and obtain a wake-up word recognition result.

[0177] In a possible implementation manner, the first probability determination unit includes:

[0178] A first probability determination subunit, configured to determine a first probability distribution of the speech signal from each of the multiple audio acquisition devices based on the respective first Gaussian mixture models of the multiple audio acquisition devices, where the first Gaussian mixture model of the audio acquisition device is a Gaussian mixture model constructed based on the audio characteristics of the audio signal acquired by the audio acquisition device;

[0179] The second probability determination unit includes:

[0180] A second probability determination subunit, configured to determine a second probability distribution of the speech signal from each of the multiple audio acquisition environments based on the respective second Gaussian mixture models of the multiple audio acquisition environments, where the second Gaussian mixture model of the audio acquisition environment is a Gaussian mixture model constructed based on the audio characteristics of the audio signal acquired in the audio acquisition environment.

[0181] In another possible implementation manner, the wake-up recognition unit includes:

[0182] A wake-up recognition subunit, configured to input at least one of the first probability distribution and the second probability distribution of the speech signal, and the speech signal into a wake-up word recognition model, and obtain a wake-up word recognition result output by the wake-up word recognition model;

[0183] Among them, the wake word recognition model is obtained through supervised training based on multiple audio training samples and at least one of the first probability distribution and the second probability distribution of the audio training samples. The first probability distribution of the audio training samples is the probability distribution of the audio training samples respectively from the multiple audio acquisition devices, and the second probability distribution of the audio training samples is the probability distribution of the audio training samples respectively from the multiple audio acquisition environments.

[0184] In an alternative manner, the first probability distribution of the speech signal determined by the first probability determination unit includes the first probability of the speech signal from each of the multiple audio acquisition devices.

[0185] The second probability distribution of the speech signal determined by the second probability determination unit includes: the second probability of the speech signal from each of the multiple audio acquisition environments.

[0186] The wake word recognition sub-unit includes:

[0187] A vector construction unit, configured to form a first vector with each of the first probabilities in the first probability distribution of the speech signal, and form a second vector with each of the second probabilities in the second probability distribution of the speech signal.

[0188] A model recognition sub-unit, configured to input at least one of the first vector and the second vector corresponding to the speech signal, and the speech signal into the wake word recognition model, to obtain the wake word recognition result output by the wake word recognition model.

[0189] In a possible implementation manner, the device further includes: a model training unit, configured to train the wake word recognition model in the following manner:

[0190] Obtain the first audio training sets respectively collected by the multiple audio acquisition devices, where the first audio training set collected by the audio acquisition device includes at least one audio training sample;

[0191] Obtain the second audio training sets of the multiple audio acquisition environments respectively, where the second audio training set of the audio acquisition environment includes at least one audio training sample collected in the audio acquisition environment;

[0192] Based on the first audio distribution features of the multiple audio acquisition devices respectively, determine the first probability distribution of the audio training samples respectively from the multiple audio acquisition devices;

[0193] Based on the second audio distribution features of the multiple audio acquisition environments respectively, determine the second probability distribution of the audio training samples respectively from the multiple audio acquisition environments;

[0194] Train the wake word recognition model to be trained based on at least one of the first probability distribution and the second probability distribution of the audio training samples, and the audio training samples, until the wake word recognition model meets the training end condition.

[0195] In a possible implementation, after obtaining the first audio training set of the audio acquisition device, the model training unit is further configured to: based on the audio features of each audio training sample in the first audio training set of the audio acquisition device, construct a first Gaussian mixture model corresponding to the audio acquisition device, where the first Gaussian mixture model corresponding to the audio acquisition device is used to characterize the audio distribution characteristics of the audio signals collected by the audio acquisition device;

[0196] After obtaining the second audio training set in the audio acquisition environment, the model training unit is further configured to: based on the audio features of each audio training sample in the second audio training set of the audio acquisition environment, construct a second Gaussian mixture model corresponding to the audio acquisition environment, where the second Gaussian mixture model corresponding to the audio acquisition environment is used to characterize the audio distribution characteristics of the audio signals collected in the audio acquisition environment;

[0197] When the model training unit determines the first probability distribution that the audio training samples respectively originate from the multiple audio acquisition devices based on the first audio distribution characteristics of the multiple audio acquisition devices, specifically, based on the first Gaussian mixture models of the multiple audio acquisition devices, determine the first probability distribution that the audio training samples respectively originate from the multiple audio acquisition devices;

[0198] When the model training unit determines the second probability distribution that the audio training samples respectively originate from the multiple audio acquisition environments based on the second audio distribution characteristics of the multiple audio acquisition environments, specifically, it is used to determine the second probability distribution that the audio training samples respectively originate from the multiple audio acquisition environments based on the second Gaussian mixture models of the multiple audio acquisition environments.

[0199] In yet another possible implementation, the wake word recognition unit includes:

[0200] A feature determination subunit, configured to determine the audio features of the speech signal;

[0201] A parameter set determination subunit, configured to determine a model parameter set required for processing the speech signal based on at least one of the first probability distribution and the second probability distribution of the speech signal;

[0202] An identification subunit, configured to combine the model parameter set to determine the wake word recognition result corresponding to the audio features.

[0203] Corresponding to the model training method provided in the embodiments of the present application, the present application further provides a model training device, as Figure 7 shown, which shows a schematic structural diagram of a composition of a model training device of the present application. The device includes:

[0204] A first set obtaining unit 701, configured to obtain first audio training sets respectively collected by a variety of audio collection devices, where the first audio training sets collected by the audio collection devices include at least one audio training sample;

[0205] A second set obtaining unit 702, configured to obtain second audio training sets of a variety of audio collection environments respectively, where the second audio training sets of the audio collection environments include at least one audio training sample collected in the audio collection environments;

[0206] A first distribution determining unit 703, configured to determine a first probability distribution of the audio training samples respectively from the variety of audio collection devices based on first audio distribution characteristics of the variety of audio collection devices;

[0207] A second distribution determining unit 704, configured to determine a second probability distribution of the audio training samples respectively from the variety of audio collection environments based on second audio distribution characteristics of the variety of audio collection environments;

[0208] A training control unit 705, configured to train a wake word recognition model to be trained based on at least one of the first probability distribution and the second probability distribution of the audio training samples, and the audio training samples, until the wake word recognition model meets a training end condition.

[0209] On the other hand, the present application further provides an electronic device, as Figure 8 shown, which shows a schematic structural diagram of a composition of the electronic device. The electronic device can be any type of electronic device. The electronic device at least includes a memory 801 and a processor 802;

[0210] Wherein, the processor 802 is configured to execute the speech recognition method or the model training method in any one of the above embodiments.

[0211] The memory 801 is configured to store a program required for the processor to perform operations.

[0212] It can be understood that the electronic device may further include a display unit 803 and an input unit 804.

[0213] Of course, the electronic device may further have Figure 8 more or fewer components, and no limitation is imposed thereon.

[0214] On the other hand, the present application also provides a computer-readable storage medium storing at least one instruction, at least one program, a code set or an instruction set, which is loaded and executed by a processor to implement the speech recognition method or the model training method described in any one of the above embodiments.

[0215] The present application also provides a computer program which includes computer instructions stored in a computer-readable storage medium. When the computer program runs on an electronic device, it is configured to execute the speech recognition method or the model training method in any one of the above embodiments.

[0216] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. At the same time, the features described in the embodiments of this specification can be replaced or combined with each other, enabling those skilled in the art to implement or use the present application. For the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply. The relevant parts can refer to the partial descriptions of the method embodiments.

[0217] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0218] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A voice recognition method, comprising: obtaining a voice signal to be recognized; determining a first probability distribution of the voice signal from each of the multiple audio acquisition devices based on the respective first audio distribution characteristics of the multiple audio acquisition devices, where the first audio distribution characteristic of the audio acquisition device characterizes the audio distribution characteristic of the audio signal acquired by the audio acquisition device; determining a second probability distribution of the voice signal from each of the multiple audio acquisition environments based on the respective second audio distribution characteristics of the multiple audio acquisition environments, where the second audio distribution characteristic of the audio acquisition environment characterizes the audio distribution characteristic of the audio signal acquired in the audio acquisition environment; inputting at least one of the first probability distribution and the second probability distribution of the voice signal, and the voice signal into a wake word recognition model to obtain a wake word recognition result output by the wake word recognition model.

2. The method according to claim 1, where determining the first probability distribution of the voice signal from each of the multiple audio acquisition devices based on the respective first audio distribution characteristics of the multiple audio acquisition devices comprises: determining the first probability distribution of the voice signal from each of the multiple audio acquisition devices based on the respective first Gaussian mixture models of the multiple audio acquisition devices, where the first Gaussian mixture model of the audio acquisition device is a Gaussian mixture model constructed based on the audio characteristics of the audio signal acquired by the audio acquisition device; The determining the second probability distribution of the voice signal from each of the multiple audio acquisition environments based on the respective second audio distribution characteristics of the multiple audio acquisition environments comprises: determining the second probability distribution of the voice signal from each of the multiple audio acquisition environments based on the respective second Gaussian mixture models of the multiple audio acquisition environments, where the second Gaussian mixture model of the audio acquisition environment is a Gaussian mixture model constructed based on the audio characteristics of the audio signal acquired in the audio acquisition environment.

3. The method according to claim 1 or 2, where the wake word recognition model is obtained through supervised training based on multiple audio training samples and at least one of the first probability distribution and the second probability distribution of the audio training samples, where the first probability distribution of the audio training sample is the probability distribution of the audio training sample from each of the multiple audio acquisition devices, and the second probability distribution of the audio training sample is the probability distribution of the audio training sample from each of the multiple audio acquisition environments.

4. The method according to claim 3, where the first probability distribution of the voice signal includes the first probability of the voice signal from each of the multiple audio acquisition devices; The second probability distribution of the voice signal comprises: the second probability of the voice signal from each of the multiple audio acquisition environments; The inputting at least one of the first probability distribution and the second probability distribution of the voice signal, and the voice signal into a wake word recognition model includes: Construct a first vector from each of the first probabilities in the first probability distribution of the speech signal, and construct a second vector from each of the second probabilities in the second probability distribution of the speech signal; Input at least one of the first vector and the second vector corresponding to the speech signal, and the speech signal into a wake word recognition model.

5. The method according to claim 3, wherein the wake word recognition model is trained as follows: Obtain a first audio training set collected by each of a plurality of audio collection devices, where the first audio training set collected by the audio collection device includes at least one audio training sample; Obtain a second audio training set for each of a plurality of audio collection environments, where the second audio training set for the audio collection environment includes at least one audio training sample collected in the audio collection environment; Based on the first audio distribution characteristics of each of the plurality of audio collection devices, determine the first probability distribution from which the audio training samples respectively originate from the plurality of audio collection devices; Based on the second audio distribution characteristics of each of the plurality of audio collection environments, determine the second probability distribution from which the audio training samples respectively originate from the plurality of audio collection environments; Based on at least one of the first probability distribution and the second probability distribution of the audio training samples, and the audio training samples, train the wake word recognition model to be trained until the wake word recognition model meets the training end condition.

6. The method according to claim 5, after obtaining the first audio training set of the audio collection device, it further includes: Based on the audio characteristics of each audio training sample in the first audio training set of the audio collection device, construct a first Gaussian mixture model corresponding to the audio collection device, and the first Gaussian mixture model corresponding to the audio collection device is used to represent the audio distribution characteristics of the audio signal collected by the audio collection device; After obtaining the second audio training set in the audio collection environment, it further includes: Based on the audio characteristics of each audio training sample in the second audio training set of the audio collection environment, construct a second Gaussian mixture model corresponding to the audio collection environment, and the second Gaussian mixture model corresponding to the audio collection environment is used to represent the audio distribution characteristics of the audio signal collected in the audio collection environment; The determining the first probability distribution from which the audio training samples respectively originate from the plurality of audio collection devices based on the first audio distribution characteristics of each of the plurality of audio collection devices includes: Based on the first Gaussian mixture models of each of the plurality of audio collection devices, determine the first probability distribution from which the audio training samples respectively originate from the plurality of audio collection devices; The determining the second probability distribution from which the audio training samples respectively originate from the plurality of audio collection environments based on the second audio distribution characteristics of each of the plurality of audio collection environments includes: Based on the second Gaussian mixture models of each of the plurality of audio collection environments, determine the second probability distribution from which the audio training samples respectively originate from the plurality of audio collection environments.

7. A speech recognition method, including: Obtain a speech signal to be recognized; Based on the first audio distribution characteristics of various audio acquisition devices, determine the first probability distribution of the speech signal originating from the various audio acquisition devices respectively, where the first audio distribution characteristics of the audio acquisition device characterize the audio distribution characteristics of the audio signal collected by the audio acquisition device; Based on the second audio distribution characteristics of various audio acquisition environments, determine the second probability distribution of the speech signal originating from the various audio acquisition environments respectively, where the second audio distribution characteristics of the audio acquisition environment characterize the audio distribution characteristics of the audio signal collected in the audio acquisition environment; Determine the audio characteristics of the speech signal; Based on at least one of the first probability distribution and the second probability distribution of the speech signal, determine the set of model parameters required to process the speech signal; Combined with the set of model parameters, determine the wake-up word recognition result corresponding to the audio characteristics.

8. A model training method, including: Obtain the first audio training set collected by various audio acquisition devices respectively, where the first audio training set collected by the audio acquisition device includes at least one audio training sample; Obtain the second audio training set of various audio acquisition environments respectively, where the second audio training set of the audio acquisition environment includes at least one audio training sample collected in the audio acquisition environment; Based on the first audio distribution characteristics of various audio acquisition devices, determine the first probability distribution of the audio training sample originating from the various audio acquisition devices respectively; Based on the second audio distribution characteristics of various audio acquisition environments, determine the second probability distribution of the audio training sample originating from the various audio acquisition environments respectively; Input at least one of the first probability distribution and the second probability distribution of the audio training sample, and the audio training sample into the wake-up word recognition model to be trained, and train the wake-up word recognition model to be trained until the wake-up word recognition model meets the training end condition.

9. A speech recognition device, including: A signal acquisition unit for acquiring the speech signal to be recognized; A first probability determination unit for determining the first probability distribution of the speech signal originating from the various audio acquisition devices respectively based on the first audio distribution characteristics of the various audio acquisition devices, where the first audio distribution characteristics of the audio acquisition device characterize the audio distribution characteristics of the audio signal collected by the audio acquisition device; A second probability determination unit for determining the second probability distribution of the speech signal originating from the various audio acquisition environments respectively based on the second audio distribution characteristics of the various audio acquisition environments, where the second audio distribution characteristics of the audio acquisition environment characterize the audio distribution characteristics of the audio signal collected in the audio acquisition environment; A wake-up recognition unit for inputting at least one of the first probability distribution and the second probability distribution of the speech signal, and the speech signal into the wake-up word recognition model to obtain the wake-up word recognition result output by the wake-up word recognition model.

10. A model training device, including: A first set acquisition unit, configured to acquire a first audio training set collected by each of a plurality of audio acquisition devices, where the first audio training set collected by the audio acquisition device includes at least one audio training sample; A second set acquisition unit, configured to acquire a second audio training set of each of a plurality of audio acquisition environments, where the second audio training set of the audio acquisition environment includes at least one audio training sample collected in the audio acquisition environment; A first distribution determination unit, configured to determine a first probability distribution from which the audio training samples respectively originate from the plurality of audio acquisition devices based on the first audio distribution characteristics of the plurality of audio acquisition devices; A second distribution determination unit, configured to determine a second probability distribution from which the audio training samples respectively originate from the plurality of audio acquisition environments based on the second audio distribution characteristics of the plurality of audio acquisition environments; A training control unit, configured to input at least one of the first probability distribution and the second probability distribution of the audio training samples, and the audio training samples, into a wake-up word recognition model to be trained, and train the wake-up word recognition model to be trained until the wake-up word recognition model meets the training end condition.

Citation Information

Patent Citations

  • System and method for recognizing environmental sound

    CN103370739A

  • Method and device for recognizing wake-up words, medium and equipment

    CN110047485A