Speech Recognition Method and Related Devices, Electronic Devices, and Storage Media
Through unsupervised training of the encoding network and supervised training of the decoding network, the insufficient speech recognition performance under low signal-to-noise ratio and low resource conditions is solved, and efficient speech recognition in these scenarios is achieved.
Patent Information
- Application Number
- CN202210514378.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-05-11
AI Technical Summary
The existing speech recognition technology lacks performance in low signal-to-noise ratio and low resource scenarios, making it difficult to effectively extract effective characterization information in noisy speech.
Unsupervised training is performed using coding networks, and the comparison loss training of sample clean speech and noisy speech is performed, combined with supervised training of decoding networks, the demand for labeled data is reduced and speech recognition performance is improved.
Under the conditions of low signal-to-noise ratio and low resource, the accuracy and efficiency of speech recognition are improved and adapted to the needs of low resource scenarios.
Smart Images

Figure CN114842833B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech recognition, and in particular, to a speech recognition method, related devices, electronic devices, and storage media. Background Art
[0002] Speech recognition refers to recognizing the input speech and automatically converting the speech information into text. Currently, low signal-to-noise ratio and low resources are important problems faced by current speech recognition technologies.
[0003] Specifically, in terms of low resources, due to the lack of rich and labeled training data, it is difficult for speech recognition models to be fully learned. In terms of low signal-to-noise ratio, conventional methods are difficult to extract effective characterization information from noisy speech, and the speech recognition results drop sharply, unable to meet the low-resource speech recognition requirements under low signal-to-noise ratio conditions. However, in real-world scenarios, affected by communication conditions such as transmission media and transmission protocols, as well as environmental noise, there will inevitably be low signal-to-noise ratio speech that needs to be recognized. In view of this, how to improve speech recognition performance in low signal-to-noise ratio and low-resource scenarios has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by the present application is to provide a speech recognition method, related devices, electronic devices, and storage media, which can improve speech recognition performance in low signal-to-noise ratio and low-resource scenarios.
[0005] To solve the above technical problem, a first aspect of the present application provides a speech recognition method, including: obtaining the speech to be recognized; recognizing the speech to be recognized based on a speech recognition model to obtain a recognition text; wherein, the speech recognition model includes an encoding network and a decoding network, the encoding network is trained based on a contrast loss between frame-level first quantization features after feature clustering and quantization of a sample first clean speech and frame-level noisy speech features of a sample first noisy speech, the sample first noisy speech is obtained by adding noise to the sample first clean speech, and the decoding network is trained in a supervised manner based on a sample second noisy speech after the encoding network converges.
[0006] To solve the above technical problem, a first aspect of the present application provides a speech recognition device, including: an obtaining module and a recognition module, the obtaining module is configured to obtain the speech to be recognized; the recognition module is configured to recognize the speech to be recognized based on a speech recognition model to obtain a recognition text; wherein, the speech recognition model includes an encoding network and a decoding network, the encoding network is trained based on a contrast loss between frame-level first quantization features of a sample first clean speech after quantization and frame-level noisy speech features of a sample first noisy speech, the sample first noisy speech is obtained by adding noise to the sample first clean speech, and the decoding network is trained based on a sample second noisy speech labeled with a sample recognition text after the encoding network converges.
[0007] To solve the above technical problems, a third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other. Program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the speech recognition method in the first aspect above.
[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the speech recognition method in the first aspect above.
[0009] In the above solution, the speech to be recognized is obtained, and the speech to be recognized is recognized based on a speech recognition model to obtain a recognition text. The speech recognition model includes an encoding network and a decoding network. The encoding network is trained based on a contrast loss between frame-level first quantization features obtained by feature clustering and quantization of a sample first clean speech and frame-level noisy speech features of a sample first noisy speech, and the sample first noisy speech is obtained by adding noise to the sample first clean speech. The decoding network is supervised-trained based on a sample second noisy speech after the encoding network converges. On the one hand, the training of the encoding network can be completed through unsupervised training by speech feature contrast, and the decoding network is supervised-trained after the encoding network converges, so the demand for labeled data can be reduced as much as possible, thus adapting to low-resource scenarios. On the other hand, by constructing a contrast prediction task between the sample first noisy speech and the sample first clean speech, the encoding network can learn to extract effective speech representations from the noisy speech as much as possible for the clean speech, thus adapting to low signal-to-noise ratio scenarios. Therefore, through the unsupervised contrast prediction task of the encoding network and the supervised recognition prediction task of the decoding network, the speech recognition performance can be improved in low signal-to-noise ratio and low-resource scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a flowchart of an embodiment of the speech recognition method of the present application;
[0011] Figure 2 is a framework diagram of an embodiment of training a speech recognition model;
[0012] Figure 3 is a framework diagram of an embodiment of a deep feature extraction sub-network;
[0013] Figure 4 is a flowchart of an embodiment of training a speech recognition model;
[0014] Figure 5 is a framework diagram of an embodiment of the speech recognition device of the present application;
[0015] Figure 6It is a schematic diagram of the framework of an embodiment of the electronic device of the present application;
[0016] Figure 7 It is a schematic diagram of the framework of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners
[0017] The following will combine with the accompanying drawings of the specification to elaborate on the solutions of the embodiments of the present application in detail.
[0018] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0019] In this article, the terms "system" and "network" are often used interchangeably. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.
[0020] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the voice recognition method of the present application. Specifically, it may include the following steps:
[0021] Step S11: Obtain the voice to be recognized.
[0022] In the embodiments of the present disclosure, the voice to be recognized can be either clean voice or noisy voice, which is not limited herein. Exemplarily, when making a long-distance call through a communication device such as a mobile phone, it is inevitably affected by communication conditions such as the transmission medium and transmission protocol, as well as environmental noise, and the voice signal is often mixed with noise signals. Therefore, the voice to be recognized collected in this case is noisy voice; or, in a space with good sound insulation such as a recording studio, there is almost no environmental noise. Therefore, the voice to be recognized collected in this environment is clean voice.
[0023] It should be noted that clean voice and noisy voice can be absolute concepts, that is, clean voice can be free of any noise signals, while noisy voice can be mixed with noise signals; or, clean voice and noisy voice can also be relative concepts, that is, the noise signal mixed in clean voice is lower than a preset threshold, that is, the signal-to-noise ratio of clean voice is relatively high and will not cause any interference to the recognition voice, and the noise signal mixed in the recognition voice is not lower than the preset threshold, that is, the signal-to-noise ratio of noisy voice is relatively high and will cause interference to the recognition voice.
[0024] Step S12: Recognize the speech to be recognized based on the speech recognition model to obtain the recognized text.
[0025] In the embodiments of the present disclosure, the speech recognition model includes an encoding network and a decoding network. The encoding network is trained based on the contrast loss between the frame-level first quantization features after feature clustering and quantization of the sample first clean speech and the frame-level noisy speech features of the sample first noisy speech. The sample first noisy speech is obtained by adding noise to the sample first clean speech. The decoding network is supervised-trained based on the sample second noisy speech after the encoding network converges. In addition, it should be noted that the "frame-level... features" (such as frame-level noisy speech features, frame-level first quantization features, etc.) referred to in the embodiments of the present disclosure represent the features at each speech frame level, that is, the features extracted on the basis of the speech frame granularity, unless otherwise specified.
[0026] In an implementation scenario, the sample first clean speech can be collected in advance, and then noise data is added to the sample first clean speech to obtain the sample first noisy speech. It should be noted that the noise data can be set based on the specific scenario of speech recognition. Exemplarily, when speech recognition is applied to indoor mobile communication, the noise data can include the voices of people around; or when speech recognition is applied to outdoor mobile communication, the noise data can include car honks, wind noise, etc. Other situations can be inferred by analogy and will not be exemplified one by one here.
[0027] In an implementation scenario, the frame-level first quantization features can be obtained by clustering and quantizing the frame-level first speech features of the sample first clean speech by a clustering model, and before each round of training of the encoding network, the clustering model is pre-trained based on the frame-level second speech features of the sample second clean speech. It should be noted that the clustering model can be a mathematical model of clustering algorithms such as K-Means, or a neural network model for realizing deep clustering, which is not limited here. In the above manner, the frame-level first quantization features are obtained by clustering and quantizing the frame-level first speech features of the sample first clean speech by the clustering model, and before each round of training of the encoding network, the clustering model is pre-trained based on the frame-level second speech features of the sample second clean speech. Therefore, the clustering model can assist in the training of the encoding network, which helps to promote the performance improvement of the encoding network during the training process.
[0028] In a specific implementation scenario, the sample second clean speech can be different from the aforementioned sample first clean speech. Exemplarily, at least 100 hours of clean speech can be collected in advance and the speech format can be regularized, such as uniformly processed to 16kHz 16bit, as the sample second clean speech. Of course, at least part of the sample first clean speech can also be selected as the sample second clean speech, which is not limited here.
[0029] In a specific implementation scenario, before the first-round training of the encoding network, in order to guide the encoding network to accurately segment acoustic units subsequently, the frame-level second speech feature can be the frame-level acoustic feature of the sample second clean speech (e.g., Mel cepstral coefficient feature), and during the first-round training of the encoding network, the frame-level first speech feature can be the frame-level acoustic feature of the sample first clean speech (e.g., Mel cepstral coefficient feature). In addition, the feature dimension of the frame-level acoustic feature is not limited here. Exemplarily, it can be set to 39 dimensions, etc. In the above manner, before the first-round training of the encoding network, the frame-level second speech feature is the frame-level acoustic feature of the sample second clean speech, and during the first-round training of the encoding network, the frame-level first speech feature is the frame-level acoustic feature of the sample first clean speech, which can guide the encoding network to segment the pronunciation unit boundary as accurately as possible during the training phase and improve the refinement of the acoustic unit modeling of the encoding network.
[0030] In a specific implementation scenario, different from the first-round training of the encoding network, after the first-round training of the encoding network, the frame-level first speech feature can be extracted from the sample first clean speech by the encoding network obtained from the latest training, and the frame-level second speech feature can be extracted from the sample second clean speech by the encoding network obtained from the latest training. For the specific process of the encoding network to extract the frame-level first speech feature and the frame-level second speech feature, reference can be made to the following specific description of the encoding network, which will not be elaborated here for the time being. In the above manner, after the first-round training of the encoding network, the frame-level first speech feature is extracted from the sample first clean speech by the encoding network obtained from the latest training, and the frame-level second speech feature is extracted from the sample second clean speech by the encoding network obtained from the latest training, which can further optimize the network performance of the encoding network by combining multiple rounds of iteration of the encoding network after guiding the encoding network to learn the pronunciation unit boundary.
[0031] In a specific implementation scenario, before each round of training of the encoding network, the frame-level second speech features of the sample second-cleanest speech can be obtained first, and then these frame-level second speech features can be used as training data. The clustering model is used to perform frame-level clustering on these frame-level second speech features, and based on the difference between the frame-level clustering result and the clustering expected result, the model parameters of the clustering model are adjusted. Exemplarily, taking the mathematical model of the K-Means algorithm as the clustering model, the number of clustering centers can be preset to N (for example, 100). After using the clustering model to perform frame-level clustering on the above frame-level second speech features and obtaining N clustering sets, the model parameters of the mathematical model, such as the clustering decision threshold, can be appropriately adjusted according to the clustering result. It should be noted that the number of clustering centers N can be set according to actual application needs. For example, in the case of higher requirements for segmentation refinement, N can be set to be relatively large, or in the case of relatively loose requirements for segmentation refinement, N can be set to be slightly smaller, which is not limited here. In addition, in the case where the clustering model is a neural network model based on deep clustering, the training process can be deduced by analogy and will not be elaborated here.
[0032] In a specific implementation scenario, whether before the first round of training of the encoding network or after the first round of training of the encoding network, during each round of training of the encoding network, once the clustering model has completed the previous training and obtained the frame-level first speech features of the sample first-cleanest speech, the frame-level first speech features can be processed based on the clustering model to obtain several clustering sets, and the frame-level first speech features located in the same clustering set are subjected to feature statistics to obtain the statistical speech features of the clustering set. Then, the statistical speech features of the clustering set are used as the frame-level first quantization features of the speech frames to which the frame-level first speech features in the clustering set belong. It should be noted that the feature statistics can include, but are not limited to, taking the average of each frame-level first speech feature. In addition, for the convenience of description, the frame-level first quantization feature of the t-th speech frame in the sample first-cleanest speech can be denoted as where the superscript k indicates that the frame-level first speech feature of the t-th speech frame is clustered into the k-th clustering set. In the above manner, the frame-level first speech features are processed based on the clustering model to obtain several clustering sets, and the frame-level first speech features located in the same clustering set are subjected to feature statistics to obtain the statistical speech features of the clustering set. Then, the statistical speech features of the clustering set are used as the frame-level first quantization features of the speech frames to which the frame-level first speech features in the clustering set belong. Therefore, the accuracy of the frame-level first quantization features can be improved.
[0033] In one implementation scenario, after obtaining the paired sample first clean speech and sample first noisy speech, the frame-level deep speech features of the sample first noisy speech can be extracted. In the case of masking several speech frames in the sample first noisy speech, context encoding is performed based on the frame-level deep speech features to obtain the frame-level noisy speech features of each speech frame in the sample first noisy speech. Then, by comparing the frame-level noisy speech features and the frame-level first quantization features at the same time sequence, the contrast loss can be obtained. It should be noted that the frame-level deep speech features contain the deep semantic information of the speech frame itself, while the frame-level noisy speech features not only contain the semantic information of the speech frame itself, but also further contain the semantic information of the speech frames (such as adjacent speech frames) related to its semantics. In addition, for the convenience of description, each speech frame of the sample first noisy speech can be denoted as I = [i1, i2, … i t , …, i T , where i t represents the speech frame at time sequence t. Further, the frame-level deep speech features extracted from each speech frame can be denoted as X = [x1, x2, … x t , …, x T , where x t represents the frame-level deep speech feature of the speech frame at time sequence t. Further, the frame-level noisy speech features extracted from each speech frame can be denoted as O = [o1, o2, … o t , …, o T , where o t represents the frame-level noisy speech feature of the speech frame at time sequence t. As mentioned above, the frame-level first quantization feature of the t-th speech frame in the sample first clean speech can be denoted as Then the frame-level first quantization feature and the frame-level noisy speech feature o t represent a pair of frame-level noisy speech features and frame-level first quantization features at the same time sequence. Other cases can be deduced by analogy and will not be exemplified one by one here. In the above manner, the frame-level deep speech features of the sample first noisy speech are extracted, and in the case of masking several speech frames in the sample first noisy speech, context encoding is performed based on the frame-level deep speech features to obtain the frame-level noisy speech features of each speech frame in the sample first noisy speech, so as to compare the frame-level noisy speech features and the frame-level first quantization features at the same time sequence to obtain the contrast loss. Therefore, in the training stage of the encoding network, by constructing a contrast learning task of unsupervised speech pairs, the encoding network can learn to extract the effective information of the noisy speech, which helps to improve the network performance of the encoding network in extracting the effective information of the noisy speech.
[0034] In a specific implementation scenario, please refer to Figure 2 , Figure 2 is a schematic framework diagram of an embodiment for training a speech recognition model. As Figure 2As shown, the encoding network may specifically include a depth feature extraction sub-network and a context encoding sub-network connected in sequence. The depth feature extraction sub-network is used to extract frame-level depth speech features, while the context encoding sub-network is used to perform context encoding. Exemplarily, the depth feature extraction sub-network and the context encoding sub-network may adopt a multi-layer cascade design. Please refer to Figure 3 , Figure 3 which is a schematic diagram of the framework of an embodiment of the depth feature extraction sub-network. As Figure 3 shown, the depth feature extraction sub-network may adopt a 7-layer one-dimensional convolutional structure for speech preprocessing to extract the depth features of speech. The convolutional dimension of each layer may be set to 512, the stride may be set to [5, 2, 2, 2, 2, 2, 2], and the convolutional kernel may be set to [10, 3, 3, 3, 3, 2, 2]. The specific structure of each layer of convolution may be referred to Figure 3 shown by the right dotted box in. In addition, the context encoding sub-network may be composed of multiple layers of Transformer structures. Each layer of Transformer may include a feed-forward network, a layer dropout network, and an attention network. The Transformer may be set to 12 layers, the inner dimension of the feed-forward network may be set to 3072, the number of heads of the attention network may be set to 8, and the final output feature dimension may be set to 768. It should be noted that the specific network structures of the above depth feature extraction sub-network and context encoding sub-network are only a possible implementation manner in the actual application process, and do not limit the specific structures of the depth extraction sub-network and context encoding sub-network accordingly. In the above manner, the encoding network includes a depth feature extraction sub-network and a context encoding sub-network connected in sequence. The depth feature extraction sub-network is used to extract frame-level depth speech features, and the context encoding sub-network is used to perform context encoding. Therefore, the network structure of the encoding network can be simplified as much as possible, so that in the subsequent application process of speech recognition, it can be directly applied only by splicing with the decoding network, which can contribute to the engineering implementation and the promotion and implementation of the solution.
[0035] In a specific implementation scenario, several masked speech frames in the sample first noisy speech may be temporally continuous speech frames, and their proportion in the sample first noisy speech may not be limited. For example, it may account for 35%, or it may account for 40%, or it may also be any proportion between 35% and 40%. This is not limited here. It should be noted that the above proportion is only a possible implementation manner in the actual application process, and does not limit the actual proportion of the masked speech frames in the sample first noisy speech accordingly. In addition, when the speech frames in the sample first noisy speech are masked, it means that when performing context encoding, the frame-level depth speech features of the masked speech frames will no longer be referred to.
[0036] In a specific implementation scenario, to improve the accuracy of feature comparison, before the formal comparison, the frame-level noisy speech feature can also be projected based on the feature projection parameters corresponding to the cluster set to which the frame-level first quantization feature at the same time sequence as the frame-level noisy speech feature belongs, to obtain the frame-level noisy projection feature of the frame-level noisy speech feature. On this basis, the comparison loss can be obtained based on the feature similarity between the frame-level first quantization feature at the same time sequence as the frame-level noisy speech feature and the frame-level noisy projection feature of the frame-level noisy speech feature. It should be noted that the greater the feature similarity, the smaller the comparison loss, and the smaller the feature similarity, the greater the comparison loss. Therefore, by minimizing the comparison loss, the encoding network can be forced to learn to extract effective speech representations from the noisy speech as much as possible for the clean speech, improving the speech coding performance of the encoding network for the noisy speech. For the sake of convenience of description, still taking the frame-level noisy speech feature o t as an example, the frame-level first quantization feature at the same time sequence as it is which belongs to the k-th cluster set, and the feature projection parameters corresponding to this cluster set can be denoted as A (k) , then the frame-level noisy speech feature o (k) can be projected using this feature projection parameter A t to obtain the frame-level noisy projection feature Exemplarily, in the case of presetting N cluster centers as described above, there are also N feature projection parameters accordingly. In the case of projecting the frame-level noisy speech features at other time sequences, the specific process of feature projection can be analogized accordingly, and no further examples will be given here. It should be noted that the feature projection parameters can be adjusted during the training process, that is, during the training of the encoding network, the network parameters and feature projection parameters of the encoding network can be adjusted based on the comparison loss. In the above method, the frame-level noisy speech feature is projected based on the feature projection parameters corresponding to the cluster set to which the frame-level first quantization feature at the same time sequence as the frame-level noisy speech feature belongs, to obtain the frame-level noisy projection feature of the frame-level noisy speech feature, and the comparison loss is obtained based on the feature similarity between the frame-level first quantization feature at the same time sequence as the frame-level noisy speech feature and the frame-level noisy projection feature of the frame-level noisy speech feature, which can improve the accuracy of feature comparison. In addition, during the training of the encoding network, adjusting the network parameters and feature projection parameters of the encoding network based on the comparison loss can help improve the training accuracy of the encoding network.
[0037] In a specific implementation scenario, as described above, during the calculation of the contrast loss, several speech frames in the first noisy speech of the sample are masked. In this case, the contrast loss can be obtained by weighting the first loss and the second loss. The first loss is obtained by comparing the frame-level noisy speech features and the frame-level first quantization features at the same first time sequence. The second loss is obtained by comparing the frame-level noisy speech features and the frame-level first quantization features at the same second time sequence. The first time sequence is the time sequence where the masked speech frames in the first noisy speech of the sample are located, and the second time sequence is the time sequence where the unmasked speech frames in the first noisy speech of the sample are located. Taking the first loss as an example, for the masked speech frames, after obtaining the feature similarities corresponding to each time sequence, a normalization function such as softmax can be used to normalize the feature similarities corresponding to each time sequence, and logarithmic summation is performed to obtain the first loss. For ease of description, the first loss can be denoted as L m :
[0038]
[0039]
[0040] In the above formulas (1) and (2), f represents the network parameters of the encoding network, X represents the first noisy speech of the sample, M represents the time sequence set of several masked speech frames in the first noisy speech of the sample, {C (k)} k represents the k-th clustering set, represents the masked speech frame in the first noisy speech of the sample, represents the speech frame at time sequence t in the first noisy speech of the sample, represents the speech frame at time sequence t in the first clean speech of the sample, and the frame-level first quantization feature of this speech frame corresponds to the k-th clustering set, represents the normalized feature similarity, sim represents the function for calculating the feature similarity (e.g., cosine similarity function), τ represents the temperature coefficient, and its value can be not limited, such as it can be set to 0.8, 0.9, etc. In addition, the specific meanings of A (k) , o t , can be referred to the relevant descriptions above and will not be elaborated here. The calculation method of the second loss can refer to the above formulas (1) and (2) of the first loss and will not be elaborated here. For ease of description, the second loss can be denoted as L u . On this basis, the first loss L m and the second loss L u can be weighted to obtain the contrast loss Loss:
[0041] Loss = L m + αL u……(3)
[0042] In the above formula (3), α represents an adjustable hyperparameter, which can be specifically set according to application needs. Exemplarily, when the masked part is more important for training the encoding network, α can be set to less than 1; when the unmasked part is more important for training the encoding network, α can be set to greater than 1; or when the masked part and the unmasked part are equally important for training the encoding network, α can be set to equal 1. There is no limitation here.
[0043] In one implementation scenario, after multiple iterations of training, the training of the encoding network converges. It should be noted that the convergence condition can be preset, and the convergence condition can include but is not limited to: the number of iterations is greater than a first threshold (such as 1000 times, 5000 times, etc.), and the contrast loss is less than a second threshold. There is no limitation here.
[0044] In one implementation scenario, after the training of the encoding network converges, the decoding network can be further trained. Specifically, a small number of sample second noisy voices labeled with sample recognition texts can be collected in advance, and the converged encoding network and the decoding network can be combined and connected, and the sample second noisy voice can be input. Thus, through the deep feature extraction sub-network and the context encoding sub-network in the encoding network in sequence, the frame-level noisy voice features of each voice frame in the sample second noisy voice can be obtained. It should be noted that different from the training process of the encoding network, in the training process of the decoding network, the voice frames in the sample second noisy voice can be not masked. On this basis, the decoding network can be used to decode the frame-level noisy voice features of each voice frame in the sample second noisy voice to obtain the predicted recognition text, and then based on the difference between the sample recognition text and the predicted recognition text, the recognition loss can be obtained, and based on the recognition loss, the network parameters of the decoding network can be adjusted, and the network parameters of the encoding network can be fixed. Exemplarily, the predicted recognition text can be obtained through several decodings. And each time of decoding, the decoding network is used to decode the frame-level noisy voice features of each voice frame in the sample second noisy voice and the decoded characters obtained from the previous decoding to obtain the decoded characters of this decoding until the decoded characters obtained from this decoding are end characters. Thus, the combination of the decoded characters obtained from each decoding can be used as the predicted recognition text. In addition, the specific calculation method of the recognition loss can refer to loss functions such as CTC loss, etc., which will not be elaborated here.
[0045] In an implementation scenario, after the decoding network is also trained and converges, the speech recognition model can be used to recognize the speech to be recognized, and the recognition text of the speech to be recognized can be obtained. Specifically, the deep feature extraction sub-network of the encoding network can be used to extract features from the speech to be recognized, and the frame-level deep speech features of each speech frame of the speech to be recognized can be obtained. Then, the context encoding sub-network of the encoding network can be used to perform context encoding on the frame-level deep speech features of each speech frame, and the frame-level context features of each speech frame can be obtained. Finally, the decoding network can be used to decode the frame-level context features of each speech frame to obtain the recognition text of the speech to be recognized. Exemplarily, the recognition text can be obtained through several decodings. And each time during decoding, the decoding network decodes the frame-level context features of each speech frame and the decoded characters obtained from the previous decoding to obtain the decoded characters of this decoding, until the decoded characters obtained from this decoding are end characters. Thus, the combination of the decoded characters obtained from each decoding can be used as the recognition text.
[0046] In the above solution, the speech to be recognized is obtained, and the speech recognition model is used to recognize the speech to be recognized to obtain the recognition text. The speech recognition model includes an encoding network and a decoding network. The encoding network is trained based on the contrast loss between the frame-level first quantization features after feature clustering and quantization of the sample first clean speech and the frame-level noisy speech features of the sample first noisy speech, and the sample first noisy speech is obtained by adding noise to the sample first clean speech. The decoding network is supervised-trained based on the sample second noisy speech after the encoding network is trained and converges. On the one hand, the training of the encoding network can be completed through unsupervised training by comparing speech features, and the decoding network is supervised-trained after the encoding network is trained and converges. Therefore, the demand for labeled data can be reduced as much as possible, so as to adapt to low-resource scenarios. On the other hand, by constructing a contrast prediction task between the sample first noisy speech and the sample first clean speech, the encoding network can learn to extract effective speech representations from the noisy speech as much as possible for the clean speech, so as to adapt to low signal-to-noise ratio scenarios. Therefore, through the unsupervised contrast prediction task of the encoding network and the supervised recognition prediction task of the decoding network, the speech recognition performance can be improved in low signal-to-noise ratio and low-resource scenarios.
[0047] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an embodiment for training a speech recognition model. Specifically, it may include the following steps:
[0048] Step S401: Extract the frame-level acoustic features of the sample first clean speech as the frame-level first speech features, and extract the frame-level acoustic features of the sample second clean speech as the frame-level second speech features.
[0049] For specific details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated here.
[0050] Step S402: Train a clustering model based on the frame-level second speech features of the sample second clean speech.
[0051] For specific details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0052] Step S403: Cluster and quantize the frame-level first speech features of the sample first clean speech based on the clustering model to obtain the frame-level first quantization features of the sample first clean speech.
[0053] For specific details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0054] Step S404: Extract the frame-level deep speech features of the sample first noisy speech.
[0055] For specific details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0056] Step S405: With several speech frames of the sample first noisy speech masked, perform context encoding based on the frame-level deep speech features to obtain the frame-level noisy speech features of each speech frame of the sample first noisy speech.
[0057] For specific details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0058] Step S406: For the masked speech frames, compare the frame-level noisy speech features and the frame-level first quantization features at the same time sequence to obtain a first loss, and for the unmasked speech frames, compare the frame-level noisy speech features and the frame-level first quantization features at the same time sequence to obtain a second loss.
[0059] For specific details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0060] Step S407: Weight the first loss and the second loss to obtain a comparison loss.
[0061] For specific details, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein.
[0062] Step S408: Adjust the network parameters of the encoding network based on the comparison loss.
[0063] As described in the foregoing disclosed embodiments, whether in the process of calculating the first loss or in the process of calculating the second loss, the frame-level noisy speech features can be feature-projected based on the feature projection parameters corresponding to the cluster set to which the frame-level first quantization features at the same time sequence as the frame-level noisy speech features belong, to obtain the frame-level noisy projection features of the frame-level noisy speech features, and then the corresponding loss can be obtained based on the feature similarity between the frame-level first quantization features at the same time sequence as the frame-level noisy speech features and the frame-level noisy projection features of the frame-level noisy speech features. In this case, the feature projection parameters and the network parameters of the encoding network can be adjusted based on the comparison loss.
[0064] Step S409: Determine whether the encoding network is trained to convergence. If not, execute step S410; otherwise, execute step S412.
[0065] Specifically, a convergence condition can be preset in advance. The convergence condition is used to determine whether the encoding network is trained to convergence. For the specific setting method of the convergence condition, reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated herein. In the case where the encoding network is trained to convergence, the decoding network can be continuously trained. In the case where the encoding network is not trained to convergence, the following step S410 and subsequent steps can be continuously executed to continue training the encoding network until it is trained to convergence.
[0066] Step S410: Extract the frame-level first speech features of the sample first clean speech based on the latest trained encoding network, and extract the frame-level second speech features of the sample second clean speech based on the latest trained encoding network.
[0067] Specifically, as in the previously disclosed embodiments, the encoding network can include a deep feature extraction sub-network and a context encoding sub-network. In this case, the sample first clean speech can be input into the latest trained encoding network, and the intermediate layer result of the encoding network can be taken to obtain the frame-level first speech features. Similarly, the sample second clean speech can be input into the latest trained encoding network, and the intermediate layer result of the encoding network can be taken to obtain the frame-level second speech features. Exemplarily, if the Transformer layer of the encoding network includes 24 layers, the 12th layer Transformer result can be taken as the intermediate layer result. Of course, in a real scenario, it is not limited to taking the intermediate layer result, and any layer result can also be taken according to actual application needs. Exemplarily, the last layer result can also be taken, which is not limited herein.
[0068] Step S411: Re-execute step S402 and subsequent steps.
[0069] Specifically, as described above, after the first-round training of the encoding network, the frame-level first speech feature is extracted from the sample first clean speech by the latest trained encoding network, and the frame-level second speech feature is extracted from the sample second clean speech by the latest trained encoding network. On this basis, the above-mentioned processes such as the clustering model training and the encoding network training can be re-executed, which will not be elaborated here.
[0070] Step S412: Train the decoding network based on the sample second noisy speech labeled with the sample recognition text and the encoding network with converged training.
[0071] Specific reference can be made to the relevant descriptions in the foregoing disclosed embodiments, which will not be elaborated here.
[0072] In the above solution, on the one hand, the training of the encoding network can be completed through unsupervised training by comparing speech features, and the decoding network is supervised trained after the encoding network training converges. Therefore, the demand for labeled data can be reduced as much as possible, so as to adapt to low-resource scenarios. On the other hand, by constructing a contrast prediction task between the sample first noisy speech and the sample first clean speech, the encoding network can learn to extract effective speech representations from the noisy speech as much as possible for the clean speech, so as to adapt to low signal-to-noise ratio scenarios. Therefore, through the unsupervised contrast prediction task of the encoding network and the supervised recognition prediction task of the decoding network, the speech recognition performance can be improved in low signal-to-noise ratio and low-resource scenarios.
[0073] Please refer to Figure 5 , Figure 5 which is a schematic framework diagram of an embodiment of the speech recognition device 50 of the present application. The speech recognition device 50 includes: an acquisition module 51 and a recognition module 52. The acquisition module 51 is used to acquire the speech to be recognized; the recognition module 52 is used to recognize the speech to be recognized based on the speech recognition model to obtain the recognition text. Among them, the speech recognition model includes an encoding network and a decoding network. The encoding network is trained based on the contrast loss between the frame-level first quantization feature after quantization of the sample first clean speech and the frame-level noisy speech feature of the sample first noisy speech. The sample first noisy speech is obtained by adding noise to the sample first clean speech. The decoding network is trained based on the sample second noisy speech labeled with the sample recognition text after the encoding network training converges.
[0074] In the above solution, the speech to be recognized is obtained, and the speech to be recognized is recognized based on a speech recognition model to obtain a recognized text. The speech recognition model includes an encoding network and a decoding network. The encoding network is trained based on a contrast loss between frame-level first quantization features obtained by feature clustering and quantization of a sample first clean speech and frame-level noisy speech features of a sample first noisy speech, and the sample first noisy speech is obtained by adding noise to the sample first clean speech. The decoding network is supervised-trained based on a sample second noisy speech after the encoding network converges. On the one hand, the training of the encoding network can be completed through unsupervised training by speech feature contrast, and the decoding network is supervised-trained after the encoding network converges. Therefore, the demand for labeled data can be minimized as much as possible, thus adapting to low-resource scenarios. On the other hand, by constructing a contrast prediction task for the sample first noisy speech and the sample first clean speech, the encoding network can learn to extract effective speech representations from the noisy speech as much as possible, thus adapting to low signal-to-noise ratio scenarios. Therefore, through the unsupervised contrast prediction task of the encoding network and the supervised recognition prediction task of the decoding network, the speech recognition performance can be improved in low signal-to-noise ratio and low-resource scenarios.
[0075] In some disclosed embodiments, the speech recognition device 50 includes a deep feature extraction module for extracting frame-level deep speech features of a sample first noisy speech; the speech recognition device 50 includes a context encoding module for performing context encoding based on the frame-level deep speech features in the case of masking a plurality of speech frames in the sample first noisy speech to obtain frame-level noisy speech features of each speech frame of the sample first noisy speech; the speech recognition device 50 includes a loss metric module for comparing the frame-level noisy speech features and the frame-level first quantization features at the same time sequence to obtain a contrast loss.
[0076] Therefore, frame-level deep speech features of the sample first noisy speech are extracted, and context encoding is performed based on the frame-level deep speech features in the case of masking a plurality of speech frames in the sample first noisy speech to obtain frame-level noisy speech features of each speech frame of the sample first noisy speech, so as to compare the frame-level noisy speech features and the frame-level first quantization features at the same time sequence to obtain a contrast loss. Therefore, in the training stage of the encoding network, by constructing a contrast learning task for unsupervised speech pairs, the encoding network can learn to extract effective information from the noisy speech, which helps to improve the network performance of the encoding network in extracting effective information from the noisy speech.
[0077] In some disclosed embodiments, the speech recognition device 50 includes a feature projection module, which is configured to perform feature projection on the frame-level noisy speech features based on the feature projection parameters corresponding to the clustering set to which the frame-level first quantization features located in the same time sequence as the frame-level noisy speech features belong, so as to obtain the frame-level noisy projection features of the frame-level noisy speech features; the loss metric module is specifically configured to obtain a contrastive loss based on the feature similarity between the frame-level first quantization features located in the same time sequence as the frame-level noisy speech features and the frame-level noisy projection features of the frame-level noisy speech features.
[0078] Therefore, performing feature projection on the frame-level noisy speech features based on the feature projection parameters corresponding to the clustering set to which the frame-level first quantization features located in the same time sequence as the frame-level noisy speech features belong, so as to obtain the frame-level noisy projection features of the frame-level noisy speech features, and obtaining a contrastive loss based on the feature similarity between the frame-level first quantization features located in the same time sequence as the frame-level noisy speech features and the frame-level noisy projection features of the frame-level noisy speech features can improve the accuracy of feature comparison.
[0079] In some disclosed embodiments, during the training process of the encoding network, based on the contrastive loss, the network parameters and feature projection parameters of the encoding network are adjusted.
[0080] Therefore, during the training process of the encoding network, adjusting the network parameters and feature projection parameters of the encoding network based on the contrastive loss can help improve the training accuracy of the encoding network.
[0081] In some disclosed embodiments, the encoding network includes a depth feature extraction sub-network and a context encoding sub-network connected in sequence. The depth feature extraction sub-network is configured to extract frame-level depth speech features, and the context encoding sub-network is configured to perform context encoding.
[0082] Therefore, the encoding network includes a depth feature extraction sub-network and a context encoding sub-network connected in sequence. The depth feature extraction sub-network is configured to extract frame-level depth speech features, and the context encoding sub-network is configured to perform context encoding. Therefore, the network structure of the encoding network can be simplified as much as possible, so that in the subsequent application process of speech recognition, it only needs to be spliced with the decoding network to be directly applied, which can help with engineering implementation and the promotion and implementation of the solution.
[0083] In some disclosed embodiments, the contrastive loss is obtained by weighting a first loss and a second loss; wherein, the first loss is obtained by comparing the frame-level noisy speech features and the frame-level first quantization features located at the same first time sequence, and the second loss is obtained by comparing the frame-level noisy speech features and the frame-level first quantization features located at the same second time sequence, and the first time sequence is the time sequence where the masked speech frames in the first noisy speech of the sample are located, and the second time sequence is the time sequence where the unmasked speech frames in the first noisy speech of the sample are located.
[0084] Therefore, a contrast loss is obtained by weighting a first loss and a second loss. The first loss is obtained by comparing frame-level noisy speech features and frame-level first quantization features at the same first time sequence. The second loss is obtained by comparing frame-level noisy speech features and frame-level first quantization features at the same second time sequence. The first time sequence is the time sequence where the masked speech frames are located in the first noisy speech sample, and the second time sequence is the time sequence where the unmasked speech frames are located in the first noisy speech sample. That is, the first loss of the masked part and the second loss of the unmasked part are obtained by the same calculation method. Therefore, it is possible to distinguish the masked part and the unmasked part to improve the accuracy of the contrast loss while minimizing the complexity of the loss metric as much as possible.
[0085] In some disclosed embodiments, the frame-level first quantization features are obtained by clustering and quantizing the frame-level first speech features of the first clean speech sample by a clustering model. Before each round of training of the encoding network, the clustering model is pre-trained based on the frame-level second speech features of the second clean speech sample.
[0086] Therefore, the frame-level first quantization features are obtained by clustering and quantizing the frame-level first speech features of the first clean speech sample by a clustering model. Before each round of training of the encoding network, the clustering model is pre-trained based on the frame-level second speech features of the second clean speech sample. Therefore, it is possible to assist the encoding network training by the clustering model, which helps to improve the performance of the encoding network during the training process.
[0087] In some disclosed embodiments, before the first round of training of the encoding network, the frame-level second speech features are the frame-level acoustic features of the second clean speech sample. During the first round of training of the encoding network, the frame-level first speech features are the frame-level acoustic features of the first clean speech sample; and / or, after the first round of training of the encoding network, the frame-level first speech features are extracted from the first clean speech sample by the newly trained encoding network, and the frame-level second speech features are extracted from the second clean speech sample by the newly trained encoding network.
[0088] Therefore, before the first round of training of the encoding network, the frame-level second speech features are the frame-level acoustic features of the second clean speech sample. During the first round of training of the encoding network, the frame-level first speech features are the frame-level acoustic features of the first clean speech sample, which can guide the encoding network to accurately segment the pronunciation unit boundaries as much as possible during the training stage and improve the fineness of the acoustic unit modeling of the encoding network; after the first round of training of the encoding network, the frame-level first speech features are extracted from the first clean speech sample by the newly trained encoding network, and the frame-level second speech features are extracted from the second clean speech sample by the newly trained encoding network, which can further optimize the network performance of the encoding network by combining multiple rounds of iteration of the encoding network after guiding the encoding network to learn the pronunciation unit boundaries.
[0089] In some disclosed embodiments, the speech recognition device 50 includes a feature clustering module for processing frame-level first speech features based on a clustering model to obtain a number of clustering sets; the speech recognition device 50 includes a feature statistics module for performing feature statistics on the frame-level first speech features located in the same clustering set to obtain the statistical speech features of the clustering set; the speech recognition device 50 includes a feature selection module for using the statistical speech features of the clustering set as the frame-level first quantization features of the speech frames to which the respective frame-level first speech features in the clustering set belong.
[0090] Therefore, by processing the frame-level first speech features based on the clustering model to obtain a number of clustering sets, performing feature statistics on the frame-level first speech features located in the same clustering set to obtain the statistical speech features of the clustering set, and then using the statistical speech features of the clustering set as the frame-level first quantization features of the speech frames to which the respective frame-level first speech features in the clustering set belong, the accuracy of the frame-level first quantization features can be improved.
[0091] Please refer to Figure 6 , Figure 6 which is a schematic framework diagram of an embodiment of the electronic device 60 of the present application. The electronic device 60 includes a memory 61 and a processor 62 that are coupled to each other. Program instructions are stored in the memory 61, and the processor 62 is configured to execute the program instructions to implement the steps in any of the above-described speech recognition method embodiments. Specifically, the electronic device 60 may include, but is not limited to: a server, a desktop computer, a laptop computer, a tablet computer, a smart phone, etc., which are not limited herein.
[0092] Specifically, the processor 62 is configured to control itself and the memory 61 to implement the steps in any of the above-described speech recognition method embodiments. The processor 62 may also be referred to as a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 62 may be implemented jointly by integrated circuit chips.
[0093] In the above solution, on the one hand, the training of the encoding network can be completed through unsupervised training by comparing speech features, while the decoding network is trained in a supervised manner after the encoding network converges. Therefore, the demand for labeled data can be minimized as much as possible, thus adapting to low-resource scenarios. On the other hand, by constructing a contrast prediction task between the first noisy speech sample and the first clean speech sample, the encoding network can learn to extract effective speech representations from the noisy speech as much as possible for the clean speech, thus adapting to low signal-to-noise ratio scenarios. Therefore, by performing an unsupervised contrast prediction task through the encoding network and a supervised recognition prediction task through the decoding network, the speech recognition performance can be improved in low signal-to-noise ratio and low-resource scenarios.
[0094] Please refer to Figure 7 , Figure 7 which is a schematic framework diagram of an embodiment of the computer-readable storage medium 70 of the present application. The computer-readable storage medium 70 stores program instructions 71 that can be run by a processor, and the program instructions 71 are used to implement the steps in any of the above speech recognition method embodiments.
[0095] In the above solution, on the one hand, the training of the encoding network can be completed through unsupervised training by comparing speech features, while the decoding network is trained in a supervised manner after the encoding network converges. Therefore, the demand for labeled data can be minimized as much as possible, thus adapting to low-resource scenarios. On the other hand, by constructing a contrast prediction task between the first noisy speech sample and the first clean speech sample, the encoding network can learn to extract effective speech representations from the noisy speech as much as possible for the clean speech, thus adapting to low signal-to-noise ratio scenarios. Therefore, by performing an unsupervised contrast prediction task through the encoding network and a supervised recognition prediction task through the decoding network, the speech recognition performance can be improved in low signal-to-noise ratio and low-resource scenarios.
[0096] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0097] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.
[0098] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of apparatuses or units can be in electrical, mechanical, or other forms.
[0099] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0100] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0101] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0102] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
Claims
1. A speech recognition method, characterized in that, Including: Obtain the speech to be recognized; Recognize the speech to be recognized based on a speech recognition model to obtain a recognition text; Wherein, the speech recognition model includes an encoding network and a decoding network. The encoding network is trained based on a contrast loss between frame-level first quantization features obtained by feature clustering and quantization of a sample first clean speech and frame-level noisy speech features of a sample first noisy speech. The sample first noisy speech is obtained by adding noise to the sample first clean speech. The decoding network is supervised-trained based on a sample second noisy speech after the encoding network converges. The steps for obtaining the contrast loss include: Extract the frame-level deep speech features of the sample first noisy speech; Based on the frame-level deep speech features, perform context encoding under the condition of masking several speech frames in the sample first noisy speech to obtain frame-level noisy speech features of each of the speech frames of the sample first noisy speech; Based on feature projection parameters corresponding to a clustering set to which the frame-level first quantization features located at the same time sequence as the frame-level noisy speech features belong, perform feature projection on the frame-level noisy speech features to obtain frame-level noisy projection features of the frame-level noisy speech features; Based on the feature similarity between the frame-level first quantization features located at the same time sequence as the frame-level noisy speech features and the frame-level noisy projection features of the frame-level noisy speech features, obtain the contrast loss. And during the training process of the encoding network, based on the contrast loss, adjust the network parameters of the encoding network and the feature projection parameters.
2. The method according to claim 1, characterized in that, The encoding network includes a depth feature extraction sub-network and a context encoding sub-network connected in sequence. The depth feature extraction sub-network is used to extract the frame-level deep speech features, and the context encoding sub-network is used to perform the context encoding.
3. The method according to claim 1 or 2, characterized in that, The contrast loss is obtained by weighting a first loss and a second loss; Wherein, the first loss is obtained by comparing the frame-level noisy speech features and the frame-level first quantization features located at the same first time sequence, and the second loss is obtained by comparing the frame-level noisy speech features and the frame-level first quantization features located at the same second time sequence. The first time sequence is the time sequence where the masked speech frames in the sample first noisy speech are located, and the second time sequence is the time sequence where the unmasked speech frames in the sample first noisy speech are located.
4. The method according to claim 1, characterized in that The frame-level first quantization features are obtained by clustering and quantizing the frame-level first speech features of the sample first clean speech by a clustering model. And before each round of training of the encoding network, the clustering model is pre-trained based on the frame-level second speech features of a sample second clean speech.
5. The method according to claim 4, wherein Before the first round of training of the encoding network, the frame-level second speech features are the frame-level acoustic features of the sample second clean speech, and during the first round of training of the encoding network, the frame-level first speech features are the frame-level acoustic features of the sample first clean speech; And / or, after the first round of training of the encoding network, the frame-level first speech feature is extracted from the sample first clean speech by the latest trained encoding network, and the frame-level second speech feature is extracted from the sample second clean speech by the latest trained encoding network.
6. The method according to claim 4, characterized in that, The obtaining step of the frame-level first quantization feature includes: Processing the frame-level first speech feature based on the clustering model to obtain a plurality of clustering sets; Performing feature statistics on each of the frame-level first speech features located in the same clustering set to obtain the statistical speech feature of the clustering set; Using the statistical speech feature of the clustering set as the frame-level first quantization feature of the speech frame to which each of the frame-level first speech features in the clustering set belongs.
7. A voice recognition device, characterized in that, Including: An obtaining module, configured to obtain the speech to be recognized; A recognition module, configured to recognize the speech to be recognized based on a speech recognition model to obtain a recognition text; Wherein, the speech recognition model includes an encoding network and a decoding network. The encoding network is trained based on a contrast loss between the frame-level first quantization feature after quantization of the sample first clean speech and the frame-level noisy speech feature of the sample first noisy speech. The sample first noisy speech is obtained by adding noise to the sample first clean speech. The decoding network is trained based on the sample second noisy speech labeled with the sample recognition text after the encoding network converges. The obtaining step of the contrast loss includes: Extracting the frame-level deep speech feature of the sample first noisy speech; Performing context encoding based on the frame-level deep speech feature in the case of masking a plurality of speech frames in the sample first noisy speech to obtain the frame-level noisy speech feature of each speech frame of the sample first noisy speech; Performing feature projection on the frame-level noisy speech feature based on the feature projection parameter corresponding to the clustering set to which the frame-level first quantization feature located in the same time sequence as the frame-level noisy speech feature belongs, to obtain the frame-level noisy projection feature of the frame-level noisy speech feature; Obtaining the contrast loss based on the feature similarity between the frame-level first quantization feature located in the same time sequence as the frame-level noisy speech feature and the frame-level noisy projection feature of the frame-level noisy speech feature, and during the training process of the encoding network, adjusting the network parameters of the encoding network and the feature projection parameter based on the contrast loss.
8. An electronic device, characterized in that, Including a memory and a processor that are coupled to each other. Program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the speech recognition method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, Stored with program instructions that can be run by a processor, the program instructions are used to implement the speech recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Speech recognition model training method and speech recognition method and device
CN112086087A