Speech recognition model training method and device, electronic device and storage medium
By constructing the first and second loss functions, combining the feature comparison of original speech and noise-added speech, the speech recognition model is trained, and the problem of inaccurate speech feature recognition in noise scenarios is solved, and the recognition accuracy of the model in a noisy environment is improved.
Patent Information
- Application Number
- CN202310443829.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-23
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-04-23
AI Technical Summary
Existing speech recognition models are difficult to accurately distinguish confusing speech features in noise scenarios, resulting in low recognition accuracy and affecting user experience.
By simultaneously inputting the original speech and noise-added speech to the encoder, the target acoustic features are obtained, and the target tag sequence is decoded by the decoder to build a first loss function, and the second loss function is constructed, while comparing the original speech and the reconstructed acoustic features of the noise-added speech of the same timing to update the parameters of the speech recognition model.
The recognition accuracy of the speech recognition model in noise scenarios is improved, and the recognition ability of speech features is enhanced.
Smart Images

Figure CN116343770B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a method for training a speech recognition model, a training device for a speech recognition model, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the development of artificial intelligence (AI) technology, the number of scenarios where voice recognition is used to provide user services is increasing, such as intelligent voice assistants and speech-to-text conversion. Consequently, numerous methods for training voice recognition models have emerged. However, in related technologies, voice recognition models cannot accurately distinguish easily confused speech features in noisy environments. This results in low recognition accuracy, which impacts the user experience. Summary of the Invention
[0003] The purpose of the present disclosure is to provide a speech recognition model training method, a speech recognition model training device, an electronic device and a computer-readable storage medium, thereby improving the precision of the speech recognition model and the accuracy of the speech recognition results at least to a certain extent.
[0004] According to a first aspect of the present disclosure, a training method for a speech recognition model is provided, the speech recognition model includes an encoder and a decoder, and the training method includes: inputting speech data into the encoder to obtain target acoustic features; the speech data includes original speech and noisy speech corresponding to the original speech; using the decoder to decode the target acoustic features to obtain a target label sequence, and constructing a first loss function based on the speech data and the target label sequence; reconstructing the target acoustic features to obtain reconstructed acoustic features, and constructing a second loss function by comparing the reconstructed acoustic features of the original speech and the reconstructed acoustic features of the corresponding noisy speech at the same time sequence; updating the parameters of the speech recognition model based on the first loss function and the second loss function to obtain a trained speech recognition model.
[0005] According to a second aspect of the present disclosure, a training device for a speech recognition model is provided, the speech recognition model includes an encoder and a decoder, and the training device includes: a first data processing module, used to input speech data into the encoder to obtain target acoustic features; the speech data includes original speech and noisy speech corresponding to the original speech; a second data processing module, used to use the decoder to decode the target acoustic features to obtain a target label sequence, and construct a first loss function based on the speech data and the target label sequence; a contrastive learning module, used to reconstruct the target acoustic features to obtain reconstructed acoustic features, and construct a second loss function by comparing the reconstructed acoustic features of the original speech and the reconstructed acoustic features of the corresponding noisy speech at the same time sequence; a parameter updating module, used to update the parameters of the speech recognition model based on the first loss function and the second loss function to obtain a trained speech recognition model.
[0006] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing one or more programs, which, when executed by one or more processors, enables the one or more processors to implement the above method.
[0007] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented.
[0008] The technical solution involved in the embodiments of the present disclosure simultaneously inputs the original speech and the noisy speech corresponding to the original speech into the encoder to obtain the target acoustic features, and decodes the target acoustic features through the decoder to obtain the target label sequence, so as to construct a first loss function based on the speech data and the target label sequence, and construct a second loss function by comparing the reconstructed acoustic features of the original speech and the reconstructed acoustic features of the loaded speech at the same time sequence, so that the model can not only update the parameters according to the loss generated by the speech recognition task, but also obtain the contrast loss (i.e., the second loss function) by comparing the speech features of the original speech and the noisy speech at the same time sequence, so that the model can better learn the correlation between the original speech and the loaded speech in the same speech features, as well as the difference between different speech features, thereby improving the model's recognition accuracy of easily confused speech features in noisy scenarios.
[0009] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0011] Figure 1 A schematic diagram showing the stages involved in the technical solution of the embodiment of the present disclosure is shown;
[0012] Figure 2 A schematic diagram schematically illustrates the structure of a speech recognition model according to an exemplary embodiment of the present disclosure;
[0013] Figure 3 A flowchart schematically illustrates a method for training a speech recognition model according to an exemplary embodiment of the present disclosure;
[0014] Figure 4A schematic diagram schematically illustrates a method for determining positive and negative samples according to an exemplary embodiment of the present disclosure;
[0015] Figure 5 A schematic diagram schematically illustrates a comparative network structure in an exemplary embodiment of the present disclosure;
[0016] Figure 6 A flowchart of constructing a second loss function using the obtained reconstructed acoustic features in an exemplary embodiment of the present disclosure is schematically shown;
[0017] Figure 7 A schematic diagram schematically illustrates a method for determining positive sample pairs and negative sample pairs according to an exemplary embodiment of the present disclosure;
[0018] Figure 8 A flowchart of constructing a second loss function using the obtained reconstructed acoustic features in an exemplary embodiment of the present disclosure is schematically shown;
[0019] Figure 9 Schematically illustrates a flow chart of constructing a first loss function based on supervised speech data and a target label sequence in an exemplary embodiment of the present disclosure;
[0020] Figure 10 A schematic diagram schematically illustrates the composition of a training device for a speech recognition model in an exemplary embodiment of the present disclosure;
[0021] Figure 11 A schematic diagram of an electronic device to which the embodiments of the present disclosure can be applied is shown. DETAILED DESCRIPTION
[0022] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0023] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0024] Figure 1Schematic diagram showing the stages involved in the technical solution of the embodiment of the present disclosure, such as Figure 1 As shown, the technical solution of the embodiment of the present disclosure includes a model training stage and a model application stage.
[0025] In the embodiment of the present disclosure, the provided speech recognition model training method can be executed by a terminal device. In the manner of execution by the terminal device, each step of the speech recognition model training method in the embodiment of the present disclosure is executed by the terminal device.
[0026] For example, the processor of the terminal device performs the following steps: inputting the voice data into the encoder to obtain the target acoustic features; the voice data includes the original voice and the noisy voice corresponding to the original voice; using the decoder to decode the target acoustic features to obtain the target label sequence, and constructing a first loss function based on the voice data and the target label sequence; reconstructing the target acoustic features to obtain reconstructed acoustic features, and constructing a second loss function by comparing the reconstructed acoustic features of the original voice and the reconstructed acoustic features of the corresponding noisy voice at the same time sequence; updating the parameters of the voice recognition model based on the first loss function and the second loss function to obtain a trained voice recognition model.
[0027] The terminal device can be an intelligent device with data processing capabilities, such as a smart phone, computer, tablet computer, vehicle-mounted device, wearable device, monitoring device and other intelligent devices. The terminal device can also be called a mobile terminal, terminal, mobile device, etc. The present disclosure does not limit the type of terminal device.
[0028] Furthermore, the training method of the speech recognition model provided in the embodiment of the present disclosure can also be executed by a server. Correspondingly, in this way of execution by the server, the server can start executing the steps in the training method of the speech recognition model of the embodiment of the present disclosure in response to a trigger command, wherein the trigger command can be sent by a terminal device used by the user, or can be triggered locally by the server in response to some automated events. The server can be a background system that provides relevant services in the embodiment of the present disclosure, and can include a portable computer, desktop computer, smart phone, and other electronic devices with computing functions, or a cluster formed by multiple electronic devices.
[0029] In addition, the technical solution of the embodiment of the present disclosure can also be executed collaboratively by the terminal device and the server. In this way of collaborative execution by the electronic device and the server, some steps in the technical solution provided by the embodiment of the present disclosure are executed by the terminal device, while other steps are executed by the server. For example, the technical solution provided by the embodiment of the present disclosure can be obtained by the terminal device. Voice data (including original voice and corresponding noisy voice) is sent to the server, and the server executes the training phase of the model, and sends the trained speech recognition model to the terminal device, and the terminal device executes the model application phase.
[0030] It should be noted that in this method of collaborative execution by the terminal device and the server, the steps executed by the terminal device and the server respectively can be dynamically adjusted according to actual conditions, and there is no special restriction on this.
[0031] In the embodiments of the present disclosure, the technical solution is described by taking the execution of the technical solution by the terminal device as an example.
[0032] The speech recognition model of the related technology uses the entire target sequence to calculate the loss of the entire sentence. However, in noisy scenarios, whether in end-to-end systems or hybrid systems, the model cannot accurately distinguish easily confused speech features. For example, the same phoneme or word is difficult to distinguish in clean speech and noisy speech, which in turn affects the accuracy of the speech recognition results.
[0033] Based on one or more of the above problems, an exemplary embodiment of the present disclosure first provides a training method for a speech recognition model. While constructing a first loss function based on speech data and a target label sequence, a second loss function is constructed by comparing the reconstructed acoustic features of the original speech and the reconstructed acoustic features of the noisy speech of the same time sequence, so that the model can better learn the correlation between the original speech and the loaded speech in the same speech features, as well as the differences between different speech features, thereby improving the accuracy of the model's speech recognition in noisy scenarios.
[0034] In the disclosed embodiments, the trained speech recognition model can be applied in an end-to-end manner to various speech recognition task-related scenarios. For example, scenarios such as transcribing voice files, generating subtitles in real-time during live broadcasts, recording courtroom conversations in real-time, and enabling intelligent voice assistants to interact with smart devices via voice.
[0035] Furthermore, the technical solutions of the embodiments of the present disclosure can be applied to various speech recognition frameworks, including but not limited to wenet, kaldi2, espnet, wav2etter++, etc.
[0036] refer to Figure 2As shown, the speech recognition model of the embodiment of the present disclosure is an attention encoder-decoder framework based on continuous time classification (CTC) (full name: Connectionist temporal classification), including an encoder and a decoder. The speech recognition model can include ASR (Automatic Speech Recognition) subtasks and comparison subtasks, corresponding to classification network structures and comparison network structures, respectively, to perform speech recognition tasks and feature comparison learning tasks.
[0037] The encoder can be understood as a shared network. The encoder's output is used for the speech recognition subtask, feature comparison learning, and as acoustic features for the decoder. The speech recognition subtask can be based on linear layers and softmax layers to perform ASR tasks. This embodiment of the disclosure does not impose any specific restrictions on the classification network structure.
[0038] Among them, the encoder can be selected from encoder structures such as Conformer, LSTM, Transformer, etc., and the embodiments of the present disclosure do not make special limitations on this.
[0039] See also Figure 3 As shown, the training method of the speech recognition model of the embodiment of the present disclosure may include the following steps S310 to S340:
[0040] In step S310, speech data is input into an encoder to obtain target acoustic features; the speech data includes original speech and noisy speech corresponding to the original speech.
[0041] In the exemplary embodiments of the present disclosure, speech data refers to speech data that has been manually labeled as a reference. This manual label is described as a true label sequence in the subsequent description. The original speech in the speech data is clean speech, and the noisy speech in the speech data is obtained by adding noise to the original speech (clean speech). The true label of the noisy speech in the speech data is the true label of the corresponding clean speech (original speech). This true label is determined based on manual annotation and speech and alignment information.
[0042] Among them, the speech data can be acoustic features such as Fbank features and MFCC (Mel-scale Frequency Cepstral Coefficients) features, which are used as the input of the encoder. That is, the speech data includes the acoustic features of the original speech and the acoustic features of the corresponding noisy speech.
[0043] In step S320, the target acoustic features are decoded using a decoder to obtain a target label sequence, and a first loss function is constructed based on the speech data and the target label sequence.
[0044] In an exemplary embodiment of the present disclosure, the first loss function is based on the loss of the ASR subtask. The target acoustic features output by the encoder can be used as the acoustic features of the decoder. The target acoustic features are decoded by the decoder to obtain a target label sequence to construct the first loss function based on the speech data and the target label sequence.
[0045] In step S330, the target acoustic feature is reconstructed to obtain a reconstructed acoustic feature, and a second loss function is constructed by comparing the reconstructed acoustic feature of the original speech and the reconstructed acoustic feature of the corresponding noisy speech at the same time sequence.
[0046] In an exemplary embodiment of the present disclosure, the second loss function is a supervised contrast loss based on the contrast subtask. Reconstructing acoustic features refers to further processing the target acoustic features using a preset network structure to obtain more distinct contrast features. The preset network structure may be a contrast network structure.
[0047] The embodiment of the present disclosure enables the model to learn the correlation between the same speech features of clean speech and noisy speech, and to learn the differences between different speech features of clean speech and noisy speech. A second loss function is constructed by comparing the reconstructed acoustic features of the original speech and the corresponding reconstructed acoustic features of the noisy speech at the same time sequence.
[0048] In step S340, the parameters of the speech recognition model are updated based on the first loss function and the second loss function to obtain a trained speech recognition model.
[0049] In an exemplary embodiment of the present disclosure, the parameters in the speech recognition model may be updated in combination with the first loss function and the second loss function to obtain a trained speech recognition model.
[0050] The training method of the speech recognition model based on the embodiment of the present disclosure allows the model to update parameters according to the loss generated by the speech recognition task, and obtain the contrast loss (i.e., the second loss function) by comparing the speech features of the original speech and the noisy speech at the same time sequence, so that the model can better learn the correlation between the original speech and the loaded speech in the same speech features, as well as the differences between different speech features, thereby improving the model's recognition accuracy for easily confused speech features in noisy scenarios.
[0051] In an exemplary embodiment, before inputting the speech data into the encoder to obtain the target acoustic features, a batch set can be determined from the sample speech data, where the batch set includes speech frames corresponding to the original speech and speech frames corresponding to the noisy speech, and positive samples and negative samples of the speech frames corresponding to the speech data in the batch set are determined.
[0052] The sample speech data is a set of all training samples used for model training. A portion of the samples in the set is used to perform a back propagation parameter update on the model, and the portion of samples is a determined batch set (batch).
[0053] Exemplarily, the sample speech data includes all speech frames, and some speech frames (speech frames of original speech and speech frames of noisy speech of original speech) can be selected from the sample speech data to form a batch set for iterative training of the model. In the embodiment of the present disclosure, all speech frames of original speech and speech frames of noisy speech of original speech in the batch set are compared.
[0054] Based on this, in an exemplary embodiment, a method for determining positive samples and negative samples of speech frames corresponding to speech data is also provided, which may include:
[0055] First, for a batch set, the speech frames corresponding to the speech data in the batch set are aligned with each other to obtain feature alignment information. Then, based on the feature alignment information, the speech frames with the same speech features in the batch set are determined as positive examples, and the speech frames with different speech features in the batch set are determined as negative samples.
[0056] The speech feature may be one or more of phonemes, words, etc., and the embodiments of the present disclosure do not impose any special limitation on this.
[0057] Exemplarily, a positive example of a speech frame corresponding to speech data refers to a speech frame having the same speech features as the speech frame, whereas a negative example refers to a speech frame having different speech features from the speech frame. The speech features may be, for example, phonemes, i.e., a speech frame having the same phonemes as a speech frame is a positive example of the speech frame, while a speech frame having different phonemes from the speech frame is a negative example of the speech frame.
[0058] Taking the speech feature as phoneme as an example, if the speech frame corresponding to the speech data includes X={x1,x2,x3,……,x n}, then according to whether the phonemes of the speech frames are the same or not, determine the positive and negative samples of the speech frame x1, determine the positive and negative samples of x2, and so on, determine x n Positive samples and negative samples. Accordingly, each speech frame and its positive sample can form a positive sample pair, and each speech frame and its negative sample can form a negative sample pair. Figure 4 , x2 and x1 constitute a positive sample pair, x n and x1 form a negative sample pair, and x n It forms a negative sample pair with x2, and other speech frames are not listed one by one.
[0059] The embodiment of the present disclosure determines positive samples and negative samples by forcibly aligning the speech features of the speech frames, so that subsequent solutions can compare the original speech and the noisy speech in the same sequence based on the positive samples and the negative samples.
[0060] In an exemplary embodiment, the speech recognition model further includes a contrast head structure, see Figure 5 A schematic diagram of a comparison network structure of an exemplary embodiment of the present disclosure is shown, where the comparison network structure includes a first linear layer, an activation layer, a dropout layer, and a second linear layer.
[0061] Based on the contrast network structure, for a batch set, reconstructing the target acoustic feature to obtain the reconstructed acoustic feature includes: inputting the target acoustic feature into the contrast network structure to be processed sequentially through the first linear layer, the activation layer, the dropout layer, and the second linear layer, and obtaining the acoustic feature output by the second linear layer as the reconstructed acoustic feature.
[0062] Specifically, the reconstructed acoustic features output by the contrast network structure can be obtained through formula (1):
[0063] O=w2*Dropout((w1*x)*sigmoid(w1*x)) (1)
[0064] Among them, O is the reconstructed acoustic feature output by the contrast network structure, x is the target acoustic feature output by the encoder, x∈(B,T,C), B is the size of the batch set, T is the number of speech frames in the batch set, and C is the dimension of each speech frame. The embodiment of the present disclosure adjusts the output of the contrast network structure from O∈(B,T,C) to O∈(B*T,C) to calculate the contrast loss of a batch set; w1 and w2 are weight parameters respectively.
[0065] Based on the contrastive network structure of the embodiment of the present disclosure, two linear layers with activation layers are used to project the output of the encoder to facilitate subsequent contrastive learning of features based on reconstructed acoustic features.
[0066] It should be noted that Figure 5 The model parameters in the comparative network structure shown are only exemplary and can be adaptively adjusted according to actual needs, and there is no special limitation on this.
[0067] In an exemplary embodiment, the second loss function includes an Enhanced Supervised Contrastive Learning (ESCL) loss function. Figure 6 As shown, constructing the second loss function by comparing the reconstructed acoustic features of the original speech and the reconstructed acoustic features of the corresponding noisy speech at the same time sequence may include steps S610 and S620:
[0068] Step S610: For the batch set, based on the reconstructed acoustic features, obtain the similarity between each speech frame and the positive sample and the negative sample of the speech frame in the batch set, and determine the contrast loss of the speech frame according to the similarity.
[0069] For each speech frame in the batch set, the similarity between the speech frame and its positive sample can be obtained, and the sum of all similarities can be obtained to obtain the sum of the first similarities, and the sum of the second similarities between the speech frame and other speech frames in the batch set except the speech frame can be obtained. Finally, the contrast loss of the speech frame is constructed according to the sum of the first similarities and the sum of the second similarities.
[0070] Among them, each original speech frame is compared with each speech frame in the batch set, including the speech frame of the noisy speech corresponding to the speech frame, and the speech frame is compared with the speech frame of the noisy speech corresponding to the time sequence to learn the correlation between the same speech features of the clean speech and the noisy speech, as well as the difference between the different speech features of the clean speech and the noisy speech.
[0071] Step S620: constructing an enhanced supervised contrast loss function based on the contrast loss of each speech frame in the batch set.
[0072] After obtaining the contrastive loss of each speech frame in the batch set, an enhanced supervised contrastive loss function can be constructed according to the contrastive loss of each speech frame in the batch set (including speech frames of original speech and noisy speech).
[0073] The disclosed embodiment compares all information in the batch set simultaneously, comparing each speech frame with all its corresponding positive and negative samples. Figure 7 As shown, the batch set may include multiple subsets (such as batch1, batch2 and batch3, etc.), and the "l" at the first position and the "l" at the second to last position may be regarded as positive samples, and the "l" at the second to last position and the "iu" at the last position and the "ao" at the last position may be regarded as negative samples. It should be noted that the embodiment of the present disclosure compares the information of all speech frames (original speech and noisy speech frames) in the batch set. Figure 7 Other positive and negative examples are not fully shown.
[0074] Specifically, the enhanced supervision contrast loss function can be constructed through formula (2):
[0075]
[0076] Among them, Loss esclTo enhance the supervised contrast loss function, S is the length of the contrast network structure output, S = B*T, B is the size of the batch set, T is the number of speech frames in the batch set, O is the reconstructed acoustic feature (the output of the contrast network structure), J is the set of positive samples corresponding to the i-th speech frame, K is the set of other speech frames in the batch set except the i-th speech frame, τ is a hyperparameter for regulating the ratio, τ can be set to 0.07, of course, it can also be set according to actual needs, and there is no special limitation on this.
[0077] The similarity between two samples can be obtained by formula (3):
[0078]
[0079] Here, sim(·,·) represents the similarity between two samples, and α and β are the reconstructed acoustic features O corresponding to different samples.
[0080] The enhanced supervised contrast loss function of the embodiment of the present disclosure can enhance the robustness of feature representation by bringing the representation of positive samples closer and moving the representation of negative samples further away. Since this process includes the comparison between the positive and negative samples of the speech frames of the original speech and the speech frames of the noisy speech, the model can also better learn the correlation between the same speech features of the original speech and the loaded speech, as well as the differences between different speech features.
[0081] In an exemplary embodiment, the second loss function further includes a Sequential Supervised Contrastive Learning (SSCL) loss function.
[0082] like Figure 8 As shown, constructing the second loss function by comparing the reconstructed acoustic features of the original speech and the reconstructed acoustic features of the corresponding noisy speech at the same time sequence may include steps S810 to S840:
[0083] Step S810: for the batch set, determine positive sample pairs and negative sample pairs according to the positive sample and negative sample of each speech frame.
[0084] For each speech frame in the batch set (including speech frames of original speech and speech frames of noisy speech), a positive sample pair is determined based on the speech frame and the positive sample of the speech frame, and a negative sample pair is determined based on the speech frame and the negative sample of the speech frame. Then, the positive sample pair required for this step is obtained based on the positive sample pairs of all speech frames, and the negative sample pair required for this step is obtained based on the negative sample pairs of all speech frames.
[0085] Step S820: Obtain the sum of the third similarities of all positive sample pairs in the batch set.
[0086] Based on the reconstructed speech features output by the contrast network structure, the sum of the third similarities of all positive sample pairs in the batch set is obtained.
[0087] Step S830: Obtain the sum of the fourth similarities of all negative example pairs in the batch set.
[0088] The reconstructed speech features output by the contrast network structure are also used to obtain the sum of the fourth similarities of all negative sample pairs in the batch set.
[0089] Step S840: constructing a sequence supervision comparison loss function according to the sum of the third similarities and the sum of the fourth similarities.
[0090] Specifically, the sequence supervision contrast loss function can be constructed by formula (4):
[0091]
[0092] Among them, P is all positive sample pairs, N is all negative sample pairs, i and j are the identifiers corresponding to all positive sample pairs output by the comparison network structure, k and w are the identifiers corresponding to all negative sample pairs output by the corresponding network structure. The similarity between two samples and the method of obtaining τ can be found in the section on constructing the enhanced supervision contrast loss function, which will not be repeated here.
[0093] By constructing a sequential supervised contrast loss function and simultaneously comparing all positive and negative samples corresponding to the original speech and noisy speech in the batch set, the speech recognition model pays more attention to alignment information, improving the targeted feature comparison while enabling the speech recognition model to learn the correlation between the same speech features of clean speech and noisy speech, as well as the differences between different speech features.
[0094] In an exemplary embodiment, constructing a second loss function based on the obtained enhanced supervised contrast loss function and the sequence supervised contrast loss function includes:
[0095] First, based on the weights pre-assigned to the enhanced supervised contrast loss function and the sequence supervised contrast loss function respectively, the enhanced supervised contrast loss function and the sequence supervised contrast loss function are weighted; then, the weighted first joint loss function is used as the second loss function.
[0096] Specifically, the second loss function can be constructed by formula (5):
[0097] Loss contrast =βLoss sscl +(1-β)Loss escl (5)
[0098] Among them, Losscontrast is the second loss function, Loss ESCL To enhance the supervision contrast loss function, Loss SSCL is the sequence supervised contrast loss function, β and (1-β) are the weights pre-assigned to the enhanced supervised contrast loss function and the sequence supervised contrast loss function respectively. β is used as a hyperparameter to balance the two supervised contrast loss functions.
[0099] In an exemplary embodiment, the first loss function includes a continuous-time classification loss function and an attention loss function. The continuous-time classification loss function is a loss function that describes the loss between sequences based on the assumption of conditional independence. The continuous-time classification loss function describes the loss incurred during the process of aligning supervised speech data to target label sequences. The attention loss function is used to describe the loss incurred by the attention mechanism in speech recognition models.
[0100] like Figure 9 As shown, constructing a first loss function based on speech data and target label sequence includes steps S910 to S930:
[0101] Step S910: Taking the real label sequence of the speech data as a reference, determining the continuous-time classification loss function according to the alignment result of the target label sequence and the real label sequence.
[0102] The continuous-time classification loss function shares the encoder in the speech recognition model with the decoder in the speech recognition model, and the continuous-time classification loss function is used to make conditional independence assumptions on the various factors output by the decoder in the speech recognition model.
[0103] For example, the actual label sequence of the speech data is used as a reference, and the loss of aligning the speech data to the target label sequence is independently evaluated according to the forward-backward algorithm conditions. Specifically, the continuous-time classification loss function can be constructed by formula (6):
[0104]
[0105] The true label sequence, y is the target label sequence.
[0106] It should be noted that the continuous-time classification loss function can also be constructed based on the speech data and the target label sequence by other methods, not limited to the method shown in formula (6). The embodiment of the present disclosure does not limit the specific method of constructing the continuous-time classification loss function.
[0107] Step S920: For each true label, determine the attention loss function according to the relationship between the true label and each true label that precedes the true label in the true label sequence.
[0108] Taking the true label sequence as a reference, according to the relationship between the true label in the true label sequence and the true labels before the true label, the loss of the entire speech recognition model due to the attention mechanism is evaluated to obtain the attention loss function.
[0109] For example, the attention loss function can be constructed by formula (7):
[0110]
[0111] True labels, y* [1:n-1] It is y* n The first (n-1) true labels, x is the true label sequence of supervised speech data.
[0112] It should be noted that the attention loss function can also be constructed based on the speech data and the real label sequence in other ways, not limited to the method shown in formula (7). The embodiment of the present disclosure does not limit the specific method of constructing the attention loss function.
[0113] Step S930: Determine a first loss function based on the continuous-time classification loss function and the attention loss function.
[0114] The continuous time classification loss function and the attention loss function can be weighted based on the weights pre-assigned to the continuous time classification loss function and the attention loss function respectively, and the weighted third joint loss function can be used as the first loss function.
[0115] Specifically, the first loss function can be constructed by formula (8):
[0116] Loss ASR =θLoss CTC +(1-θ)Loss attention (8)
[0117] Among them, Loss ASR is the first loss function, θ is a hyperparameter for balancing the continuous-time classification loss function and the attention loss function, for example, 0.3 or set according to actual needs.
[0118] In an exemplary embodiment, updating the parameters in the speech recognition model based on the first loss function and the second loss function to obtain a trained speech recognition model may include:
[0119] First, based on the weights pre-assigned to the first loss function and the second loss function respectively, the first loss function and the second loss function are weighted, and then the parameters in the speech recognition model are adjusted using the weighted second joint loss function to obtain a trained speech recognition model.
[0120] Specifically, the second joint loss function can be constructed by formula (9):
[0121] Loss=Loss ASR +αLoss contrast (9)
[0122] Among them, Loss is the second joint loss function, and α is the hyperparameter used to balance the first loss function and the second loss function.
[0123] In summary, the embodiment of the present disclosure simultaneously inputs the original speech and the noisy speech into the speech recognition model, obtains the target acoustic features through the encoder, and decodes the target acoustic features through the decoder to obtain the target label sequence, so as to construct a first loss function according to the speech data and the target label sequence, and constructs a second loss function by comparing the reconstructed acoustic features of the original speech and the reconstructed acoustic features of the loaded speech at the same time sequence, so that the model can not only update the parameters according to the loss generated by the speech recognition task, but also obtain the contrast loss (i.e., the second loss function) by comparing the speech features of the original speech and the noisy speech at the same time sequence, so that the model can better learn the correlation between the same speech features of the original speech and the loaded speech, as well as the difference between different speech features, so that the same speech features between the original speech and the noisy speech are clustered more closely, and the different speech features are further away, thereby improving the recognition accuracy of the model for easily confused speech features in noisy scenarios.
[0124] It should be noted that the above figures are merely illustrative of the processes included in the methods according to exemplary embodiments of the present disclosure and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0125] For further reference, Figure 10 As shown, in an exemplary embodiment of the present disclosure, a speech recognition model training device 1000 is provided, comprising a first data processing module 1010, a second data processing module 1020, a data comparison module 1030 and a parameter updating module 1040. In which:
[0126] The first data processing module 1010 is configured to input speech data into an encoder to obtain target acoustic features; the speech data includes original speech and noisy speech corresponding to the original speech;
[0127] A second data processing module 1020 is configured to decode the target acoustic features using a decoder to obtain a target label sequence, and construct a first loss function based on the speech data and the target label sequence;
[0128] A contrastive learning module 1030 is configured to reconstruct the target acoustic feature to obtain a reconstructed acoustic feature, and construct a second loss function by comparing the reconstructed acoustic feature of the original speech and the reconstructed acoustic feature of the corresponding noisy speech at the same time sequence;
[0129] The parameter updating module 1040 is used to update the parameters of the speech recognition model based on the first loss function and the second loss function to obtain a trained speech recognition model.
[0130] In an exemplary embodiment, the original speech is clean speech; the first data processing module 1010 is further configured to execute: performing noise processing on the clean speech to obtain noisy speech; wherein the true label of the noisy speech in the speech data is the true label of the original speech corresponding to the noisy speech.
[0131] In an exemplary embodiment, the first data processing module 1010 is configured to perform: determining a batch set from the sample speech data, the batch set including speech frames corresponding to the original speech and speech frames corresponding to the noisy speech; and determining positive samples and negative samples of the speech frames corresponding to the speech data in the batch set.
[0132] In an exemplary embodiment, the first data processing module 1010 is configured to perform: for a batch set, performing speech feature alignment on the speech frames corresponding to the speech data in the batch set to obtain feature alignment information; based on the feature alignment information, determining the speech frames with the same speech features in the batch set as positive examples, and determining the speech frames with different speech features in the batch set as negative samples.
[0133] In an exemplary embodiment, the speech recognition model also includes a contrastive network structure, which includes a cascaded first linear layer, an activation layer, a discard layer, and a second linear layer; the contrastive learning module 1030 is configured to perform: inputting the target acoustic feature into the contrastive network structure for processing through the first linear layer, the activation layer, the discard layer, and the second linear layer in sequence; and obtaining the acoustic feature output by the second linear layer as the reconstructed acoustic feature.
[0134] In an exemplary embodiment, the second loss function includes an enhanced supervised contrast loss function; the contrast learning module 1030 is configured to perform: for a batch set, based on the reconstructed acoustic features, obtaining the similarity between each speech frame and the positive samples and negative samples of the speech frames in the batch set, and determining the contrast loss of the speech frame according to the similarity; constructing an enhanced supervised contrast loss function according to the contrast loss of each speech frame in the batch set.
[0135] In an exemplary embodiment, the contrastive learning module 1030 is configured to perform: for each speech frame, obtaining the sum of the first similarities between the speech frame and the positive sample of the speech frame; obtaining the sum of the second similarities between the speech frame and other speech frames in the batch set except the speech frame; and constructing the contrast loss of the speech frame based on the sum of the first similarities and the sum of the second similarities.
[0136] In an exemplary embodiment, the second loss function also includes a sequence supervised contrast loss function; the contrast learning module 1030 is configured to perform: for a batch set, determining positive sample pairs and negative sample pairs based on the positive samples and negative samples of each speech frame; obtaining the sum of the third similarities of all positive sample pairs in the batch set; obtaining the sum of the fourth similarities of all negative sample pairs in the batch set; and constructing a sequence supervised contrast loss function based on the sum of the third similarities and the sum of the fourth similarities.
[0137] In an exemplary embodiment, the contrastive learning module 1030 is configured to perform: weighting the enhanced supervised contrastive loss function and the sequence supervised contrastive loss function based on weights pre-assigned to the enhanced supervised contrastive loss function and the sequence supervised contrastive loss function respectively; and using the weighted first joint loss function as the second loss function.
[0138] In an exemplary embodiment, the first loss function includes a continuous-time classification loss function and an attention loss function; the second data processing module 1020 is configured to execute: taking the real label sequence of the speech data as a reference, determining the continuous-time classification loss function according to the alignment result of the target label sequence and the real label sequence; for each real label, determining the attention loss function according to the relationship between the real label and the real labels located in front of the real label in the real label sequence; determining the first loss function based on the continuous-time classification loss function and the attention loss function.
[0139] In an exemplary embodiment, the parameter update module 1040 is configured to perform: weighting the first loss function and the second loss function based on the weights pre-assigned to the first loss function and the second loss function respectively; using the weighted second joint loss function to update the parameters in the speech recognition model to obtain a trained speech recognition model.
[0140] The specific details of each module in the above device have been described in detail in the implementation method part. For details not disclosed, please refer to the implementation method part, and they will not be repeated here.
[0141] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods, or program products. Therefore, various aspects of the present disclosure may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."
[0142] The exemplary embodiments of the present disclosure further provide an electronic device for the above method, which may be the above imaging device or server. Generally, the electronic device includes at least a processor and a memory, the memory being used to store executable instructions of the processor, and the processor being configured to execute the above method by executing the executable instructions.
[0143] Below is Figure 11 Taking the mobile terminal 1100 in the embodiment as an example, the structure of the electronic device in the embodiment of the present disclosure is exemplarily described. It should be understood by those skilled in the art that in addition to the components specifically used for mobile purposes, Figure 11 The structure in FIG. 1 can also be applied to fixed type devices. In other embodiments, the mobile terminal 1100 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware. The interface connection relationship between the components is only shown schematically and does not constitute a structural limitation of the mobile terminal 1100. In other embodiments, the mobile terminal may also adopt the same Figure 11 Different interface connection methods, or a combination of multiple interface connection methods.
[0144] like Figure 11 As shown, the mobile terminal 1100 may specifically include: a processor 1101, a memory 1102, a bus 1103, a mobile communication module 1104, an antenna 1, a wireless communication module 1105, an antenna 2, a display screen 1106, a camera module 1107, an audio module 1108, a power module 1109, and a sensor module 1110.
[0145] The processor 1101 may include one or more processing units, for example: the processor 1101 may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor and / or an NPU (Neural-Network Processing Unit), etc.
[0146] The encoder can encode (i.e., compress) an image or video to reduce the data size for easy storage or transmission. The decoder can decode (i.e., decompress) the encoded data of the image or video to restore the image or video data. The mobile terminal 1100 can support one or more encoders and decoders, such as: image formats such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap), and video formats such as MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).
[0147] The processor 1101 may be connected to the memory 1102 or other components via a bus 1103 .
[0148] Memory 1102 can be used to store computer-executable program code, which includes instructions. Processor 1101 executes various functional applications and data processing of mobile terminal 1100 by running the instructions stored in memory 1102. Memory 1102 can also store application data, such as images, videos, and other files.
[0149] The communication functions of mobile terminal 1100 are implemented through mobile communication module 1104, antenna 1, wireless communication module 1105, antenna 2, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1104 can provide 3G, 4G, and 5G mobile communication solutions for mobile terminal 1100. Wireless communication module 1105 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1100.
[0150] The display screen 1106 is used to implement display functions, such as displaying user interfaces, images, videos, etc., as well as displaying abnormality prompts. The camera module 1107 is used to implement shooting functions, such as capturing images and videos to capture scene images. The audio module 1108 is used to implement audio functions, such as playing audio and capturing voice. The power module 1109 is used to implement power management functions, such as charging the battery, powering the device, and monitoring battery status. The sensor module 1110 may include one or more sensors to implement corresponding sensing and detection functions.
[0151] In addition, the exemplary embodiments of the present disclosure further provide a computer-readable storage medium on which is stored a program product capable of implementing the methods described above in this specification. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product comprising program code. When the program product is executed on a terminal device, the program code is configured to cause the terminal device to execute the steps described in the "Exemplary Methods" section above according to various exemplary embodiments of the present disclosure.
[0152] It should be noted that the computer-readable medium shown in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0153] In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the foregoing.
[0154] In addition, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0155] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow from the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.
Claims
1. A method for training a speech recognition model, characterized in that: The speech recognition model includes an encoder and a decoder, and the training method includes: Inputting speech data into the encoder to obtain target acoustic features; the speech data includes original speech and noisy speech corresponding to the original speech; Decoding the target acoustic features using the decoder to obtain a target label sequence, and constructing a first loss function based on the speech data and the target label sequence; Reconstructing the target acoustic feature to obtain a reconstructed acoustic feature, and constructing a second loss function by comparing the reconstructed acoustic feature of the original speech and the corresponding reconstructed acoustic feature of the noisy speech at the same time sequence; The parameters of the speech recognition model are updated based on the first loss function and the second loss function to obtain a trained speech recognition model.
2. The method according to claim 1, characterized in that The original speech is a clean speech; before inputting the speech data into the encoder to obtain the target acoustic features, the method further includes: Noising the clean speech to obtain the noisy speech; The true label of the noisy speech in the speech data is the true label of the original speech corresponding to the noisy speech.
3. The method according to claim 1, characterized in that Before inputting the speech data into the encoder to obtain the target acoustic features, the method further includes: Determining a batch set from the sample speech data, the batch set including speech frames corresponding to the original speech and speech frames corresponding to the noisy speech; Determine positive samples and negative samples of speech frames corresponding to the speech data in the batch set.
4. The method according to claim 3, characterized in that The determining of positive samples and negative samples of speech frames corresponding to the speech data in the batch set includes: For the batch set, performing speech feature alignment on speech frames corresponding to the speech data in the batch set to obtain feature alignment information; Based on the feature alignment information, speech frames with the same speech features in the batch set are determined as positive examples, and speech frames with different speech features in the batch set are determined as negative examples.
5. The method according to claim 3, characterized in that The speech recognition model further includes a contrast network structure comprising a cascaded first linear layer, an activation layer, a dropout layer, and a second linear layer; The reconstructing the target acoustic feature to obtain the reconstructed acoustic feature includes: Inputting the target acoustic feature into the comparison network structure to be processed sequentially through the first linear layer, the activation layer, the dropout layer, and the second linear layer; Acoustic features output by the second linear layer are obtained as the reconstructed acoustic features.
6. The method according to claim 3, characterized in that The second loss function includes an enhanced supervised contrast loss function; The constructing of the second loss function by comparing the reconstructed acoustic features of the original speech and the corresponding reconstructed acoustic features of the noisy speech at the same time sequence includes: For the batch set, based on the reconstructed acoustic features, obtaining similarities between each speech frame and positive samples and negative samples of the speech frames in the batch set, and determining a contrast loss of the speech frame according to the similarities; The enhanced supervised contrastive loss function is constructed according to the contrastive loss of each speech frame in the batch set.
7. The method according to claim 6, characterized in that The step of obtaining, for the batch set, a similarity between each speech frame and a positive sample and a negative sample of the speech frame in the batch set based on the reconstructed acoustic features, and determining a contrast loss of the speech frame according to the similarity, comprising: For each of the speech frames, obtaining a sum of first similarities between the speech frame and a positive example of the speech frame; Obtaining a sum of second similarities between the speech frame and other speech frames in the batch set except the speech frame; The contrast loss of the speech frame is constructed according to the sum of the first similarities and the sum of the second similarities.
8. The method according to claim 6, characterized in that The second loss function also includes a sequence supervision contrast loss function; The constructing of the second loss function by comparing the reconstructed acoustic features of the original speech and the corresponding reconstructed acoustic features of the noisy speech at the same time sequence includes: For the batch set, determining positive sample pairs and negative sample pairs according to the positive sample and the negative sample of each speech frame; Obtaining the sum of the third similarities of all positive example pairs in the batch set; Obtaining the sum of the fourth similarities of all negative example pairs in the batch set; The sequence supervision comparison loss function is constructed according to the sum of the third similarities and the sum of the fourth similarities.
9. The method according to claim 8, characterized in that The constructing of the second loss function by comparing the reconstructed acoustic features of the original speech and the corresponding reconstructed acoustic features of the noisy speech at the same time sequence includes: weighting the enhanced supervisory contrast loss function and the sequence supervisory contrast loss function based on weights pre-assigned to the enhanced supervisory contrast loss function and the sequence supervisory contrast loss function, respectively; The weighted first joint loss function is used as the second loss function.
10. The method according to claim 1, characterized in that The first loss function includes a continuous time classification loss function and an attention loss function; The constructing a first loss function based on the speech data and the target label sequence includes: Determining the continuous-time classification loss function based on an alignment result between the target label sequence and the true label sequence, taking the true label sequence of the speech data as a reference; For each of the true labels, determining the attention loss function according to the relationship between the true label and each true label preceding the true label in the true label sequence; The first loss function is determined based on the continuous-time classification loss function and the attention loss function.
11. The method according to any one of claims 1 to 10, characterized in that The updating of the parameters of the speech recognition model based on the first loss function and the second loss function to obtain a trained speech recognition model includes: Weighting the first loss function and the second loss function based on weights pre-assigned to the first loss function and the second loss function respectively; The parameters in the speech recognition model are updated using the weighted second joint loss function to obtain the trained speech recognition model.
12. A training device for a speech recognition model, characterized in that: The speech recognition model includes an encoder and a decoder, and the training device includes: a first data processing module, configured to input speech data into the encoder to obtain target acoustic features; the speech data including original speech and noisy speech corresponding to the original speech; a second data processing module, configured to decode the target acoustic features using the decoder to obtain a target label sequence, and construct a first loss function based on the speech data and the target label sequence; A contrastive learning module, configured to reconstruct the target acoustic feature to obtain a reconstructed acoustic feature, and construct a second loss function by comparing the reconstructed acoustic feature of the original speech and the corresponding reconstructed acoustic feature of the noisy speech at the same time sequence; A parameter updating module is used to update the parameters of the speech recognition model based on the first loss function and the second loss function to obtain a trained speech recognition model.
13. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 11 by executing the executable instructions.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Real-time speech enhancement method
CN109215674A
Mixed sound signal separation method and device, electronic equipment and readable medium
CN109801644A