Speech recognition model training method and device, electronic device, and storage medium
Patent Information
- Application Number
- CN202310444054.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-23
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-04-23
AI Technical Summary
相关技术中,通过利用整个目标序列计算整个句子的损失来对模型参数进行更新,但存在语音识别准确性低的缺点,进而影响了用户的使用体验
[0008]本公开实施例所提供的语音识别模型的训练方法,将有监督的语音数据输入至语音识别模型的编码器得到目标声学特征,一方面,利用解码器对目标声学特征进行解码,得到目标标签序列,以基于有监督的语音数据和目标标签序列构建第一损失函数,另一方面,对目标声学特征进行重构得到重构声学特征,以将有监督的语音数据对应的语音帧的正例样本和负例样本进行对比,进而利用得到的重构声学特征构建第二损失函数,使得模型既可以根据语音识别任务产生的损失进行参数更新,并且通过语音帧的正例样本和负例样本的对比学习,通过提取更鲁棒的声学表示,使得模型更好的学习到相同语音特征之间的相关性和不同语音特征之间的差异性,进而提升模型的训练精度。
Smart Images

Figure CN116364067B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, specifically to a training method for a speech recognition model, a training device for a speech recognition model, an electronic device, and a computer-readable storage medium. Background Technology
[0002] With the development of artificial intelligence technology, scenarios that use speech recognition to provide user services are gradually increasing, such as intelligent voice assistants and speech-to-text conversion. Consequently, many training methods for speech recognition models have emerged. Among these techniques, one method updates model parameters by calculating the loss of the entire sentence using the entire target sequence; however, this method suffers from low speech recognition accuracy, thus affecting the user experience. Summary of the Invention
[0003] The purpose of this disclosure is to provide a training method for a speech recognition model, a training device for a speech recognition model, an electronic device, and a computer-readable storage medium, thereby improving the accuracy of the speech recognition model and the accuracy of the speech recognition results to at least a certain extent.
[0004] According to a first aspect of this disclosure, a method for training a speech recognition model is provided. The speech recognition model includes an encoder and a decoder. The training method includes: inputting supervised speech data into the encoder to obtain target acoustic features; wherein the speech frames corresponding to the supervised speech data have their own positive and negative samples; decoding the target acoustic features using the decoder to obtain a target label sequence, and constructing a first loss function based on the supervised speech data and the target label sequence; reconstructing the target acoustic features to obtain reconstructed acoustic features, and comparing the positive and negative samples of the speech frames corresponding to the supervised speech data to construct a second loss function using the obtained reconstructed acoustic features; adjusting the parameters in the speech recognition model based on the first and second loss functions to obtain a trained speech recognition model.
[0005] According to a second aspect of this disclosure, a training apparatus for a speech recognition model is provided. The speech recognition model includes an encoder and a decoder. The training apparatus includes: a first data processing module for inputting supervised speech data into the encoder to obtain target acoustic features; wherein the speech frames corresponding to the supervised speech data have their own positive and negative samples; a second data processing module for decoding the target acoustic features using the decoder to obtain a target label sequence, and constructing a first loss function based on the supervised speech data and the target label sequence; a data comparison module for reconstructing the target acoustic features to obtain reconstructed acoustic features, and comparing the positive and negative samples of the speech frames corresponding to the supervised speech data to construct a second loss function using the obtained reconstructed acoustic features; and a parameter update module for adjusting the parameters in the speech recognition model based on the first and second loss functions to obtain a trained speech recognition model.
[0006] According to a third aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to perform the method described above.
[0007] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the method described above.
[0008] The speech recognition model training method provided in this disclosure involves inputting supervised speech data into the encoder of the speech recognition model to obtain target acoustic features. On one hand, the target acoustic features are decoded using a decoder to obtain a target label sequence, and a first loss function is constructed based on the supervised speech data and the target label sequence. On the other hand, the target acoustic features are reconstructed to obtain reconstructed acoustic features. Positive and negative samples of the speech frames corresponding to the supervised speech data are compared, and a second loss function is constructed using the obtained reconstructed acoustic features. This allows the model to update its parameters based on the loss generated by the speech recognition task, and through comparative learning of positive and negative samples of the speech frames, extracts more robust acoustic representations, enabling the model to better learn the correlation between the same speech features and the differences between different speech features, thereby improving the training accuracy of the model.
[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0011] Figure 1 A schematic diagram illustrating the stages involved in the technical solution of this disclosure embodiment is shown;
[0012] Figure 2 This schematic diagram illustrates the structure of a speech recognition model according to an exemplary embodiment of the present disclosure.
[0013] Figure 3 A flowchart illustrating a training method for a speech recognition model according to an exemplary embodiment of the present disclosure is shown.
[0014] Figure 4 This illustration schematically shows an exemplary embodiment of the present disclosure for determining positive and negative sample samples;
[0015] Figure 5 This schematic diagram illustrates a comparative network structure in an exemplary embodiment of the present disclosure.
[0016] Figure 6 This schematically illustrates a flowchart of a method for constructing a second loss function using the obtained reconstructed acoustic features in an exemplary embodiment of this disclosure;
[0017] Figure 7 This schematic diagram illustrates an exemplary embodiment of the present disclosure for determining positive sample pairs and negative sample pairs;
[0018] Figure 8 This schematically illustrates a flowchart of a method for constructing a second loss function using the obtained reconstructed acoustic features in an exemplary embodiment of this disclosure;
[0019] Figure 9 The flowchart illustrating an exemplary embodiment of the present disclosure is shown in which a first loss function is constructed based on supervised speech data and a target label sequence.
[0020] Figure 10 This schematic diagram illustrates the composition of a training apparatus for a speech recognition model in an exemplary embodiment of the present disclosure.
[0021] Figure 11 A schematic diagram of an electronic device to which embodiments of the present disclosure may be applied is shown. Detailed Implementation
[0022] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0023] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0024] Figure 1 The diagram illustrates the stages involved in the technical solution of this disclosure embodiment, as shown below. Figure 1 As shown, the technical solution of this disclosure includes a model training stage and a model application stage.
[0025] In this embodiment of the disclosure, the training method for the provided speech recognition model can be executed by a terminal device. In the case where the training method for the speech recognition model in this embodiment of the disclosure is executed by the terminal device, each step is executed by the terminal device.
[0026] For example, the processor of the terminal device performs the following steps: supervising speech data is input into the encoder to obtain target acoustic features; the target acoustic features are decoded using the decoder to obtain a target label sequence, and a first loss function is constructed based on the supervised speech data and the target label sequence; the target acoustic features are reconstructed to obtain reconstructed acoustic features, and positive and negative samples of the speech frames corresponding to the supervised speech data are compared to construct a second loss function using the obtained reconstructed acoustic features; the parameters in the speech recognition model are adjusted based on the first and second loss functions to obtain the trained speech recognition model.
[0027] The terminal device can be an intelligent device with data processing capabilities, such as a smartphone, computer, tablet, in-vehicle device, wearable device, monitoring device, etc. The terminal device can also be referred to as a mobile terminal, terminal, mobile device, etc. This disclosure does not limit the type of terminal device.
[0028] Furthermore, the speech recognition model training method provided in this embodiment can also be executed by a server. Correspondingly, in this server-executed method, the server can respond to a trigger command to begin executing the steps in the speech recognition model training method of this embodiment. This trigger command can be sent by a user's terminal device or triggered locally by the server in response to some automated event. The server can be a backend system providing the relevant services in this embodiment, and may include a single electronic device with computing capabilities, such as a portable computer, desktop computer, or smartphone, or a cluster of multiple electronic devices.
[0029] Furthermore, the technical solutions of this disclosure can also be executed collaboratively by a terminal device and a server. In this method of collaborative execution by an electronic device and a server, some steps in the technical solutions provided by this disclosure are executed by the terminal device, while other steps are executed by the server. For example, the technical solutions provided by this disclosure can involve the terminal device acquiring supervised speech data and sending it to the server, the server performing the model training phase, and sending the trained speech recognition model back to the terminal device, whereby the terminal device performs the model application phase.
[0030] It should be noted that in this method where the terminal device and the server work together, the steps executed by the terminal device and the server can be dynamically adjusted according to the actual situation, and no special restrictions are imposed on this.
[0031] In this embodiment of the disclosure, the technical solution is described as being executed by a terminal device.
[0032] In related technologies, speech recognition models are based on an encoder-decoder structure to perform speech recognition end-to-end. Typically, the loss of the entire sentence is calculated by using the entire target sequence. However, this approach weakens the relationship between each frame of the model output and its actual corresponding speech features (such as phonemes). Consequently, in some application scenarios, speech features (such as phonemes) that are difficult to distinguish are often misidentified, leading to a decrease in the accuracy of speech recognition.
[0033] Based on one or more of the above-mentioned problems, the exemplary embodiments of this disclosure first provide a training method for a speech recognition model. While constructing a first loss function based on supervised speech data and target label sequences, a contrastive loss function (i.e., a second loss function) is constructed by comparing positive and negative samples of the speech frames corresponding to the supervised speech data and utilizing reconstructed acoustic features. This supervised contrastive learning enables the model to better learn the correlation between the same speech features (such as phonemes) and the differences between different speech features (such as phonemes), bringing the representations of the same speech features (such as phonemes) closer together and widening the representations of different speech features (such as phonemes) further apart. This improves the model's ability to recognize speech features that are difficult to distinguish, thereby improving the accuracy of the speech recognition model.
[0034] In this embodiment of the disclosure, the trained speech recognition model can be applied in an end-to-end manner to various speech recognition task-related scenarios. Examples include transcribing audio files, generating subtitles in real-time during live broadcasts, recording court proceedings in real-time, and enabling intelligent voice assistants to interact with smart devices via voice.
[0035] Furthermore, the technical solutions of the embodiments disclosed herein can be applied to various speech recognition frameworks, including but not limited to wenet, kaldi2, espnet, wav2etter++, etc.
[0036] refer to Figure 2 As shown, the speech recognition model in this embodiment is based on a continuous-temporal classification (CTC) attention encoder-decoder framework, including an encoder and a decoder. This speech recognition model can include an ASR (Automatic Speech Recognition) subtask and a contrastive subtask, corresponding to a classification network structure and a contrastive network structure, respectively, to perform speech recognition and feature contrastive learning tasks.
[0037] The encoder can be understood as a shared network. The encoder's output is used for speech recognition subtasks, feature contrast learning tasks, and as acoustic features for the decoder. Speech recognition subtasks can be performed based on linear layers and softmax layers, etc. The embodiments of this disclosure do not impose special restrictions on the classification network structure.
[0038] The encoder can be selected from encoder structures such as Conformer, LSTM, and Transformer, and this disclosure does not impose any special limitations on this.
[0039] See Figure 3As shown, the training method for the speech recognition model in this embodiment may include the following steps S310 to S340:
[0040] In step S310, supervised speech data is input into the encoder to obtain target acoustic features; wherein, the speech frames corresponding to the supervised speech data have their own positive and negative sample samples.
[0041] In an exemplary embodiment of this disclosure, supervised speech data refers to speech data that has been labeled with artificial tags as a reference benchmark, which will be described in the following description as a sequence of real tags.
[0042] The speech data can be acoustic features such as Fbank features and MFCC (Mel-scale Frequency Cepstral Coefficients) features, which serve as input to the encoder, and each speech data has a corresponding speech frame. Taking Fbank features as an example, speech samples are first acquired, then the Fbank features of the speech samples are extracted as speech data, and the true label sequence is determined based on manual annotation and speech alignment information.
[0043] In supervised speech data, positive examples of speech frames refer to speech frames with the same speech features as the given speech frame, while negative examples refer to speech frames with different speech features. Speech features can be, for example, phonemes; that is, speech frames with the same phonemes as a given speech frame are positive examples of that speech frame, and speech frames with different phonemes are negative examples of that speech frame.
[0044] Taking speech features as phonemes as an example, if the supervised speech data corresponds to speech frames including X = {x1, x2, x3, ..., x...} n Then, based on whether the phonemes of the speech frames are the same, the positive and negative samples of speech frame x1 are determined, the positive and negative samples of x2 are determined, and so on, to determine x n The speech frame contains both positive and negative examples. Correspondingly, each speech frame can form a positive example pair with its positive examples, and each speech frame can form a negative example pair with its negative examples. For example... Figure 4 x2 and x1 form a positive sample pair, x n x1 forms a negative sample pair, and x n x2 forms a negative sample pair; other speech frames will not be listed one by one.
[0045] Of course, the speech features in this disclosure can also be words, etc., and the positive sample, negative sample, positive sample pair and negative sample pair of each speech frame can be obtained in the manner of using phonemes as speech features. This disclosure includes, but is not limited to, the above-mentioned speech features.
[0046] In step S320, the target acoustic features are decoded using a decoder to obtain a target label sequence, and a first loss function is constructed based on the supervised speech data and the target label sequence.
[0047] In an exemplary embodiment of this disclosure, the first loss function is a loss based on the ASR subtask, which can use the target acoustic features output by the encoder as the acoustic features of the decoder, and decode the target acoustic features to obtain the target label sequence, so as to construct the first loss function based on the supervised speech data and the target label sequence.
[0048] In step S330, the target acoustic features are reconstructed to obtain reconstructed acoustic features, and the positive and negative samples of the speech frames corresponding to the supervised speech data are compared to construct a second loss function using the obtained reconstructed acoustic features.
[0049] In an exemplary embodiment of this disclosure, the second loss function is a supervised contrastive loss based on a contrastive subtask. Reconstructing acoustic features refers to obtaining acoustic features by further processing the target acoustic features through a preset network structure in order to obtain more obvious contrastive features. The preset network structure can be a contrastive network structure.
[0050] In this embodiment, the representations of the same speech features are brought closer together and the representations of different speech features are pulled further apart. Positive and negative samples obtained by forced alignment of speech features are compared, and a second loss function is constructed using reconstructed acoustic features.
[0051] In step S340, the parameters in the speech recognition model are adjusted based on the first loss function and the second loss function to obtain the trained speech recognition model.
[0052] In an exemplary embodiment of this disclosure, the parameters in the speech recognition model can be updated by combining a first loss function and a second loss function to obtain a trained speech recognition model.
[0053] The training method for the speech recognition model based on the embodiments of this disclosure enables the model to update parameters according to the loss generated by the speech recognition task, and to learn more robust acoustic representations by comparing positive and negative samples of speech frames, thereby improving the model's training accuracy by extracting more robust acoustic representations.
[0054] In one exemplary embodiment, before inputting supervised speech data into the encoder to obtain target acoustic features, a batch set can be determined from the sample speech data, the batch set including the speech frames corresponding to the supervised speech data.
[0055] The sample speech data is the set of all training samples used for model training. The parameters of the model are updated by backpropagation using a portion of the samples in the set. This portion of the samples is the determined batch.
[0056] For example, if the sample speech data includes all speech frames, then a portion of the speech frames can be selected from the sample speech data to form a batch set for one iteration of model training. That is, the batch set includes the speech frames corresponding to the supervised speech data. In this embodiment of the disclosure, all speech frames in the batch set are compared.
[0057] Based on this, in an exemplary embodiment, an implementation method for determining positive and negative samples of speech frames corresponding to supervised speech data is also provided, which may include:
[0058] First, the speech frames in the batch set are processed for phoneme alignment to obtain phoneme alignment information. Then, based on the phoneme alignment information, speech frames with the same phonemes in the batch set are identified as positive samples, and speech frames with different phonemes in the batch set are identified as negative samples. For details, please refer to the example in step S310.
[0059] The embodiments of this disclosure determine positive and negative samples based on forced alignment, so that subsequent schemes can compare positive and negative samples, cluster the same phonemes more closely, and keep different phonemes further apart, thereby reducing replacement errors.
[0060] In one exemplary embodiment, the speech recognition model further includes a contrast head structure, see [link to example]. Figure 5 A schematic diagram of a contrast network structure according to an exemplary embodiment of the present disclosure is shown. The contrast network structure includes a first linear layer, an activation layer, a dropout layer, and a second linear layer.
[0061] Based on the contrastive network structure, for a batch set, the target acoustic features are reconstructed to obtain the reconstructed acoustic features by: inputting the target acoustic features into the contrastive network structure, processing them sequentially through the first linear layer, the activation layer, the dropout layer, and the second linear layer, and obtaining the acoustic features output by the second linear layer as the reconstructed acoustic features.
[0062] Specifically, the reconstructed acoustic features of the comparison network structure output can be obtained through formula (1):
[0063] O=w2*Dropout((w1*x)*sigmoid(w1*x)) (1)
[0064] Wherein, O is the reconstructed acoustic feature output by the contrast network structure, x is the target acoustic feature output by the encoder, x∈(B,T,C), B is the size of the batch set, T is the number of speech frames in the batch set, and C is the dimension of each speech frame. In this embodiment of the present disclosure, the output of the contrast network structure is adjusted from O∈(B,T,C) to O∈(B*T,C) to calculate the contrast loss of a batch set; w1 and w2 are weight parameters, respectively.
[0065] Based on the comparison network structure of this disclosure embodiment, two linear layers with activation layers are used to project the encoder output to facilitate subsequent comparative learning of features based on reconstructed acoustic features.
[0066] It should be noted that, Figure 5 The model parameters in the network structures shown for comparison are merely illustrative and can be adaptively adjusted according to actual needs; no special limitations are imposed on them.
[0067] In one exemplary embodiment, the second loss function includes an Enhanced Supervised Contrastive Learning (ESCL) loss function. For example... Figure 6 As shown, comparing positive and negative samples of the speech frames corresponding to the supervised speech data to construct a second loss function using the obtained reconstructed acoustic features may include steps S610 and S620:
[0068] Step S610: For the batch set, based on the reconstructed acoustic features, obtain the similarity between each speech frame and the positive and negative sample samples of the speech frames in the batch set, and determine the contrast loss of the speech frame according to the similarity.
[0069] For each speech frame in the batch set, the similarity between the speech frame and its positive sample can be obtained, and the sum of all similarities can be obtained to get the first similarity sum. The second similarity sum between the speech frame and other speech frames in the batch set other than the speech frame can be obtained. Finally, the contrast loss of the speech frame is constructed based on the first similarity sum and the second similarity sum.
[0070] Step S620: Construct an enhanced supervised contrastive loss function based on the contrastive loss of each speech frame in the batch set.
[0071] After obtaining the contrast loss for each speech frame in the batch set, an enhanced supervised contrast loss function can be constructed based on the contrast loss of each speech frame in the batch set.
[0072] This disclosure embodiment simultaneously compares all information in the batch set, comparing each speech frame with all its corresponding positive and negative sample instances. For example, see... Figure 7As shown, the batch set may include multiple subsets (such as batch1, batch2, and batch3, etc.). The first "l" and the second-to-last "l" can be considered positive samples, and the second-to-last "l", the last "iu", and the last "ao" can be considered negative samples. It should be noted that this embodiment compares all information in the batch set. Figure 7 Other positive and negative samples are not fully shown.
[0073] Specifically, the enhanced supervised contrastive loss function can be constructed using formula (2):
[0074]
[0075] Among them, Loss escl To enhance the supervised contrastive loss function, S is the length of the contrastive network structure output, S = B*T, where B is the size of the batch set, T is the number of speech frames in the batch set, O is the reconstructed acoustic features (output of the contrastive network structure), J is the set of positive samples corresponding to the i-th speech frame, K is the set of other speech frames in the batch set besides the i-th speech frame, and τ is a hyperparameter for adjusting the ratio. τ can be set to 0.07, or it can be set according to actual needs without special restrictions.
[0076] The similarity between two samples can be obtained by formula (3):
[0077]
[0078] Where sim(·,·) represents the similarity between two samples, and α and β are the reconstructed acoustic features O corresponding to different samples.
[0079] The enhanced supervised contrastive loss function of this disclosure can enhance the robustness of feature representation by bringing the representation of positive examples closer together and pulling the representation of negative examples further apart.
[0080] In one exemplary embodiment, the second loss function further includes a Sequential Supervised Contrastive Learning (SSCL) loss function.
[0081] like Figure 8 As shown, comparing positive and negative samples of the speech frames corresponding to the supervised speech data to construct a second loss function using the obtained reconstructed acoustic features may include steps S810 to S840:
[0082] Step S810: For the batch set, determine the positive sample pairs and negative sample pairs based on the positive and negative sample pairs of the speech frames corresponding to the supervised speech data.
[0083] For each speech frame in the batch set, a positive sample pair is determined based on the speech frame and its positive sample pairs, and a negative sample pair is determined based on the speech frame and its negative sample pairs. Then, the positive sample pairs required for this step are obtained based on the positive sample pairs of all speech frames, and the negative sample pairs required for this step are obtained based on the negative sample pairs of all speech frames.
[0084] Step S820: Obtain the sum of the third similarities of all positive sample pairs in the batch set.
[0085] Based on the reconstructed speech features output by the contrastive network structure, the sum of the third similarity of all positive sample pairs in the batch set is obtained.
[0086] Step S830: Obtain the sum of the fourth similarities of all negative sample pairs in the batch set.
[0087] Similarly, based on the reconstructed speech features output by the contrastive network structure, the sum of the fourth similarity of all negative sample pairs in the batch set is obtained.
[0088] Step S840: Construct a sequence-supervised contrastive loss function based on the sum of the third similarity and the sum of the fourth similarity.
[0089] Specifically, the sequence-supervised contrastive loss function can be constructed using formula (4):
[0090]
[0091] Where P represents all positive sample pairs, N represents all negative sample pairs, i and j are the identifiers corresponding to all positive sample pairs output by the contrast network structure, and k and w are the identifiers corresponding to all negative sample pairs output by the corresponding network structure. The similarity between two samples and the method of obtaining τ can be found in the section on constructing the enhanced supervised contrastive loss function, and will not be repeated here.
[0092] By constructing a sequence-supervised contrastive loss function, which compares all positive and negative samples in the batch, the speech recognition model pays more attention to alignment information and improves the targeting of feature comparison.
[0093] In one exemplary embodiment, constructing a second loss function based on the obtained enhanced supervision contrastive loss function and sequence supervision contrastive loss function includes:
[0094] First, the enhanced supervision contrastive loss function and the sequence supervision contrastive loss function are weighted based on the weights pre-assigned to them respectively; then, the first joint loss function obtained by weighting is used as the second loss function.
[0095] Specifically, the second loss function can be constructed using formula (5):
[0096] Loss contrast =βLoss sscl +(1-β)Loss escl (5)
[0097] Among them, Loss contrast For the second loss function, Loss ESCL To enhance the supervised contrastive loss function, Loss SSCL Let be the sequence-supervised contrastive loss function, and β and (1-β) be the weights pre-assigned to the augmented supervision contrastive loss function and the sequence-supervised contrastive loss function, respectively. β is used as a hyperparameter to balance the two supervision contrastive loss functions.
[0098] In an exemplary embodiment, the first loss function includes a continuous-time classification loss function and an attention loss function. The continuous-time classification function is a loss function that describes the loss that occurs between sequences based on the assumption of conditional independence. The continuous-time classification loss function describes the loss that occurs in the process of supervising speech data to align to the target label sequence. The attention loss function is used to describe the loss that occurs in the attention mechanism in the speech recognition model.
[0099] like Figure 9 As shown, constructing the first loss function based on supervised speech data and target label sequences includes steps S910 to S930:
[0100] Step S910: Using the real label sequence of supervised speech data as a reference, determine the continuous-time classification loss function based on the alignment result between the target label sequence and the real label sequence.
[0101] The continuous-time classification loss function shares the encoder in the speech recognition model with the decoder in the speech recognition model. The continuous-time classification loss function makes a conditional assumption on the factors output by the decoder in the speech recognition model.
[0102] For example, using the real label sequence of supervised speech data as a reference, the loss in aligning supervised speech data to the target label sequence is independently evaluated according to the forward-backward algorithm conditions. Specifically, the continuous-time classification loss function can be constructed using formula (6):
[0103]
[0104] Among them, Loss CTC Let x be the real label sequence of supervised speech data, and y be the target label sequence.
[0105] It should be noted that, based on supervised speech data and target label sequences, continuous-time classification loss functions can also be constructed in other ways, not limited to the method shown in formula (6). The specific method of constructing continuous-time classification loss functions is not limited in the embodiments of this disclosure.
[0106] Step S920: For each real label, determine the attention loss function based on the relationship between the real label sequence and each real label preceding the real label in the real label sequence.
[0107] Using the real label sequence as a reference, the loss of the entire speech recognition model due to the attention mechanism is evaluated based on the relationship between the real labels in the real label sequence and the real labels preceding the real label, thus obtaining the attention loss function.
[0108] For example, the attention loss function can be constructed using formula (7):
[0109]
[0110] Among them, Loss attention Let y* be the attention loss function. n It is the nth real label in the real label sequence, y* [1:n-1] It is y* n The preceding (n-1) real labels, x is the sequence of real labels for supervised speech data.
[0111] It should be noted that attention loss functions can also be constructed in other ways based on supervised speech data and real label sequences, and are not limited to the method shown in formula (7). The specific method of constructing attention loss functions is not limited in the embodiments of this disclosure.
[0112] Step S930: Determine the first loss function based on the continuous-time classification loss function and the attention loss function.
[0113] The continuous-time classification loss function and the attention loss function can be weighted based on the weights pre-assigned to them, and the resulting third joint loss function can be used as the first loss function.
[0114] Specifically, the first loss function can be constructed using formula (8):
[0115] Loss ASR =θLoss CTC +(1-θ)Loss attention (8)
[0116] Among them, Loss ASRθ is the first loss function, and θ is a hyperparameter that balances the continuous-time classification loss function and the attention loss function, for example, 0.3 or set according to actual needs.
[0117] In an exemplary embodiment, adjusting the parameters in the speech recognition model based on a first loss function and a second loss function to obtain the trained speech recognition model may include:
[0118] First, based on the weights pre-assigned to the first and second loss functions respectively, the first and second loss functions are weighted. Then, the parameters in the speech recognition model are adjusted using the weighted second joint loss function to obtain the trained speech recognition model.
[0119] Specifically, the second joint loss function can be constructed using formula (9):
[0120] Loss = Loss ASR +αLoss contrast (9)
[0121] Where Loss is the second joint loss function, and α is a hyperparameter used to balance the first and second loss functions.
[0122] Table 1 shows the comparison results of various models in the embodiments of this disclosure. The four models are listed from top to bottom as follows: Model 1: Open source continuous time classification CTC attention model (Conformer CTC prefix beam search); Model 2: Baseline version of the open source continuous time classification CTC attention model based on preset conditions (Conformer attention rescoring); Model 3: Model after adding the enhanced contrast loss and sequence supervision contrast loss of the embodiments of this disclosure to Model 2 (Conformer+supvisied contrast loss+CTC prefix beam search); Model 4: Model after adding attention rescoring to Model 3 (CTC-Conformer+supvisied contrast loss+attention rescoring).
[0123] Table 1
[0124]
[0125]
[0126] As shown in Table 1, the Phoneme Error Rate (PER) for each model was determined based on substitution errors, deletion errors, and insertion errors. Specifically, the PERs for Model 3 and Model 4 were 15.01 and 15.08, respectively. The PER for Model 3 was 1.46% lower than that for Model 1, and the PER for Model 4 was 1.27% lower than that for Model 2.
[0127] Furthermore, compared to Model 1 or Model 2, the deletion errors of Model 3 and Model 4 are significantly reduced, indicating that the embodiments of this disclosure, by adding an enhanced supervised contrastive loss function and a sequence supervised contrastive loss function in the calculation of the loss function, successfully attract the model's attention to the speech frames that are easy to ignore the alignment information generated by forced alignment. By contrastive learning, the model brings the representations of the same phonemes closer together and widens the representations of different phonemes apart, strengthening the relationship between each frame output by the model and its actual corresponding phoneme, enabling the model to learn phonemes that are difficult to distinguish.
[0128] In summary, the embodiments of this disclosure update the parameters in the speech recognition model according to the second joint loss function. This allows the model to update its parameters based on the first loss function generated by the speech recognition task, and through comparative learning of positive and negative samples of speech frames, extract more robust acoustic representations. This enables the model to better learn the correlation between the same speech features and the differences between different speech features based on the second loss function, allowing the model to focus on easily confused phonemes, reduce substitution errors, and thus improve the training accuracy of the model.
[0129] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0130] Further reference Figure 10 As shown, an exemplary embodiment of this disclosure provides a speech recognition model training device 1000, including a first data processing module 1010, a second data processing module 1020, a data comparison module 1030, and a parameter update module 1040. Wherein:
[0131] The first data processing module 1010 is used to input supervised speech data into the encoder to obtain target acoustic features; wherein, the speech frames corresponding to the supervised speech data have their own positive and negative sample samples.
[0132] The second data processing module 1020 is used to decode the target acoustic features using a decoder to obtain a target label sequence, and to construct a first loss function based on supervised speech data and the target label sequence.
[0133] The data comparison module 1030 is used to reconstruct the target acoustic features to obtain reconstructed acoustic features, and compare the positive and negative samples of the speech frames corresponding to the supervised speech data to construct a second loss function using the obtained reconstructed acoustic features.
[0134] The parameter update module 1040 is used to adjust the parameters in the speech recognition model based on the first loss function and the second loss function to obtain the trained speech recognition model.
[0135] In one exemplary embodiment, the first data processing module 1010 is further configured to perform: determining a batch set from the sample speech data, the batch set including speech frames corresponding to supervised speech data.
[0136] In an exemplary embodiment, the speech recognition model further includes a contrast network structure, which includes a cascaded first linear layer, an activation layer, a dropout layer, and a second linear layer; the data contrast module 1030 is configured to perform: inputting target acoustic features into the contrast network structure for processing through the first linear layer, the activation layer, the dropout layer, and the second linear layer in sequence; and obtaining the acoustic features output by the second linear layer as reconstructed acoustic features.
[0137] In an exemplary embodiment, the second loss function includes an enhanced supervised contrastive loss function; the data comparison module 1030 is configured to perform: for a batch set, based on reconstructed acoustic features, obtain the similarity between each speech frame and the positive and negative sample samples of the speech frames in the batch set, and determine the contrastive loss of the speech frames according to the similarity; and construct an enhanced supervised contrastive loss function based on the contrastive loss of each speech frame in the batch set.
[0138] In an exemplary embodiment, the data comparison module 1030 is configured to perform the following: for each speech frame, obtain the sum of a first similarity between the speech frame and positive samples of the speech frame; obtain the sum of a second similarity between the speech frame and other speech frames in the batch set other than the speech frame; and construct a contrast loss for the speech frame based on the sum of the first similarity and the sum of the second similarity.
[0139] In an exemplary embodiment, the second loss function further includes a sequence-supervised contrastive loss function; the data contrast module 1030 is configured to perform: for a batch set, determining positive sample pairs and negative sample pairs based on positive and negative sample pairs of the speech frames corresponding to the supervised speech data; obtaining the sum of third similarities of all positive sample pairs in the batch set; obtaining the sum of fourth similarities of all negative sample pairs in the batch set; and constructing a sequence-supervised contrastive loss function based on the sum of third and fourth similarities.
[0140] In an exemplary embodiment, the data comparison module 1030 is configured to perform: weighting the enhanced supervision comparison loss function and the sequence supervision comparison loss function based on weights pre-assigned to the enhanced supervision comparison loss function and the sequence supervision comparison loss function respectively; and using the weighted first joint loss function as the second loss function.
[0141] In one exemplary embodiment, the first loss function includes a continuous-time classification loss function and an attention loss function; the second data processing module 1020 is configured to perform: using the real label sequence of supervised speech data as a reference, determining the continuous-time classification loss function based on the alignment result between the target label sequence and the real label sequence; for each real label, determining the attention loss function based on the relationship between the real label and each real label preceding the real label in the real label sequence; and determining the first loss function based on the continuous-time classification loss function and the attention loss function.
[0142] In an exemplary embodiment, the first data processing module 1010 is configured to perform the following process: for a batch set, determining positive and negative samples of speech frames corresponding to supervised speech data: performing phoneme alignment processing on the speech frames in the batch set to obtain phoneme alignment information; based on the phoneme alignment information, determining speech frames with the same phonemes in the batch set as positive samples, and determining speech frames with different phonemes in the batch set as negative samples.
[0143] In an exemplary embodiment, the parameter update module 1040 is configured to perform: weighting the first loss function and the second loss function based on the weights pre-assigned to the first loss function and the second loss function respectively; and adjusting the parameters in the speech recognition model using the weighted second joint loss function to obtain the trained speech recognition model.
[0144] The specific details of each module in the above-mentioned device have been described in detail in the method section of the implementation. For any undisclosed details, please refer to the implementation content of the method section, and therefore will not be repeated here.
[0145] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0146] An exemplary embodiment of this disclosure also provides an electronic device for the above-described method. This electronic device may be the aforementioned imaging device or a server. Generally, the electronic device includes at least a processor and a memory, the memory for storing executable instructions of the processor, and the processor configured to perform the above-described method by executing the executable instructions.
[0147] The following is based on Figure 11 Taking the mobile terminal 1100 as an example, the construction of the electronic device in this embodiment of the present disclosure will be described by way of example. Those skilled in the art should understand that, apart from components specifically designed for mobile purposes, Figure 11 The structure shown can also be applied to fixed-type devices. In other embodiments, the mobile terminal 1100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware. The interface connections between the components are only schematic and do not constitute a limitation on the structure of the mobile terminal 1100. In other embodiments, the mobile terminal may also adopt a similar design to... Figure 11 Different interface connection methods, or combinations of multiple interface connection methods.
[0148] like Figure 11 As shown, the mobile terminal 1100 may specifically include: a processor 1101, a memory 1102, a bus 1103, a mobile communication module 1104, an antenna 1, a wireless communication module 1105, an antenna 2, a display screen 1106, a camera module 1107, an audio module 1108, a power module 1109, and a sensor module 1110.
[0149] The processor 1101 may include one or more processing units, such as an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit).
[0150] An encoder encodes (compresses) images or videos to reduce data size for easier storage or transmission. A decoder decodes (decompresses) the encoded data to restore the original image or video data. The mobile terminal 1100 can support one or more encoders and decoders, such as image formats like JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), and BMP (Bitmap), and video formats like MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).
[0151] The processor 1101 can be connected to the memory 1102 or other components via the bus 1103.
[0152] The memory 1102 can be used to store computer executable program code, which includes instructions. The processor 1101 executes various functional applications and data processing of the mobile terminal 1100 by running the instructions stored in the memory 1102. The memory 1102 can also store application data, such as images, videos, and other files.
[0153] The communication function of mobile terminal 1100 can be implemented through mobile communication module 1104, antenna 1, wireless communication module 1105, antenna 2, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1104 can provide 3G, 4G, 5G and other mobile communication solutions for mobile terminal 1100. Wireless communication module 1105 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1100.
[0154] The display screen 1106 is used to implement display functions, such as displaying the user interface, images, videos, etc., and displaying abnormal prompts. The camera module 1107 is used to implement shooting functions, such as capturing images and videos to acquire scene images. The audio module 1108 is used to implement audio functions, such as playing audio and capturing voice. The power module 1109 is used to implement power management functions, such as charging the battery, supplying power to the device, and monitoring battery status. The sensor module 1110 may include one or more sensors for implementing corresponding sensing and detection functions.
[0155] Furthermore, exemplary embodiments of this disclosure also provide a computer-readable storage medium storing a program product capable of implementing the methods described above. In some possible embodiments, various aspects of this disclosure may also be implemented as a program product including program code that, when run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure.
[0156] It should be noted that the computer-readable medium disclosed herein may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0157] In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0158] Furthermore, program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0159] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
Claims
1. A method for training a speech recognition model, characterized in that, The speech recognition model includes an encoder and a decoder, and the training method includes: Supervised speech data is input into the encoder to obtain target acoustic features; wherein, the speech frames corresponding to the supervised speech data have their own positive and negative sample samples; The target acoustic features are decoded using the decoder to obtain a target label sequence, and a first loss function is constructed based on the supervised speech data and the target label sequence. The target acoustic features are reconstructed to obtain reconstructed acoustic features, and the positive and negative samples of the speech frames corresponding to the supervised speech data are compared to construct a second loss function using the obtained reconstructed acoustic features. The parameters in the speech recognition model are adjusted based on the first loss function and the second loss function to obtain the trained speech recognition model. The method further includes: determining a batch set from the sample speech data, the batch set including speech frames corresponding to the supervised speech data; the speech recognition model further includes a contrastive network structure, the contrastive network structure including a cascaded first linear layer, an activation layer, a dropout layer, and a second linear layer; for the batch set, reconstructing the target acoustic features to obtain reconstructed acoustic features includes: The target acoustic features are input into the contrast network structure and processed sequentially through the first linear layer, the activation layer, the dropout layer, and the second linear layer; the acoustic features output by the second linear layer are obtained as the reconstructed acoustic features; For the batch set, the process of determining the positive and negative examples of the speech frames corresponding to the supervised speech data includes: The speech frames in the batch set are processed to perform phoneme alignment to obtain phoneme alignment information; based on the phoneme alignment information, speech frames with the same phonemes in the batch set are identified as positive samples, and speech frames with different phonemes in the batch set are identified as negative samples.
2. The method according to claim 1, characterized in that, The second loss function includes the enhanced supervised contrastive loss function; The step of comparing positive and negative samples of the speech frames corresponding to the supervised speech data to construct a second loss function using the obtained reconstructed acoustic features includes: For the batch set, based on the reconstructed acoustic features, the similarity between each speech frame and the positive and negative samples of the speech frames in the batch set is obtained, and the contrast loss of the speech frame is determined according to the similarity. The enhanced supervised contrastive loss function is constructed based on the contrastive loss of each speech frame in the batch set.
3. The method according to claim 2, characterized in that, For the batch set, based on the reconstructed acoustic features, the similarity between each speech frame and the positive and negative samples of the speech frames in the batch set is obtained, and the contrast loss of the speech frame is determined based on the similarity, including: For each of the speech frames, the sum of the first similarities between the speech frame and the positive sample of the speech frame is obtained; Obtain the sum of the second similarities between the speech frame and other speech frames in the batch set excluding the speech frame; The contrast loss of the speech frame is constructed based on the sum of the first similarity and the sum of the second similarity.
4. The method according to claim 2, characterized in that, The second loss function also includes a sequence-supervised contrastive loss function; The step of comparing positive and negative samples of the speech frames corresponding to the supervised speech data to construct a second loss function using the obtained reconstructed acoustic features includes: For the batch set, positive sample pairs and negative sample pairs are determined based on the positive and negative sample pairs of the speech frames corresponding to the supervised speech data; Obtain the sum of the third similarities of all positive sample pairs in the batch set; Obtain the sum of the fourth similarity scores for all negative sample pairs in the batch set; The sequence-supervised contrastive loss function is constructed based on the sum of the third similarity and the sum of the fourth similarity.
5. The method according to claim 4, characterized in that, The process of reconstructing the target acoustic features to obtain reconstructed acoustic features, and comparing positive and negative samples of the speech frames corresponding to the supervised speech data to construct a second loss function using the obtained reconstructed acoustic features, includes: The enhanced supervision contrast loss function and the sequence supervision contrast loss function are weighted based on the weights pre-assigned to them respectively. The first joint loss function obtained by weighting is used as the second loss function.
6. The method according to claim 1, characterized in that, The first loss function includes a continuous-time classification loss function and an attention loss function; The construction of the first loss function based on the supervised speech data and the target label sequence includes: Using the real label sequence of the supervised speech data as a reference, the continuous-time classification loss function is determined based on the alignment result between the target label sequence and the real label sequence; For each real label, the attention loss function is determined based on the relationship between the real label and each real label preceding it in the real label sequence; The first loss function is determined based on the continuous-time classification loss function and the attention loss function.
7. The method according to any one of claims 1 to 6, characterized in that, The step of adjusting the parameters in the speech recognition model based on the first loss function and the second loss function to obtain the trained speech recognition model includes: The first loss function and the second loss function are weighted based on the weights pre-assigned to the first loss function and the second loss function, respectively. The parameters in the speech recognition model are adjusted using the weighted second joint loss function to obtain the trained speech recognition model.
8. A training device for a speech recognition model, characterized in that, The speech recognition model includes an encoder and a decoder, and the training device includes: The first data processing module is used to input supervised speech data into the encoder to obtain target acoustic features; wherein the speech frames corresponding to the supervised speech data have their own positive and negative sample samples; The second data processing module is used to decode the target acoustic features using the decoder to obtain a target label sequence, and to construct a first loss function based on the supervised speech data and the target label sequence. The data comparison module is used to reconstruct the target acoustic features to obtain reconstructed acoustic features, and compare the positive and negative samples of the speech frames corresponding to the supervised speech data to construct a second loss function using the obtained reconstructed acoustic features. The parameter update module is used to adjust the parameters in the speech recognition model based on the first loss function and the second loss function to obtain the trained speech recognition model. The first data processing module is further configured to perform: determining a batch set from the sample speech data, the batch set including the speech frames corresponding to the supervised speech data; The speech recognition model further includes a contrastive network structure, which comprises a cascaded first linear layer, an activation layer, a dropout layer, and a second linear layer; for the batch set, the reconstruction of the target acoustic features to obtain reconstructed acoustic features includes: The target acoustic features are input into the contrast network structure and processed sequentially through the first linear layer, the activation layer, the dropout layer, and the second linear layer; the acoustic features output by the second linear layer are obtained as the reconstructed acoustic features; For the batch set, the process of determining the positive and negative examples of the speech frames corresponding to the supervised speech data includes: The speech frames in the batch set are processed to perform phoneme alignment to obtain phoneme alignment information; based on the phoneme alignment information, speech frames with the same phonemes in the batch set are identified as positive samples, and speech frames with different phonemes in the batch set are identified as negative samples.
9. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Voice recognition model training method and device, electronic equipment and storage medium
CN111916067A
Speech recognition method for comparative predictive coding self-supervised structure joint training
CN112767922A
Contrast type context understanding enhancement method for dialogue emotion recognition
CN113946670A