Model training, speech recognition method and device, electronic equipment and storage medium

By initializing and fine-tuning symmetric convolutional networks and causal convolutional networks, the problem of not being able to train a high-accuracy streaming speech recognition model in existing technologies is solved, and a speech recognition model suitable for streaming tasks is realized, which has versatility.

CN116246613BActive Publication Date: 2026-04-14JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies cannot train speech recognition models that are suitable for streaming tasks and have high recognition accuracy.

Method used

By acquiring a pre-trained model and the original recognition model, and taking advantage of the structural differences between symmetric convolutional networks and causal convolutional networks, initialization is performed for each convolutional kernel, and fine-tuning is carried out based on multiple sets of training samples to obtain a speech recognition model suitable for streaming tasks.

Benefits of technology

We have successfully trained a speech recognition model that is suitable for streaming tasks and has a high recognition accuracy. It has a certain degree of versatility and can be applied to both streaming and non-streaming tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246613B_ABST
    Figure CN116246613B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a model training method, a speech recognition method, a device, an electronic device and a storage medium. The model training method comprises: obtaining a pre-training model and an original recognition model for speech recognition, the symmetric convolution network in the pre-training model comprising a first convolution kernel corresponding to a current time or a past time before the current time, and the original recognition model comprising a causal convolution network; pre-training the pre-training model based on a plurality of first audio data; for each first convolution kernel, determining a second convolution kernel in the causal convolution network corresponding to the first convolution kernel, and initializing a second parameter of the second convolution kernel based on a first parameter of the pre-trained first convolution kernel to obtain an initialized model; and training the initialized model based on a plurality of training samples to obtain a speech recognition model. The technical solution of the embodiments of the present application trains a speech recognition model applicable to streaming tasks and having a high recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and in particular to a model training, speech recognition method, device, electronic device and storage medium. Background Technology

[0002] With the rapid development of artificial intelligence, automatic speech recognition (ASR) technology based on end-to-end deep neural networks has been widely used in various business scenarios such as e-commerce, logistics and finance.

[0003] In the process of realizing this invention, the inventors discovered the following technical problems in the prior art: it is currently impossible to train a speech recognition model that is applicable to streaming tasks and has a high recognition accuracy. Summary of the Invention

[0004] This invention provides a model training, speech recognition method, device, electronic device, and storage medium to train a speech recognition model that is applicable to streaming tasks and has a high recognition accuracy.

[0005] According to one aspect of the present invention, a model training method is provided, which may include:

[0006] Obtain a pre-trained model and an original recognition model for speech recognition, wherein the pre-trained model includes a symmetric convolutional network, the symmetric convolutional network includes a first convolutional kernel corresponding to the current time or a past time before the current time, and the original recognition model includes a causal convolutional network.

[0007] The pre-trained model is pre-trained based on multiple first audio data sets;

[0008] For each first convolutional kernel, determine the second convolutional kernel corresponding to the first convolutional kernel in the causal convolutional network, and initialize the second parameters of the second convolutional kernel based on the first parameters of the pre-trained first convolutional kernel to obtain the initialized model;

[0009] The initialization model is trained based on multiple sets of training samples to obtain a speech recognition model. The training samples include the second audio data and the corresponding audio annotation results.

[0010] According to another aspect of the present invention, a speech recognition method is provided, which may include:

[0011] Acquire the audio data to be recognized, and the speech recognition model trained according to the model training method provided in any embodiment of the present invention;

[0012] The audio data to be recognized is input into the speech recognition model, and the audio recognition result of the audio data to be recognized is obtained based on the output of the speech recognition model.

[0013] According to another aspect of the present invention, a model training apparatus is provided, which may include:

[0014] The original recognition model acquisition module is used to acquire a pre-trained model and an original recognition model for speech recognition. The pre-trained model includes a symmetric convolutional network, which includes a first convolutional kernel corresponding to the current time or a past time before the current time. The original recognition model includes a causal convolutional network.

[0015] The model pre-training module is used to pre-train the pre-trained model based on multiple first audio data.

[0016] The initialization model module is used to determine the second convolution kernel corresponding to the first convolution kernel in the causal convolutional network for each first convolution kernel, and to initialize the second parameters of the second convolution kernel based on the first parameters of the pre-trained first convolution kernel to obtain the initialization model;

[0017] The speech recognition model acquisition module is used to train the initialization model based on multiple sets of training samples to obtain the speech recognition model. The training samples include second audio data and audio annotation results corresponding to the second audio data.

[0018] According to another aspect of the present invention, a voice recognition device is provided, which may include:

[0019] The speech recognition model acquisition module is used to acquire the audio data to be recognized and the speech recognition model trained according to the model training method provided in any embodiment of the present invention.

[0020] The audio recognition result acquisition module is used to input the audio data to be recognized into the speech recognition model, and obtain the audio recognition result of the audio data to be recognized based on the output of the speech recognition model.

[0021] According to another aspect of the present invention, an electronic device is provided, which may include:

[0022] At least one processor; and

[0023] A memory that is communicatively connected to at least one processor; wherein,

[0024] The memory stores a computer program that can be executed by at least one processor, such that when the at least one processor executes the program, it implements the model training method or speech recognition method provided in any embodiment of the present invention.

[0025] According to another aspect of the present invention, a computer-readable storage medium is provided, on which computer instructions are stored, which are used to cause a processor to execute and implement the model training method or speech recognition method provided in any embodiment of the present invention.

[0026] The technical solution of this invention involves obtaining a pre-trained model and an original recognition model for speech recognition. The pre-trained model includes a symmetric convolutional network that can achieve good pre-training results, and the symmetric convolutional network includes a first convolutional kernel corresponding to the current time or a past time before the current time. The original recognition model includes a causal convolutional network that matches the streaming task. The pre-trained model is pre-trained based on multiple first audio data (i.e., unlabeled audio data). Furthermore, since the network structures of the symmetric convolutional network and the causal convolutional network are different, in order to solve the problem that the original recognition model cannot be trained based on the pre-trained model due to the different network structures, for each first convolutional kernel, a second convolutional kernel corresponding to the first convolutional kernel in the causal convolutional network can be determined, and the second parameters of the second convolutional kernel are initialized based on the first parameters of the pre-trained first convolutional kernel, thereby obtaining an initialized model that can reflect the pre-training results. The initialized model is fine-tuned based on multiple sets of training samples (i.e., labeled audio data) to obtain the final speech recognition model. The above technical solution achieves the effect of training a speech recognition model that is applicable to streaming tasks and has a high recognition accuracy.

[0027] It should be understood that the description in this section is not intended to identify key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a flowchart of a model training method provided according to an embodiment of the present invention;

[0030] Figure 2a This is a schematic diagram of a symmetric convolutional network in a model training method provided in an embodiment of the present invention;

[0031] Figure 2b This is a schematic diagram of a causal convolutional network in a model training method provided in an embodiment of the present invention;

[0032] Figure 3 This is a flowchart of another model training method provided according to an embodiment of the present invention;

[0033] Figure 4 This is a flowchart of an optional example of another model training method provided according to an embodiment of the present invention;

[0034] Figure 5 This is a flowchart of a speech recognition method provided according to an embodiment of the present invention;

[0035] Figure 6 This is a structural block diagram of a model training device provided according to an embodiment of the present invention;

[0036] Figure 7 This is a structural block diagram of a speech recognition device according to an embodiment of the present invention;

[0037] Figure 8 This is a schematic diagram of the structure of an electronic device that implements the model training method or speech recognition method of the embodiments of the present invention. Detailed Implementation

[0038] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0039] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The same applies to "target," "original," etc., and will not be repeated here. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0040] Before introducing the embodiments of the present invention, the application scenarios of the embodiments of the present invention will be illustrated by way of example: In order to train a speech recognition model with high recognition accuracy, one option is to rely on a large amount of labeled audio data; another option is to first pre-train the model based on a large amount of unlabeled audio data, and then fine-tune the pre-trained model based on a small amount of labeled audio data. The former has too high a labeling cost, and the recognition accuracy of the streaming ASR model and non-streaming ASR model trained in this way is not very ideal; the latter has a lower labeling cost, and can train a non-streaming ASR model with relatively ideal recognition accuracy, but the recognition accuracy of the trained streaming ASR model is not very ideal, that is, the application effect on streaming ASR models is not good. The specific reasons are as follows:

[0041] When non-streaming ASR models compute audio features at each time step, they need to utilize both past and future audio features of that time step (i.e., the current time step). Therefore, the convolutional layers in the encoder of a non-streaming ASR model are symmetric convolutional networks. In contrast, when streaming ASR models compute audio features at each time step, they only need to utilize the audio features of the past time step and do not need to utilize the audio features of the future time step. Therefore, the convolutional layers in the encoder of a streaming ASR model are causal convolutional networks.

[0042] Based on this, since the pre-training mechanism relies on global features for comparative learning to ensure the recognition accuracy of the model, the convolutional layer in the encoder of the pre-trained model needs to be implemented using a symmetric convolutional network. This means that the pre-trained model obtained from the pre-training is different from the model structure of the streaming ASR model, making it difficult to fine-tune the training based on the pre-trained model to obtain the streaming ASR model.

[0043] Alternatively, to address the difficulty of fine-tuning training due to different model structures, causal symmetric networks can be used in pre-trained models. However, as mentioned above, since causal symmetric networks cannot rely on global features for comparative learning, the effectiveness of the pre-training mechanism will be significantly reduced, thus failing to obtain a streaming ASR model with high recognition accuracy.

[0044] In summary, it is currently impossible to train a speech recognition model that is applicable to streaming tasks and has a high recognition accuracy (i.e., a streaming ASR model with high recognition accuracy).

[0045] Figure 1This is a flowchart of a model training method provided in an embodiment of the present invention. This embodiment is applicable to training a speech recognition model that is suitable for streaming tasks and has a high recognition accuracy. The method can be executed by the model training device provided in this embodiment of the present invention. The device can be implemented by software and / or hardware and can be integrated into an electronic device, which may include various user terminals or servers.

[0046] See Figure 1 The method of this invention specifically includes the following steps:

[0047] S110. Obtain a pre-trained model and an original recognition model for speech recognition. The pre-trained model includes a symmetric convolutional network, which includes a first convolutional kernel corresponding to the current time or a past time before the current time. The original recognition model includes a causal convolutional network.

[0048] The pre-trained model can be a machine learning model to be pre-trained for speech recognition, which may include at least a symmetric convolutional network. This symmetric convolutional network may include multiple convolutional kernels, some of which may correspond to the current time step, some to past time steps before the current time step, and some to future time steps after the current time step. Here, the convolutional kernel corresponding to the current time step or a past time step is referred to as the first convolutional kernel; that is, at least one of these first convolutional kernels corresponds to the current time step, and the remaining first convolutional kernels (excluding the at least one first convolutional kernel) correspond to past time steps. It should be noted that the current time step is not a fixed time step. As illustrated in the example above, the pre-trained model needs to calculate the audio features at each time step, so the time step it is currently calculating can be considered the current time step. Furthermore, the convolutional kernel corresponding to the current time step can be understood as the convolutional kernel among the multiple convolutional kernels used to calculate the audio features related to the current time step. The meaning of the convolutional kernel corresponding to past time steps is similar and will not be elaborated further here.

[0049] The original recognition model can be a machine learning model to be fine-tuned for speech recognition, which may include at least a causal convolutional network that does not need to be associated with audio features from future times, thus adapting to streaming tasks in speech recognition. This causal convolutional network may include multiple convolutional kernels, some of which may correspond to the current time and some to past times. It should be noted that for each first convolutional kernel, the convolutional kernel corresponding to the first convolutional kernel in the causal convolutional network is called the second convolutional kernel; in other words, each first convolutional kernel may have its own corresponding second convolutional kernel.

[0050] In practical applications, optionally, taking symmetric convolution as an example, this symmetric convolutional network can be considered as part or all of the encoder in the pre-trained model. The encoder here can be an encoder based on Transformer or Conformer. Compared with Transformer, Conformer can combine global feature extraction of the attention mechanism with local feature extraction of the convolutional layer, thereby obtaining audio feature representation that takes into account multiple granularities to ensure the recognition accuracy of the model.

[0051] S120. Pre-train the pre-trained model based on multiple first audio data.

[0052] The first audio data can be pre-obtained unlabeled audio data. In practical applications, it can optionally be the first original audio obtained by direct sampling, or it can be the first audio features obtained after feature extraction from the first original audio. No specific limitation is made here. The pre-trained model is pre-trained based on multiple first audio data. Considering the application scenarios that may be involved in the embodiments of the present invention, the pre-training here can optionally be unsupervised training based on a contrastive learning strategy. On this basis, the contrastive learning loss function in the pre-training process can optionally refer to the pre-training scheme of wav2vec2.0. Furthermore, when implementing the encoder based on Conformer in the embodiments of the present invention, Conformer can be used to replace Transformer used in wav2vec2.0.

[0053] It should be noted that since the convolutional layers in the encoder of the pre-trained model are represented by symmetric convolutional networks, the symmetric convolutional network can improve the receptive field of the model features during the pre-training process based on multiple unlabeled audio data, thereby making full use of large-scale unlabeled audio data and thus improving the pre-trained model's ability to represent unlabeled audio data.

[0054] S130. For each first convolutional kernel, determine the second convolutional kernel in the causal convolutional network corresponding to the first convolutional kernel, and initialize the second parameters of the second convolutional kernel based on the first parameters of the pre-trained first convolutional kernel to obtain the initialized model.

[0055] One issue is that the symmetric convolutional network in the pre-trained model and the causal convolutional network in the original recognition model have different network structures, which prevents the pre-training results from being directly applied to the original recognition model. To address this, considering a key difference between symmetric and causal convolutional networks—that when calculating audio features at the current time, symmetric convolutional networks apply convolutional kernels corresponding to future timeframes, while causal convolutional networks either lack or do not apply convolutional kernels corresponding to future timeframes—a new approach is adopted. For each first convolutional kernel, a corresponding second convolutional kernel can be determined first, and then the second parameters of the second convolutional kernel can be initialized based on the first parameters of the pre-trained first convolutional kernel. In other words, at the convolutional layer level, only the first parameters of the first convolutional kernel corresponding to the current or future timeframe in the symmetric convolutional network need to be loaded into the causal convolutional network; the kernel parameters of the convolutional kernel corresponding to the future timeframe in the symmetric convolutional network do not need to be loaded. This solves the problem of not being able to apply the pre-training results to streaming tasks due to the inconsistency in their network structures.

[0056] Based on the above steps, each first convolutional kernel is processed, thereby loading the first parameters of each pre-trained first convolutional kernel onto the causal convolutional network of the original recognition model, thus obtaining the initialized model.

[0057] S140. The initialization model is trained based on multiple sets of training samples to obtain a speech recognition model. The training samples include the second audio data and the audio annotation results corresponding to the second audio data.

[0058] Each training sample includes second audio data and audio annotation results corresponding to the second audio data. The initial model can be trained based on multiple training samples, that is, the initial model can be trained based on the annotated audio data (i.e., fine-tuning training) to obtain the final speech recognition model.

[0059] In practical applications, the loss function optionally used during the fine-tuning training phase can be the Connectionist Temporal Classification (CTC) loss function, for the i-th second audio data x. i The audio recognition result e output by the initial model can be obtained. i Combined with the corresponding audio annotation results y i The obtained CTC loss function l ctc This can be expressed by the following formula:

[0060]

[0061] Wherein, φ(y) i ,ei ) for e i All decoding yields y i The set of CTC paths, p(π|e i ) represents the probability of the CTC path π.

[0062] Based on this, optionally, in the task of optimizing the CTC loss function, the adaptive moment estimation (ADAM) optimization algorithm can be used. ADAM adopts a warm-up strategy with a peak learning rate of 0.001 and a warm-up strategy step size of 15,000 steps, and trains the initial model until convergence.

[0063] It should be noted that since the speech recognition model is obtained by fine-tuning the initial model, and the convolutional layers in the initial model are represented by causal convolutional networks, this means that the speech recognition model trained in this way can be applied not only to streaming tasks but also to non-streaming tasks. That is, it can be used as a streaming ASR model or a non-streaming ASR model, and has a certain degree of versatility.

[0064] The technical solution of this invention involves obtaining a pre-trained model and an original recognition model for speech recognition. The pre-trained model includes a symmetric convolutional network that can achieve good pre-training results, and the symmetric convolutional network includes a first convolutional kernel corresponding to the current time or a past time before the current time. The original recognition model includes a causal convolutional network that matches the streaming task. The pre-trained model is pre-trained based on multiple first audio data (i.e., unlabeled audio data). Furthermore, since the network structures of the symmetric convolutional network and the causal convolutional network are different, in order to solve the problem that the original recognition model cannot be trained based on the pre-trained model due to the different network structures, for each first convolutional kernel, a second convolutional kernel corresponding to the first convolutional kernel in the causal convolutional network can be determined, and the second parameters of the second convolutional kernel are initialized based on the first parameters of the pre-trained first convolutional kernel, thereby obtaining an initialized model that can reflect the pre-training results. The initialized model is fine-tuned based on multiple sets of training samples (i.e., labeled audio data) to obtain the final speech recognition model. The above technical solution achieves the effect of training a speech recognition model that is applicable to streaming tasks and has a high recognition accuracy.

[0065] An optional technical solution includes a pre-trained model further comprising a first recognition network, and an original recognition model further comprising a second recognition network. Before obtaining the initialized model, the above model training method may further include: initializing the fourth parameters of the second recognition network based on the third parameters of the pre-trained first recognition network. The first recognition network can be understood as the network in the pre-trained model used for speech recognition or speech classification; in practical applications, it can optionally be represented by a fully connected layer. The second recognition network is similar to the first recognition network and will not be described further here. To effectively load the third parameters of the pre-trained first recognition network onto the second recognition network, the network structures of the two recognition networks can be identical. That is, in addition to initializing the second parameters based on the first parameters, the fourth parameters of the second recognition network can also be initialized based on the third parameters, thereby allowing the resulting initialized model to also possess the speech recognition functionality of the pre-trained model. It's important to note that the reason we can directly load all the third parameters into the second recognition network, instead of just loading a portion of the first parameters of the convolutional kernel (i.e., the first convolutional kernel) as in convolutional layers, is that the third parameters are not directly related to streaming and non-streaming tasks, or rather, the processing flow is the same for both. Of course, if the pre-trained model includes other parameters besides the third parameter outside the convolutional kernel level, they can also be directly loaded into the original recognition model, just like the third parameter, so that the resulting initialized model has the corresponding functionality of the pre-trained model.

[0066] Another optional technical solution is that the symmetric convolutional network may include an odd number of convolutional kernels arranged sequentially. The first convolutional kernel corresponding to the current time step is the convolutional kernel located in the middle position among the odd number of convolutional kernels, and the first convolutional kernel corresponding to past time steps is the convolutional kernel located before or after the middle position among the odd number of convolutional kernels. For example, assuming that the symmetric convolutional network includes 2k+1 convolutional kernels, where k is a positive integer, then the first convolutional kernel corresponding to the current time step is the convolutional kernel located in the (k+1)th position (i.e., the middle position), and the first convolutional kernel corresponding to past time steps is the convolutional kernel located in the first k positions or the last k positions. Based on this, and considering the application scenarios that may be involved in the embodiments of the present invention, in order to more vividly understand this technical solution, the following is an exemplary description with specific examples.

[0067] For example, see Figure 2aThe symmetric convolutional network shown (taking k=3 as an example) has the first k convolutional kernels corresponding to past time points, the (k+1)th convolutional kernel corresponding to the current time point, and the last k convolutional kernels corresponding to future time points. Based on this, during the convolution calculation, the audio features at each time point of the Nth layer can be calculated using a convolutional kernel of size 2k+1 and 2k+1 audio features from the (N-1)th layer. These 2k+1 audio features consist of the audio features from the k past time points, the k audio features from the k future time points, and the audio features from the current time point. See also... Figure 2b The causal convolutional network shown can be initialized using the left half of a symmetric convolutional network. Specifically, the kernel size in the causal convolutional network is k+1 (or 2k+1, by setting the parameters of the last k kernels to 0 to disable them). During parameter loading, the first parameters of the first k+1 kernels from the 2k+1 kernels of the symmetric convolutional network can be loaded into the causal convolutional network. Figure 2b In the N-1 layer, white boxes represent convolutional kernels that have been loaded with parameters, and gray boxes represent convolutional kernels that have not been loaded with parameters or do not exist.

[0068] Another optional technical solution, the above model training method, may further include: acquiring a first original audio, extracting features from the first original audio to obtain first audio features, and using the first audio features as first audio data; acquiring a second original audio, extracting features from the second original audio to obtain second audio features, and using the second audio features as second audio data; and using the second audio data and the audio annotation results corresponding to the second original audio as a set of training samples. Considering the application scenarios that may be involved in the embodiments of this invention, the first and second audio features here can be 80-dimensional Mel-spectral features, where the duration of each frame can be 25ms and the step size can be 10ms. Of course, the above is only an example and is not a specific limitation on the first and second audio features. Furthermore, the set U composed of multiple first audio features can be represented by U = {s} i |i∈[1,N U The set L, consisting of multiple training samples, can be represented as L = {x}. i ,y i |i∈[1,N L ]} is used to represent, where s i and x i Let y represent the i-th audio feature in the corresponding set, respectively. i For audio features x i The corresponding audio annotation results (which can also be called text annotation results), N U and N LThese represent the number of audio features in the corresponding set. Typically, considering the cost of manual annotation, N... U >>N L .

[0069] Before introducing the embodiments of the present invention, the application scenarios of the embodiments of the present invention will be described by way of example. For example, when training streaming ASR models and non-streaming ASR models based on pre-training mechanisms, one option is to optimize them separately, that is, to perform fine-tuning training on the non-streaming ASR model after non-streaming pre-training, and to perform fine-tuning training on the streaming ASR model after streaming pre-training. However, the ASR models trained based on the above schemes are difficult to have good versatility on both streaming and non-streaming tasks, which is a technical problem that urgently needs to be solved.

[0070] Figure 3 This is a flowchart of another model training method provided in this embodiment of the invention. This embodiment is based on the above-mentioned technical solutions and optimized. In this embodiment, optionally, the original recognition model further includes an attention layer connected to a causal convolutional network; training the initial model based on multiple sets of training samples to obtain a speech recognition model includes: for each set of training samples, inputting the second audio data in the training samples into the attention layer to obtain attention features; determining the model training task corresponding to the training samples and obtaining the mask features matching the model training task, wherein the model training task is used to train a speech recognition model applicable to streaming or non-streaming tasks in speech recognition tasks; obtaining target features based on the attention features and mask features; inputting the target features into the causal convolutional network to combine the output of the causal convolutional network and the audio annotation results in the training samples to train a speech recognition model. The explanations of terms that are the same as or corresponding to those in the above embodiments are not repeated here.

[0071] See Figure 3 The method in this embodiment may specifically include the following steps:

[0072] S210. Obtain a pre-trained model and an original recognition model for speech recognition. The pre-trained model includes a symmetric convolutional network, which includes a first convolutional kernel corresponding to the current time or a past time before the current time. The original recognition model includes a causal convolutional network and an attention layer connected to the causal convolutional network.

[0073] The original recognition model may include interconnected attention layers and causal convolutional networks. These attention layers and causal convolutional networks can be understood as part or all of the encoder described above. In other words, the original recognition model may include an encoder, which may include interconnected attention layers and causal convolutional networks.

[0074] S220. Pre-train the pre-trained model based on multiple first audio data.

[0075] S230. For each first convolutional kernel, determine the second convolutional kernel in the causal convolutional network corresponding to the first convolutional kernel, and initialize the second parameters of the second convolutional kernel based on the first parameters of the pre-trained first convolutional kernel to obtain the initialized model.

[0076] S240. For each training sample, input the second audio data in the training sample into the attention layer to obtain attention features. The training sample includes the second audio data and the audio annotation results corresponding to the second audio data.

[0077] S250. Determine the model training task corresponding to the training samples and obtain the mask features that match the model training task. The model training task is used to train a speech recognition model that can be applied to streaming or non-streaming tasks in speech recognition.

[0078] Each training sample corresponds to its own model training task. The model training tasks corresponding to any two training samples may be the same or different, without specific limitations. For each training sample, the corresponding model training task can indicate whether the training sample is used to train a speech recognition model suitable for streaming tasks or a speech recognition model suitable for non-streaming tasks.

[0079] Mask features can be understood as features used to help training samples complete the corresponding model training tasks. The values ​​can be 0 and / or 1. The specific value setting matches the model training task to be completed, and no specific limitation is made here.

[0080] In practical applications, optionally, when the model training task is a non-streaming task, the mask features can be represented by a first mask matrix, in which each first element is 1. After being combined with the attention features, the subsequent causal convolutional network can be applied to the audio features at each time step, thereby training a speech recognition model that is applicable to non-streaming tasks.

[0081] Alternatively, when the model training task is a streaming task, the mask features can be represented by a second mask matrix. Specifically, for multiple second elements in each row of the second mask matrix, the second element corresponding to the target number is 1, and all other elements are 0. The target number can be greater than or equal to the row number of the multiple second elements in the second mask matrix, and the remaining elements correspond to future times after the current time. In other words, the second elements corresponding to some or all future times are 0, and the second elements corresponding to at least the current time and all past times are 1. This means that the second elements may have 1 corresponding to some future times and 0 corresponding to some future times, or they may have all 0 corresponding to all future times, depending on the specific value of the target number, which is not specifically limited here. For example, when the target number is equal to the row number, the second elements corresponding to all future times are 0; when the target number is greater than the row number, the second elements corresponding to some future times are 0 and the second elements corresponding to some future times are 1. In this way, by applying the second mask matrix to the attention features, the subsequent causal convolutional network can be applied to the audio features corresponding to the current time and all past time, but cannot be applied to or can only be applied to some audio features corresponding to future time. This limits the observation range of the attention mechanism for audio features of future time, thereby training a speech recognition model that can be applied to streaming tasks.

[0082] S260. Obtain the target features based on attention features and mask features.

[0083] This process involves combining attention features and mask features to process the attention features based on the mask features, thereby obtaining target features (i.e., processed attention features) that match the model training task. In practical applications, optionally, when both attention features and mask features are represented by matrices, their product can be used as the target feature.

[0084] S270. Input the target features into a causal convolutional network to combine the output of the causal convolutional network with the audio annotation results in the training samples to train a speech recognition model.

[0085] The target features are input into a causal convolutional network, which combines the output of the causal convolutional network with the audio annotations from the training samples to train a speech recognition model. For example, the output of the causal convolutional network can be input into a second recognition network to obtain the audio recognition result; then, a pre-set loss function is used to adjust the network parameters in the original recognition model based on the audio recognition result and the audio annotation result, thereby training the speech recognition model.

[0086] It should be noted that during the fine-tuning training process, a portion of the training samples were used to train a speech recognition model applicable to streaming tasks, while another portion was used to train a speech recognition model applicable to non-streaming tasks. In other words, fine-tuning training for both streaming and non-streaming tasks was carried out. This allows the final trained speech recognition model to be applied to both streaming and non-streaming tasks, demonstrating good versatility.

[0087] The technical solution of this invention involves, for each set of training samples, inputting the second audio data in the training samples into an attention layer to obtain attention features; then, determining whether the model training task corresponding to the training samples is for training a speech recognition model applicable to streaming or non-streaming tasks in speech recognition, and obtaining mask features matching the model training task; then, obtaining target features based on the attention features and mask features, and inputting the target features into a causal convolutional network, thereby combining the output of the causal convolutional network and the audio annotation results in the training samples to train a speech recognition model that is universal for both streaming and non-streaming tasks.

[0088] An optional technical solution for determining the model training task corresponding to a training sample includes: acquiring random data following a uniform distribution, defined based on two pre-set distribution parameters, and parameter thresholds pre-set according to the two distribution parameters; determining the model training task corresponding to the training sample based on the numerical relationship between the random data and the parameter thresholds. Here, the two distribution parameters can be understood as the minimum and maximum data among the random data following a uniform distribution, and the parameter thresholds can be thresholds determined based on these two distribution parameters, such as using the mean of the two distribution parameters as the parameter threshold, or using a data point closer to the minimum or maximum data as the parameter threshold, etc., without specific limitations. When fine-tuning training based on any training sample, random data following a uniform distribution can be acquired for that training sample, thereby determining the model training task corresponding to the training sample based on the numerical relationship between the random data and the parameter thresholds. For example, when the random data is less than or equal to the parameter threshold, the model training task can be the task of training a speech recognition model applicable to churn tasks; otherwise, the model training task can be the task of training a speech recognition model applicable to non-churn tasks. Of course, the reverse is also possible, without specific limitations. The above technical solution can be understood as a process of randomly fine-tuning training for streaming and non-streaming tasks using dynamic window masks.

[0089] Building upon this, to more vividly illustrate how to train a speech recognition model with good generality, a specific example is provided below. For instance, assuming the number of frames for the second audio feature is D, the attention feature obtained after inputting this second audio feature into the attention layer can be represented by a D*D attention matrix. Furthermore, assuming the two distribution parameters are 0 and 1, and the parameter threshold is 0.5, then when the random data obtained for the second audio feature is greater than 0.5, the mask feature can be represented by a D*D mask matrix where all first elements are 1. In this case, the attention mechanism is global attention, thus enabling fine-tuning training for non-dropout tasks. Correspondingly, when the random data obtained for the second audio feature is less than or equal to 0.5, the mask feature can be represented by a D*D second mask matrix. For each row d (1<=d<=D) in the second mask matrix, the first d+T second elements are 1 and the last DdT second elements are 0, where T is a random integer representing the maximum streaming latency from 0ms to 1s. In this case, the attention mechanism is local attention, allowing for fine-tuning training of the churn task. In other words, during the calculation of the attention mechanism, the window of the attention mechanism has a 50% probability of representing the second audio feature at all times, and a 50% probability of seeing the second audio feature within T frames in the future. This enables the random execution of the fine-tuning training process for both churn and non-churn tasks, thereby improving the versatility of the trained speech recognition model.

[0090] To better understand the various technical solutions described above, specific examples are provided below for illustration. For examples, see [link to example]. Figure 4 :

[0091] 1) Training data preparation: Calculate the first audio features using unlabeled original audio (i.e., the first original audio), and calculate the second audio features using labeled original audio (i.e., the second original audio).

[0092] 2) Model pre-training: The pre-trained model is pre-trained using the first audio features. The Conformer in the pre-trained model is implemented based on a symmetric convolutional network.

[0093] 3) Parameter initialization: Construct the original recognition model and load the first and third parameters from the pre-trained model into the original recognition model. The Conformer in the original recognition model is implemented based on a causal convolutional network.

[0094] 4) Fine-tuning training of the general model: The dynamic window mask method is used to fine-tune the training for streaming and non-streaming tasks, so as to obtain a speech recognition model with good generality.

[0095] To verify the effectiveness of the above scheme, relevant experiments were conducted, and the results are as follows. Pre-training was performed on 100,000 hours of unlabeled raw audio, followed by fine-tuning on 100 hours of labeled raw audio. On the test machine, the word error rates for non-streaming and streaming tests were 12.35% and 13.07%, respectively, significantly better than the speech recognition model trained directly on the 100 hours of labeled raw audio, which had word error rates of 14.56% and 15.70% in non-streaming and streaming tests, respectively. Furthermore, the experimental results show that the speech recognition performance obtained using the above scheme, with a streaming loss of only 0.72% compared to the non-streaming loss, is significantly better than the 1.14% streaming loss under supervised learning. This indicates that the generalizability of the speech recognition model trained based on the above scheme is significantly improved.

[0096] Figure 5 This is a flowchart of a speech recognition method provided in an embodiment of the present invention. This embodiment is applicable to situations requiring accurate speech recognition under churn tasks. The method can be executed by the speech recognition device provided in this embodiment of the present invention. This device can be implemented in software and / or hardware, and can be integrated into an electronic device, which can be various user terminals or servers.

[0097] See Figure 5 The method of this invention specifically includes the following steps:

[0098] S310. Obtain the audio data to be recognized and the speech recognition model trained according to the model training method provided in any embodiment of the present invention.

[0099] The audio data to be recognized can be the audio data to be recognized, such as the raw audio data directly collected, or the audio features to be recognized obtained after feature extraction from the raw audio data; no specific limitation is made here. The speech recognition model can be a machine learning model that has been trained and is capable of accurately performing speech recognition under churn tasks.

[0100] S320. Input the audio data to be recognized into the speech recognition model, and obtain the audio recognition result of the audio data to be recognized based on the output of the speech recognition model.

[0101] It should be noted that the audio data to be recognized can be real-time audio data, in which case the speech recognition model is used for churned task speech recognition; the audio data to be recognized can also be offline audio data, in which case the speech recognition model is used for non-churned task speech recognition.

[0102] The technical solution of this invention, by inputting the audio data to be recognized into a speech recognition model, can obtain the audio recognition result of the audio data to be recognized, thereby achieving accurate speech recognition under churned tasks, and of course, also achieving accurate speech recognition under non-churned tasks.

[0103] Figure 6 This is a structural block diagram of a model training apparatus provided in an embodiment of the present invention. This apparatus is used to execute the model training method provided in any of the above embodiments. This apparatus and the model training methods of the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the model training apparatus can be referred to the embodiments of the above model training methods. See also... Figure 6 The device may specifically include: an original recognition model acquisition module 410, a model pre-training module 420, an initialization model acquisition module 430, and a speech recognition model acquisition module 440.

[0104] The original recognition model acquisition module 410 is used to acquire a pre-trained model and an original recognition model for speech recognition. The pre-trained model includes a symmetric convolutional network, which includes a first convolutional kernel corresponding to the current time or a past time before the current time. The original recognition model includes a causal convolutional network.

[0105] The model pre-training module 420 is used to pre-train the pre-trained model based on multiple first audio data.

[0106] The initialization model obtains module 430, which is used to determine the second convolution kernel in the causal convolutional network corresponding to the first convolution kernel for each first convolution kernel, and initialize the second parameters of the second convolution kernel based on the first parameters of the pre-trained first convolution kernel to obtain the initialization model;

[0107] The speech recognition model acquisition module 440 is used to train the initialization model based on multiple sets of training samples to obtain the speech recognition model. The training samples include the second audio data and the audio annotation results corresponding to the second audio data.

[0108] Optionally, the pre-trained model further includes a first recognition network, and the original recognition model further includes a second recognition network; the above-mentioned model training device may further include:

[0109] The parameter initialization module is used to initialize the fourth parameter of the second recognition network based on the third parameter of the pre-trained first recognition network before obtaining the initialization model.

[0110] Optionally, the symmetric convolutional network may include an odd number of convolutional kernels arranged in sequence. The first convolutional kernel corresponding to the current time step is the convolutional kernel that is in the middle position among the odd number of convolutional kernels, and the first convolutional kernel corresponding to the past time step is the convolutional kernel that is in the position before or after the middle position among the odd number of convolutional kernels.

[0111] Optionally, the original recognition model also includes an attention layer connected to a causal convolutional network; the speech recognition model module 440 may include:

[0112] The attention feature acquisition unit is used to input the second audio data in the training samples into the attention layer for each training sample to obtain the attention features;

[0113] The mask feature acquisition unit is used to determine the model training task corresponding to the training sample and obtain the mask features matching the model training task. The model training task is used to train a speech recognition model that can be applied to streaming or non-streaming tasks in speech recognition tasks.

[0114] The target feature acquisition unit is used to obtain target features based on attention features and mask features;

[0115] The speech recognition model obtains a unit, which is used to input the target features into the causal convolutional network, and combine the output of the causal convolutional network with the audio annotation results in the training samples to train the speech recognition model.

[0116] Based on this, an optional mask feature obtaining unit may include:

[0117] The parameter threshold acquisition subunit is used to acquire random data that follows a uniform distribution and parameter thresholds that are pre-set based on two pre-defined distribution parameters, for a uniform distribution defined by two pre-defined distribution parameters.

[0118] The model training task determination subunit is used to determine the model training task corresponding to the training samples based on the numerical relationship between random data and parameter thresholds.

[0119] Alternatively, if the model training task is a non-streaming task, the mask features are represented by a first mask matrix, where each first element of the first mask matrix is ​​1.

[0120] And / or,

[0121] When the model training task is a streaming task, the mask features are represented by the second mask matrix. For each row of the second mask matrix, the second element with the number of targets is 1, and the other elements are 0. The number of targets is greater than or equal to the row number of the second elements in the second mask matrix. The other elements correspond to future times after the current time.

[0122] Optionally, the above-mentioned model training device may further include:

[0123] The first audio data acquisition module is used to acquire the first raw audio, extract features from the first raw audio to obtain the first audio features, and use the first audio features as the first audio data.

[0124] The second audio data acquisition module is used to acquire the second original audio, extract features from the second original audio to obtain the second audio features, and use the second audio features as the second audio data.

[0125] The training sample acquisition module is used to take the second audio data and the audio annotation results corresponding to the second original audio as a set of training samples.

[0126] The model training apparatus provided in this embodiment of the invention acquires a pre-trained model and an original recognition model for speech recognition through an original recognition model acquisition module. The pre-trained model includes a symmetric convolutional network that can achieve good pre-training results. The symmetric convolutional network includes a first convolutional kernel corresponding to the current time or a past time before the current time. The original recognition model includes a causal convolutional network that matches the streaming task. The model pre-training module pre-trains the pre-trained model based on multiple first audio data (i.e., unlabeled audio data). Furthermore, since the network structures of the symmetric convolutional network and the causal convolutional network are different, to solve the problem that the original recognition model cannot be trained based on the pre-trained model due to the different network structures, the initialization model acquisition module determines the second convolutional kernel corresponding to the first convolutional kernel in the causal convolutional network for each first convolutional kernel, and initializes the second parameters of the second convolutional kernel based on the first parameters of the pre-trained first convolutional kernel, thereby obtaining an initialization model that can reflect the pre-training results. The speech recognition model acquisition module fine-tunes the initialization model based on multiple sets of training samples (i.e., labeled audio data) to obtain the final speech recognition model. The aforementioned device achieves the effect of training a speech recognition model that is applicable to streaming tasks and has a high recognition accuracy.

[0127] The model training apparatus provided in this embodiment of the invention can execute the model training method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0128] It is worth noting that in the embodiments of the above-mentioned model training device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0129] Figure 7 This is a structural block diagram of a speech recognition device provided in an embodiment of the present invention. This device is used to execute the speech recognition method provided in any of the above embodiments. This device and the speech recognition methods of the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the speech recognition device can be found in the embodiments of the above speech recognition methods. See also... Figure 7 The device may specifically include: a speech recognition model acquisition module 510 and an audio recognition result acquisition module 520.

[0130] The speech recognition model acquisition module 510 is used to acquire the audio data to be recognized and the speech recognition model trained according to the model training method provided in any embodiment of the present invention.

[0131] The audio recognition result acquisition module 520 is used to input the audio data to be recognized into the speech recognition model, and obtain the audio recognition result of the audio data to be recognized based on the output result of the speech recognition model.

[0132] The speech recognition device provided in this embodiment of the invention uses a speech recognition model acquisition module and an audio recognition result acquisition module to cooperate with each other to input the audio data to be recognized into the speech recognition model, thereby obtaining the audio recognition result of the audio data to be recognized. This achieves accurate speech recognition under churned tasks, and of course, it also achieves accurate speech recognition under non-churned tasks.

[0133] The speech recognition device provided in the embodiments of the present invention can execute the speech recognition method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0134] It is worth noting that in the embodiments of the above-mentioned voice recognition device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0135] Figure 8A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0136] like Figure 8 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0137] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0138] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as model training methods or speech recognition methods.

[0139] In some embodiments, the model training method or speech recognition method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the model training method or speech recognition method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the model training method or speech recognition method by any other suitable means (e.g., by means of firmware).

[0140] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0141] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0142] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0143] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0144] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0145] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0146] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0147] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A model training method, characterized in that, include: Obtain a pre-trained model and an original recognition model for speech recognition, wherein the pre-trained model includes a symmetric convolutional network, the symmetric convolutional network includes a first convolutional kernel corresponding to the current time or a past time before the current time, and the original recognition model includes a causal convolutional network; The pre-trained model is pre-trained based on multiple first audio data; For each of the first convolutional kernels, a second convolutional kernel corresponding to the first convolutional kernel in the causal convolutional network is determined, and the second parameters of the second convolutional kernel are initialized based on the first parameters of the pre-trained first convolutional kernel to obtain an initialized model; The initialization model is trained based on multiple sets of training samples to obtain a speech recognition model, wherein the training samples include second audio data and the audio annotation results corresponding to the second audio data.

2. The method according to claim 1, characterized in that, The pre-trained model further includes a first recognition network, and the original recognition model further includes a second recognition network; Prior to obtaining the initialized model, the method further includes: The fourth parameters of the second recognition network are initialized based on the third parameters of the first recognition network after pre-training.

3. The method according to claim 1, characterized in that, The symmetric convolutional network includes an odd number of convolutional kernels arranged in sequence. The first convolutional kernel corresponding to the current time step is the convolutional kernel that is arranged in the middle position among the odd number of convolutional kernels. The first convolutional kernel corresponding to the past time step is the convolutional kernel that is arranged before or after the middle position among the odd number of convolutional kernels.

4. The method according to claim 1, characterized in that, The original recognition model also includes an attention layer connected to the causal convolutional network; The process of training the initialization model based on multiple sets of training samples to obtain a speech recognition model includes: For each set of training samples, the second audio data in the training samples is input into the attention layer to obtain attention features; The model training task corresponding to the training sample is determined, and the mask features matched by the model training task are obtained. The model training task is used to train a speech recognition model that can be applied to streaming or non-streaming tasks in speech recognition. The target features are obtained based on the attention features and the mask features; The target features are input into the causal convolutional network, and the speech recognition model is trained by combining the output of the causal convolutional network with the audio annotation results in the training samples.

5. The method according to claim 4, characterized in that, Determining the model training task corresponding to the training sample includes: For a uniform distribution defined based on two pre-set distribution parameters, obtain random data that follows the uniform distribution and parameter thresholds pre-set based on the two distribution parameters; Based on the numerical relationship between the random data and the parameter threshold, the model training task corresponding to the training sample is determined.

6. The method according to claim 4, characterized in that, When the model training task is the non-streaming task, the mask feature is represented by a first mask matrix, where each first element of the first mask matrix is ​​1. And / or, When the model training task is the streaming task, the mask feature is represented by a second mask matrix. For each row of the second mask matrix, the second element of the target number is 1, and the remaining elements other than the target number of the second elements are 0. The target number is greater than or equal to the row number of the multiple second elements in the second mask matrix. The remaining elements correspond to future times after the current time.

7. The method according to claim 1, characterized in that, Also includes: A first original audio file is acquired, and features are extracted from the first original audio file to obtain a first audio feature. The first audio feature is then used as the first audio data. Acquire the second original audio and extract features from it to obtain the second audio features, and use the second audio features as the second audio data; The second audio data and the audio annotation results corresponding to the second original audio are used as a set of training samples.

8. A speech recognition method, characterized in that, include: Acquire the audio data to be recognized and the speech recognition model trained according to the model training method of any one of claims 1-7; The audio data to be recognized is input into the speech recognition model, and the audio recognition result of the audio data to be recognized is obtained based on the output result of the speech recognition model.

9. A model training device, characterized in that, include: The original recognition model acquisition module is used to acquire a pre-trained model and an original recognition model for speech recognition. The pre-trained model includes a symmetric convolutional network, which includes a first convolutional kernel corresponding to the current time or a past time before the current time. The original recognition model includes a causal convolutional network. The model pre-training module is used to pre-train the pre-trained model based on multiple first audio data. The initialization model module is used to determine, for each first convolutional kernel, a second convolutional kernel in the causal convolutional network corresponding to the first convolutional kernel, and to initialize the second parameters of the second convolutional kernel based on the first parameters of the pre-trained first convolutional kernel, so as to obtain the initialization model; The speech recognition model acquisition module is used to train the initialization model based on multiple sets of training samples to obtain a speech recognition model, wherein the training samples include second audio data and audio annotation results corresponding to the second audio data.

10. A voice recognition device, characterized in that, include: A speech recognition model acquisition module is used to acquire audio data to be recognized and a speech recognition model trained according to the model training method of any one of claims 1-7. The audio recognition result acquisition module is used to input the audio data to be recognized into the speech recognition model, and obtain the audio recognition result of the audio data to be recognized based on the output result of the speech recognition model.

11. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to cause the at least one processor to perform the model training method as described in any one of claims 1-7, or the speech recognition method as described in claim 8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute and implement the model training method as described in any one of claims 1-7, or the speech recognition method as described in claim 8.

Citation Information

Patent Citations

  • Voice processing method and device, electronic equipment and storage medium

    CN113823272A

  • Speech recognition method and system, storage medium and terminal equipment

    CN114360503A