Speech recognition model data processing system and method, speech recognition method
Through the coordinated training of cloud and terminal devices, the speech recognition model is pre-trained using Chinese phoneme unit prediction and text prediction tasks, which solves the dependence on labeled data in Chinese speech recognition, improves training efficiency and accuracy, and adapts to the needs of Chinese speech recognition tasks.
Patent Information
- Application Number
- JP2025504340
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-27
- Filing Date
- 2023-10-24
- Publication Date
- 2025-08-20
AI Technical Summary
In the prior art, the training of speech recognition models requires a large number of marked speech text data, which leads to large consumption of manpower and material resources. The existing pre-training methods have limited effects in Chinese speech recognition, especially for ideographic text languages such as Chinese, the difference between text and speech modalities is large, making it difficult to effectively utilize unlabeled data.
The voice data is pre-trained by the implementation of the Chinese phoneme unit prediction task and text prediction task, and the encoder and decoder are trained respectively. The unlabeled Chinese voice data and text data are used to optimize the model parameters to adapt to the speech recognition task.
It reduces the demand for labeling data, reduces the cost of labeling, improves training efficiency and accuracy, makes the model more adaptable to downstream speech recognition tasks, and improves the accuracy of Chinese speech recognition.
Smart Images

Figure 2025527183000001_ABST
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a data processing system and method for speech recognition models, and a speech recognition method. [Background technology]
[0002] A speech recognition model is a model that converts input speech into text and is generally obtained by training based on speech-text labeled data. In order to improve the accuracy of the trained speech recognition model, a large amount of speech-text labeled data is usually required. However, data labeling is usually done manually, which requires a lot of manpower and material resources, making it inefficient and difficult to achieve.
[0003] Therefore, there is an urgent need for efficient data processing methods for speech recognition models. Summary of the Invention [Problem to be solved by the invention]
[0004] In view of the above, embodiments of the present specification provide a data processing system for a speech recognition model. To solve the technical deficiencies in the prior art, one or more embodiments of the present specification relate to a data processing method for a speech recognition model, a speech recognition method, a data processing device for a speech recognition model, a speech recognizer, a computing device, a computer-readable storage medium, and a computer program. [Means for solving the problem]
[0005] According to a first aspect of an embodiment of the present specification, there is provided a data processing system for a speech recognition model, the data processing system including a cloud-side device and a terminal-side device, the cloud-side device is used for: obtaining a sample set including a plurality of sample pairs each including sample speech data and sample Chinese text; using an encoder for pre-training by executing a Chinese phonetic unit prediction task on the pre-training speech data to encode the sample speech data and obtain speech features of the sample speech data; inputting the speech features into a decoder for pre-training by executing a text prediction task on the pre-training Chinese phonetic units to obtain predicted Chinese text; pre-training a model including an encoder and a decoder according to the predicted Chinese text and the sample Chinese text; and obtaining model parameters of the speech recognition model obtained by pre-training when a pre-training stopping condition is reached; the cloud-side device is further used to transmit model parameters of the speech recognition model obtained by pre-training to the terminal-side device; The terminal device is used to perform speech recognition on the target speech data using a speech recognition model, and obtain target text corresponding to the target speech data.
[0006] According to a second aspect of the present specification, there is provided a data processing method for a speech recognition model applied to a cloud-side device connected to a plurality of terminal-side devices, the data processing method comprising: obtaining a sample set including a plurality of sample pairs including sample speech data and sample Chinese text; Encoding sample speech data using an encoder that is pre-trained by performing a Chinese phonetic unit prediction task on pre-training speech data to obtain speech features of the sample speech data; inputting speech features into a decoder that is pre-trained by performing a text prediction task on pre-training Chinese phonetic units to obtain predicted Chinese text; Pre-training a model including an encoder and a decoder according to the Chinese text to be predicted and the sample Chinese text, and obtaining model parameters of the pre-trained speech recognition model when a pre-training stopping condition is reached; and transmitting model parameters of the speech recognition model obtained by pre-training to a first terminal-side device that is one of the plurality of terminal-side devices.
[0007] According to a third aspect of the present specification, there is provided a speech recognition method applied to a terminal-side device connected to a cloud-side device, the speech recognition method comprising: acquiring speech data to be recognized; A step of encoding speech data to be recognized using an encoder of the speech recognition model obtained by pre-training the cloud-side device using the data processing method for a speech recognition model provided in the second aspect, and acquiring speech features of the speech data to be recognized; inputting the speech features into a decoder of a speech recognition model to obtain target text corresponding to the speech data to be recognized.
[0008] According to a fourth aspect of the present specification, there is provided a data processing device for a speech recognition model applied to a cloud-side device connected to a plurality of terminal-side devices, the data processing device comprising: a first acquiring module configured to acquire a sample set including a plurality of sample pairs including sample speech data and sample Chinese text; a first encoding module configured to encode sample speech data using an encoder that is pre-trained by performing a Chinese phonetic unit prediction task on pre-training speech data to obtain speech features of the sample speech data; a first decoding module configured to input the speech features to a decoder that is pre-trained by performing a text prediction task on pre-training Chinese phonetic units to obtain predicted Chinese text; a pre-training module configured to pre-train a model including an encoder and a decoder based on a Chinese text to be predicted and a sample Chinese text, and obtain model parameters of the speech recognition model obtained by the pre-training when a pre-training stopping condition is reached; and a first transmission module configured to transmit model parameters of the speech recognition model obtained by pre-training to a first terminal-side device, which is one of the plurality of terminal-side devices.
[0009] According to a fifth aspect of the present specification, there is provided a speech recognition device applied to a terminal-side device connected to a cloud-side device, the speech recognition device comprising: a second acquisition module configured to acquire speech data to be recognized; a second encoding module configured to encode the speech data to be recognized using an encoder of the speech recognition model obtained by pre-training the speech recognition model in a cloud-side device using the data processing method for the speech recognition model provided in the second aspect, and to obtain speech features of the speech data to be recognized; a second decoding module configured to input the speech features into a decoder of the speech recognition model to obtain target text corresponding to the speech data to be recognized.
[0010] According to a sixth aspect of the present embodiment, there is provided a computing device, the computing device comprising: It has memory and a processor, The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, and when the computer-executable instructions are executed by the processor, steps of the speech recognition model data processing method according to the second aspect or steps of the speech recognition method according to the third aspect are implemented.
[0011] The first example of the present specification 7According to the second aspect, there is provided a computer-readable storage medium having stored thereon computer-executable instructions, which, when executed by a processor, implement steps of the speech recognition model data processing method according to the second aspect or steps of the speech recognition method according to the third aspect.
[0012] The first example of the present specification 8 According to the second aspect, a computer program is provided, which, when executed on a computer, causes the computer to perform steps of the speech recognition model data processing method according to the second aspect or steps of the speech recognition method according to the third aspect. [Effects of the Invention]
[0013] A data processing system for a speech recognition model provided by an embodiment of the present specification includes a terminal-side device and a cloud-side device, wherein the cloud-side device is used to: obtain a sample set including a plurality of sample pairs each including sample speech data and sample Chinese text; use an encoder for pre-training by performing a Chinese phonetic unit prediction task on the pre-training speech data to encode the sample speech data to obtain speech features of the sample speech data; input the speech features into a decoder for pre-training by performing a text prediction task on the pre-training Chinese phonetic units to obtain predicted Chinese text; pre-train a model including an encoder and a decoder based on the predicted Chinese text and the sample Chinese text; and obtain model parameters of the speech recognition model obtained by the pre-training when a pre-training stopping condition is reached; the cloud-side device is further used to send the model parameters of the speech recognition model obtained by the pre-training to the terminal-side device; and the terminal-side device is used to speech recognize the speech data to be recognized using the speech recognition model and obtain target text corresponding to the speech data to be recognized. That is, in the pre-training stage of the speech recognition model, the present technical solution at least performs a task of predicting Chinese text based on speech data, a Chinese pronunciation unit prediction task of predicting Chinese pronunciation units based on speech data, and a text prediction task of predicting Chinese text based on Chinese pronunciation units. Therefore, when training the speech recognition model in actual application, fewer sample speech data and sample Chinese texts are required for labeling, which reduces the burden on labelers and makes it easier to obtain labeled data.In addition, the encoder is pre-trained by performing a Chinese pronunciation unit prediction task on pre-training speech data, so that the encoder has the ability to predict Chinese pronunciation units based on speech data. Because Chinese is an ideographic language and there is a large difference between Chinese text and speech data, Chinese pronunciation units can be used as a bridge between Chinese text and speech data to reduce the difference between the two. By pre-training the encoder, the speech data can be converted into Chinese pronunciation units, so that an excellent encoder for speech recognition tasks can be obtained. Finally, the decoder is pre-trained by performing a text prediction task on the pre-training Chinese pronunciation units, so that the decoder can acquire the ability to construct Chinese text based on the Chinese pronunciation units, so that the language modeling ability of the decoder is improved. That is, the encoder and decoder are pre-trained to have a certain speech recognition ability. For pre-training, a speech-text prediction task is performed on a model consisting of an encoder and a decoder, and the model parameters are adjusted in a direction more suitable for the speech recognition task, so that training efficiency and training accuracy can be improved. Furthermore, the input of the model used in pre-training is pre-training speech data or pre-training Chinese phonetic units, both of which are similar in modality to the speech data input when the speech recognition model is applied, making the speech recognition model more suitable for downstream speech recognition tasks and improving the recognition accuracy of the speech recognition model obtained by pre-training to a certain extent. [Brief explanation of the drawings]
[0014] [Figure 1] 1 shows a flowchart of a data processing method and a speech recognition method for a speech recognition model in a data processing system architecture of the speech recognition model provided according to one embodiment of the present specification. [Figure 2] 1 shows a schematic diagram of a data processing system for a speech recognition model provided in accordance with one embodiment of the present disclosure. [Figure 3]1 illustrates a data flow diagram of a data processing method for a speech recognition model provided by an embodiment of the present specification. [Figure 4a] 2 illustrates a data flow diagram of a method for determining Chinese phonetic units provided by an embodiment of the present application; [Figure 4b] FIG. 1 shows a data flow diagram of one of the encoder pre-training methods provided by one embodiment of the present application. [Figure 5] FIG. 10 illustrates another data flow diagram for pre-training an encoder provided by an embodiment of the present application. [Figure 6] 1 shows a data flow diagram of one of the decoder pre-training provided by one embodiment of the present application. [Figure 7] FIG. 10 illustrates another data flow diagram for decoder pre-training provided by an embodiment of the present disclosure. [Figure 8] 1 illustrates a data flow diagram of a method for fine-tuning a speech recognition model provided by one embodiment of the present disclosure. [Figure 9] 1 illustrates a flowchart of a data processing method for a speech recognition model applied to a cloud-side device provided by an embodiment of the present specification. [Figure 10] 1 shows a flowchart of a speech recognition method applied to a terminal-side device provided by an embodiment of the present specification; [Figure 11] 1 illustrates a data flow diagram of the performance of a speech recognition task by a speech recognition model provided by an embodiment of the present specification. [Figure 12] 1 shows a flowchart of a processing process of a data processing method for a speech recognition model provided by an embodiment of the present specification. [Figure 13] 1 shows a data flow diagram during combined training of a speech recognition model provided by an embodiment of the present application. [Figure 14] 1 illustrates a structural schematic diagram of a data processing device for a voice recognition model applied to a cloud-side device provided by an embodiment of the present specification; [Figure 15]1 shows a structural schematic diagram of a speech recognition device provided by an embodiment of the present specification; [Figure 16] 1 illustrates a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0015] In order to facilitate a thorough understanding of the present specification, many specific details are set forth in the following description. However, the present specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar applications without violating the content of the present specification. Therefore, the present specification is not limited by the specific implementations disclosed below.
[0016] The terms used in one or more embodiments herein are used only for the purpose of describing a particular embodiment and are not intended to limit one or more embodiments herein. As used in one or more embodiments herein and in the appended claims, the singular forms of "a," "the," and "the" are intended to be used interchangeably unless the context clearly dictates otherwise. multiple It should be further understood that the term "and / or," as used in one or more examples herein, refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0017] In one or more embodiments herein, the term 「 No. 1 」 , 「 No. 2 」While terms such as "predetermined" and "predetermined" may be used, it should be understood that such information should not be limited to these terms. These terms are used only to distinguish between similar types of information. For example, a first could be referred to as a second, and similarly, a second could be referred to as a first, without departing from the scope of one or more embodiments herein. As used herein, the term "suppose" can be interpreted as "when," "with," or "in response to a determination," depending on the context.
[0018] It should be noted that user-related information and user-related data in relation to the embodiments herein are all information and data that are authorized by the user or fully authorized by the parties.
[0019] First, a description of noun terms relating to one or more embodiments of this specification will be provided.
[0020] The speech recognition model is used to recognize input speech data and obtain text.
[0021] The encoder is used to encode the input speech, characters, Chinese phonetic units, etc., to represent the input in the form of a feature vector.
[0022] The decoder is used to decode the input feature vector to obtain voice, characters, etc.
[0023] The feature encoding layer is used to encode the input features and capture the association relationships between the features.
[0024] Speech features are vectorized representations of speech.
[0025] A speech encoder is used to encode speech to obtain a vectorized representation of the speech, i.e., speech features.
[0026] CTC (Connectionist temporal classification) is an algorithm that is mainly used to deal with problems where the input array is longer than the output array, and achieves alignment of the labels of the input and output arrays.
[0027] Speech recognition models, which convert speech data into text, often require large amounts of labeled data for training. Pre-training on unlabeled data allows models to achieve better results at lower cost and with greater ease. End-to-end speech recognition models commonly use two types of architectures: the conjunctive correlation prediction (CTC) architecture and the encoder-decoder architecture. While the encoder-decoder architecture considers both speech data and text grammatical information, the CTC architecture is superior to the CTC architecture because it only considers speech data. Because Chinese is an ideographic language compared to other phonetic languages such as English, the support of text grammatical information is crucial during speech recognition. However, training an end-to-end speech recognition model requires a large amount of labeled speech-text pairs, making end-to-end recognition more challenging, especially for ideographic languages like Chinese.
[0028] Currently, unsupervised pretraining has brought significant improvements to downstream tasks in various fields. Regarding pretraining of speech recognition models, a series of methods for encoder pretraining using unlabeled speech data, such as HuBERT and Data2Vec, have been proposed. However, these pretraining methods are often only applicable to CTC-structured models due to optimization, and problems arise when applied to encoder-decoder structures. This is because the decoder is not involved in the pretraining. STPT and Speech2C further propose pretraining the decoder using unlabeled text or unlabeled speech data, while SpeechT5 proposes pretraining an encoder-decoder model using both unlabeled text and unlabeled speech data. However, these methods have problems: the complementarity between different pretraining tasks has not been verified; they are designed based on the characteristics of English phonetic languages, not taking into account the characteristics of Chinese ideographic languages; and unlabeled text data is not sufficiently utilized, resulting in limited improvements.
[0029] For speech recognition tasks, currently mainstream and effective speech pre-training methods can be categorized into two types:
[0030] The first type is unimodal speech representation pre-training, such as Wav2vec 2.0, HuBERT, Data2vec, and Speech2C. This type of method uses only unlabeled speech data and achieves better speech modeling capabilities through masked prediction. However, this type of method lacks pre-training for modeling text semantic information, resulting in poor performance when applied to encoder-decoder speech recognition models.
[0031] The second type is speech-text multimodal pre-training, such as STPT and SpeechT5, which introduces unlabeled text data in addition to unlabeled speech data for pre-training, enabling the encoder-decoder structure model to acquire the ability to model both speech information and text grammatical information.
[0032] The deficiencies of STPT are 1) that unlabeled speech data does not participate in decoder parameter updates, and 2) that it is prone to model collapse problems.
[0033] Since pre-training using SpeechT5 is primarily intended to obtain a general-purpose speech model, the design of the pre-training task does not solely consider the speech recognition task. Therefore, the defects of pre-training of speech recognition models are: 1) the sequence-to-sequence task designed using unlabeled speech data is a speech frame reconstruction task, which is advantageous for speech synthesis tasks but disadvantageous for speech recognition tasks; and 2) when using unlabeled text data, the model input is text, and for pictographic characters like Chinese, the difference between the two modalities of text and speech is large, making combined training difficult.
[0034] Therefore, this specification provides a data processing system for a voice recognition model that can solve the above technical problems, and for specific implementations thereof, please refer to the relevant descriptions in the following embodiments.
[0035] This specification provides a data processing system for a speech recognition model, and the specification relates to a data processing method for a speech recognition model, a speech recognition method, a data processing apparatus for a speech recognition model, a speech recognition apparatus, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0036] Please refer to Figure 1. Figure 1 shows a flowchart of a data processing method and a speech recognition method of a speech recognition model in a data processing system architecture of the speech recognition model provided according to one embodiment of the present specification.
[0037] The system may include a cloud-side device 101 and a terminal-side device 102, where the cloud-side device 101 is used to train a speech recognition model and the terminal-side device 102 is used to perform a speech recognition task based on the speech recognition model obtained by training.
[0038] The cloud-side device 101 may be a central cloud device in a distributed cloud architecture, and the terminal-side device 102 may be an edge cloud device in the distributed cloud architecture. The cloud-side device 101 and the terminal-side device 102 may be server-side devices such as traditional servers, cloud servers, or server arrays, or may be terminal devices, and the embodiments of this specification are not limited thereto. Furthermore, the cloud-side device 101 provides very large computing and storage capabilities and is far away from users, while the terminal-side device 102 is widely deployed and close to users. The terminal-side device 102 is an extension of the cloud-side device 101, and the computing capabilities of the cloud-side device 101 can be transferred to the terminal-side device 102, thereby solving business requirements that cannot be met by centralized cloud computing through end-cloud integration and collaborative management.
[0039] In one or more embodiments of the present specification, the cloud-side device 101 obtains a sample set including a plurality of sample pairs each including sample speech data and sample Chinese text, uses an encoder for pre-training by performing a Chinese phonetic unit prediction task on the pre-training speech data, encodes the sample speech data to obtain speech features of the sample speech data, inputs the speech features into a decoder for pre-training by performing a text prediction task on the pre-training Chinese phonetic units to obtain predicted Chinese text, pre-trains a model including an encoder and a decoder according to the predicted Chinese text and the sample Chinese text, and when a pre-training stopping condition is reached, obtains model parameters of a speech recognition model obtained by pre-training, and sends the model parameters of the speech recognition model to the terminal-side device 102.
[0040] After receiving the speech recognition model, the terminal-side device 102 uses the speech recognition model to perform speech recognition on the speech data to be recognized, and obtains target text corresponding to the speech data to be recognized. Specifically, the terminal-side device 102 obtains the speech data to be recognized, inputs the speech data to be recognized into an encoder of the speech recognition model to obtain speech features of the speech data to be recognized, and inputs the speech features into a decoder of the speech recognition model to obtain target text corresponding to the speech data to be recognized.
[0041] In the data processing system for a speech recognition model provided by the embodiments of this specification, during the pre-training stage of a speech recognition model, the cloud-side device at least performs a task of predicting Chinese text based on speech data, a Chinese pronunciation unit prediction task of predicting Chinese pronunciation units based on speech data, and a text prediction task of predicting Chinese text based on Chinese pronunciation units. Therefore, when training a speech recognition model in actual application, fewer sample speech data and sample Chinese texts are required for labeling, which reduces the burden on labelers and makes it easier to obtain labeled data. In addition, the encoder is pre-trained by performing a Chinese pronunciation unit prediction task on pre-training speech data, so that the encoder has the ability to predict Chinese pronunciation units based on speech data. Because Chinese is an ideographic language and there is a large difference between Chinese text and speech data, Chinese pronunciation units can be used as a bridge between Chinese text and speech data. In order to reduce the difference between the two, the encoder can be pre-trained to convert speech data into Chinese pronunciation units, thereby obtaining an excellent encoder for speech recognition tasks. The decoder is pre-trained by performing a text prediction task on the pre-training Chinese pronunciation units, so that the decoder can acquire the ability to construct Chinese text based on the Chinese pronunciation units, thereby improving the language modeling ability of the decoder. That is, the encoder and decoder are pre-trained to have a certain speech recognition ability. For pre-training, a speech-text prediction task is performed on a model consisting of an encoder and a decoder, and the model parameters are adjusted in a direction more suitable for the speech recognition task, thereby improving training efficiency and training accuracy. Furthermore, the input of the model used in pre-training is pre-training speech data or pre-training Chinese phonetic units, both of which are similar in modality to the speech data input when the speech recognition model is applied, making the speech recognition model more suitable for downstream speech recognition tasks and improving the recognition accuracy of the speech recognition model obtained by pre-training to a certain extent.
[0042] 2 is a schematic diagram of a data processing system for a speech recognition model provided according to an embodiment of the present specification. Referring to FIG. 2, the data processing system for a speech recognition model provided according to an embodiment of the present specification includes a terminal-side device 201 and a cloud-side device 202 communicably connected to the terminal-side device 201.
[0043] the cloud-side device 202 is used for: obtaining a sample set including a plurality of sample pairs each including sample speech data and sample Chinese text; using an encoder for pre-training by executing a Chinese phonetic unit prediction task on the pre-training speech data to encode the sample speech data and obtain speech features of the sample speech data; inputting the speech features into a decoder for pre-training by executing a text prediction task on the pre-training Chinese phonetic units to obtain predicted Chinese text; pre-training a model including an encoder and a decoder according to the predicted Chinese text and the sample Chinese text; and obtaining model parameters of the speech recognition model obtained by pre-training when a pre-training stopping condition is reached; The cloud-side device 202 is further used to transmit model parameters of the pre-trained speech recognition model to the terminal-side device; The terminal-side device 201 is used to perform speech recognition on the speech data to be recognized using a speech recognition model, and to obtain target text corresponding to the speech data to be recognized.
[0044] In one or more embodiments of the present specification, the task to be performed by the speech recognition model is a speech recognition task of predicting Chinese text based on speech data, so the training of the speech recognition model is supervised training, and therefore, sample pairs consisting of sample speech data and sample text need to be obtained. In addition, since the speech recognition model provided by the present specification is a speech recognition model for the Chinese language, the obtained sample text is sample Chinese text. Furthermore, there is a correspondence between the sample speech data and the sample Chinese text in each sample pair.
[0045] For example, the cloud-side device 202 may obtain a plurality of sample pairs from an open-source sample library, and the plurality of sample pairs constitute a sample set. Illustratively, the sample audio data may be any audio data, such as a few sentences or a single word. For example, the sample audio data may be audio chat, speech, music, conference recording, etc.
[0046] In an embodiment of the present specification, the cloud-side device 202 obtains a sample set including a plurality of sample pairs each including sample speech data and sample Chinese text, pre-trains a model with an encoder and a decoder based on the sample set, obtains a speech recognition model, and sends model parameters of the pre-trained speech recognition model to the terminal-side device 201, so as to enable the terminal-side device 201 to perform speech recognition based on the speech recognition model.
[0047] The cloud-side device 202 is used to encode sample speech data and obtain speech features of the sample speech data using an encoder that performs pre-training by executing a Chinese phonetic unit prediction task on the pre-training speech data.
[0048] Here, the Chinese pronunciation unit prediction task is a task to predict corresponding Chinese pronunciation units based on input pre-training speech data, where the Chinese pronunciation units are constituent units of Chinese pronunciation, and may be pinyin, syllables, or Chinese phonemes, for example, yin, in, i, y, etc.
[0049] Here, the speech features of the sample speech data are vectorized representations of the sample speech data, and are obtained by taking into account the contextual relationships between each word of the sample speech data.
[0050] Note that the pre-training is performed by the encoder executing a Chinese phonetic unit prediction task on the pre-training speech data, but the pre-training may be performed before pre-training the model consisting of the decoder and the encoder, or may be performed during the process of pre-training the model consisting of the decoder and the encoder, and the examples in this specification are not limited to this.
[0051] Exemplarily, the encoder may be an encoder in any model with encoding capabilities. For example, the encoder may be an encoder in a transformer model, or an encoder in a BERT model. (Bidirectional Encoder Representation from Transformer) The encoder may be a model such as a CNN (Convolutional Neural Network), a LSTM (Long Short Term Memory), or a GRU (Gate Recurrent Unit). Furthermore, the encoder may include M+N blocks, where M and N are both positive integers greater than 1.
[0052] In a first possible implementation of the present specification, sample audio data may be input to an encoder, which may then encode the sample audio data using M+N blocks to obtain audio features corresponding to the sample audio data.
[0053] In a second possible implementation form of the present specification, the encoder may include a speech encoding layer and a feature encoding layer. Illustratively, the speech encoding layer includes M blocks, and the feature encoding layer includes N blocks. Sample speech data is first input to the speech encoding layer to obtain initial speech features of the sample speech data, and the initial speech features are then input to the feature encoding layer to obtain speech features of the sample speech data. Here, the difference between the initial speech features and the speech features is that the initial speech features are speech context features processed by M blocks, and the speech features are speech context features processed by M+N blocks.
[0054] In a third possible implementation form of the present specification, the encoder may include a feature extraction layer connected to the audio encoding layer, where the sample audio data is input to the feature extraction layer for audio representation extraction and downsampling processing, features of each word of the sample audio data are extracted to obtain an audio representation vector of the sample audio data, and the audio representation vector is input to the audio encoding layer to obtain initial audio features of the sample audio data, and the initial audio features are input to the feature encoding layer to obtain audio features of the sample audio data.
[0055] Here, the feature extraction layer may be referred to as a feature extractor. For example, the feature extraction layer may be a component of an encoder, and in this case, the feature extraction layer participates in the pre-training process of the encoder.
[0056] In a fourth possible implementation form of the present specification, before encoding the sample speech data using an encoder, spectral features of the sample speech data are first extracted, and the spectral features are input to the encoder to obtain speech features of the sample speech data that take into account the contextual semantic relationships of the sample speech data.
[0057] As an example, a linear predictive cepstrum coefficient algorithm or a Mel-frequency cepstrum coefficient algorithm Alternatively, a pre-trained spectral feature extraction model may be used to extract the spectral features of the sample audio data. Furthermore, the linear predictive cepstrum coefficient algorithm or the Mel-frequency cepstrum coefficient algorithm is based on the cepstrum and is more in line with the principles of human hearing, making it an effective spectral feature extraction algorithm.
[0058] In the embodiments of the present specification, by pre-training the encoder by performing a Chinese phonetic unit prediction task on pre-training speech data, the encoder can acquire the ability to predict Chinese phonetic units using speech, i.e., the ability to encode speech data; furthermore, by pre-training the encoder and pre-training a model consisting of an encoder and a decoder at the same time, the training speed can be increased and the training efficiency can be improved.
[0059] The cloud-side device 202 is used to input the speech features into a decoder that is pre-trained by performing a text prediction task on the pre-training Chinese phonetic units to obtain predicted Chinese text.
[0060] Here, the text prediction task is a task of predicting corresponding Chinese text based on input pre-training Chinese phonetic units.
[0061] Exemplarily, the decoder may be a decoder in any model having a decoding function, for example, the decoder may be a decoder of a transformer model, or the decoder may be a decoder of a model such as BERT, CNN, LSTM, GRU, etc. Furthermore, the decoder may have X blocks, where X is a positive integer greater than or equal to 1.
[0062] In one or more embodiments herein, the decoder may include a decoding layer and a text embedding layer, in which the speech features output from the encoder are first input into the decoding layer to obtain predicted text features, and the predicted text features are input into the text embedding layer to map the predicted text features as a probability distribution vector indicating the probability that the sample speech data is a specific Chinese text, and the Chinese text with the highest probability is determined as the predicted Chinese text corresponding to the sample speech data.
[0063] For example, the decoder uses an autoregressive method to decode the speech features and predict each word to obtain the predicted Chinese text. In the pre-training stage, when the decoder predicts the current Chinese text, its input includes the speech features output from the encoder and the previous text ground-truth features. In the testing stage, when the decoder predicts the current Chinese text, its input includes the speech features output from the encoder and the decoding result of the previous word output from the text embedding layer.
[0064] That is, since the decoder performs decoding while taking into account inter-contextual connections, more accurate decoding results can be obtained, and furthermore, predicted Chinese texts with a higher accuracy rate can be obtained.
[0065] In one or more embodiments of the present specification, a decoder is pre-trained by performing a text prediction task on pre-training Chinese phonetic units, so that the decoder can acquire the ability to predict Chinese text based on the Chinese phonetic units, and since the Chinese phonetic units are also phonetic features, the decoder has the ability to construct text based on the phonetic features. In addition, when pre-training a model consisting of an encoder and a decoder, the training speed can be increased and the training efficiency can be improved. Furthermore, since the Chinese language is an ideographic language and it is difficult to determine pronunciation from Chinese text, and the pre-training Chinese phonetic units are closer to speech data in modality than Chinese text, the pre-training text data is first converted into pre-training Chinese phonetic units as input for the model, and the Chinese text is predicted, and the pre-training task is made similar to the input of the training task of the speech recognition model, so that the pre-trained text data obtained by pre-training can be Speech Recognition Model becomes more adaptive to the speech recognition task, and the recognition accuracy of the trained speech recognition model is improved.
[0066] The cloud-side device 202 is used to pre-train a model consisting of an encoder and a decoder based on the predicted Chinese text and the sample Chinese text, and obtain model parameters of the pre-trained speech recognition model when a pre-training stopping condition is reached.
[0067] In one or more embodiments of the present specification, a loss value can be determined based on the predicted Chinese text and the sample Chinese text, and if the loss value is equal to or greater than a loss threshold, the parameters of the model consisting of an encoder and a decoder are adjusted based on the loss value, i.e., the parameters of the encoder and the decoder are adjusted, and then the process of encoding the sample speech data using the encoder is performed again, and if a pre-training stopping condition is reached, the pre-training is stopped and the model parameters of the speech recognition model are obtained.
[0068] In some examples herein, the pre-training stopping condition may be a condition where the loss value is less than a loss threshold, or the pre-training stopping condition may be a condition where the number of pre-training iterations is greater than or equal to a number threshold.
[0069] For example, if the loss value obtained after a certain pre-training is smaller than the loss threshold, it is proven that the model consisting of the encoder and decoder can already perfectly complete speech recognition and there is no need to continue training, so the pre-training is stopped and the model parameters of the speech recognition model obtained through pre-training are obtained.
[0070] As another example, we record the number of pre-training iterations and predict the Chinese text Each time is determined, the number of pre-training iterations may be incremented by 1. If the number of pre-training iterations is greater than the threshold number, it is proven that the pre-training for the model consisting of the encoder and decoder is sufficient, and continued pre-training may not achieve better results. Therefore, the pre-training is stopped and the model parameters of the speech recognition model obtained by pre-training are obtained.
[0071] In an embodiment of the present application, the cloud-side device 202 performs pre-training to obtain a speech recognition model, and then transmits the model parameters of the speech recognition model to the terminal-side device 201 so that the terminal-side device 201 can perform a speech recognition task based on the speech recognition model.
[0072] Please refer to Figure 3. Figure 3 shows a data flow diagram of a data processing method for a speech recognition model provided by one embodiment of the present specification. The speech recognition model includes an encoder and a decoder, where the encoder includes a feature extraction layer, a speech encoding layer, and a feature encoding layer. The decoder includes a decoding layer and a text embedding layer. Sample speech data is input to the feature extraction layer to obtain a speech representation vector of the sample speech data, and the speech representation vector is input to the speech encoding layer for encoding. The encoding result is input to the feature encoding layer to obtain speech features. The speech features are input to the decoding layer to obtain predicted text features. The predicted text features are input to the text embedding layer to obtain predicted Chinese text. After the predicted Chinese text is obtained, the parameters of each component in the model including the encoder and decoder are adjusted based on the predicted Chinese text and the sample Chinese text. If a pre-training stopping condition is met, a speech recognition model is obtained. In addition, since the decoder performs decoding using an autoregressive method, in the pre-training stage, the decoder's input includes the speech features output from the encoder and the previous text ground-truth features, and in the testing stage, the decoder's input includes the speech features output from the encoder and the decoding result output from the text embedding layer.
[0073] The above content is a process of pre-training a speech recognition model according to a speech recognition task, but as described above, the encoder and decoder are pre-trained using the Chinese phonetic unit prediction task and the text prediction task, respectively, so the pre-training process of the encoder and decoder will be described below. In the embodiments of this specification, the encoder and decoder may be pre-trained simultaneously.
[0074] The first part describes the process of pre-training the encoder.
[0075] In one or more embodiments herein, the cloud-side device 202 may: The method further includes obtaining a first pre-training speech set including a plurality of unsupervised first pre-training speech data; encoding the first pre-training speech data using an encoder to obtain first speech features corresponding to the first pre-training speech data, and determining first pronunciation units based on the first speech features; masking the first pre-training speech data; encoding the masked first pre-training speech data using an encoder to obtain second speech features corresponding to the masked first pre-training speech data, and determining second pronunciation units based on the second speech features; and pre-training the encoder based on the first pronunciation units and the second pronunciation units corresponding to the first pre-training speech data.
[0076] Here, the first pre-training speech data is unlabeled speech data.
[0077] In the embodiments herein, the pre-training task used to pre-train the encoder is an audio mask prediction task, in which audio features are output from the encoder; i.e., the audio mask prediction task is used to determine audio features based on audio data, thereby determining corresponding phonetic units, and adjusting the parameters of the encoder so that the encoder can output more accurate audio features.
[0078] For example, the cloud-side device 202 can obtain a plurality of unsupervised first pre-training speech data from an open-source database, organize the first pre-training speech data into a first pre-training speech set, and pre-train the encoder based on the first pre-training speech data. The obtained first pre-training speech data is unlabeled, which reduces the cost of manual labeling.
[0079] In some embodiments herein, first pre-training speech data may be input to an encoder to obtain first speech features corresponding to the first pre-training speech data.
[0080] In some other embodiments herein, the cloud-side device 202 is further used for extracting spectral features of the first pre-training speech data and inputting the spectral features of the first pre-training speech data into an encoder to obtain first speech features corresponding to the first pre-training speech data.
[0081] That is, before encoding the first pre-training speech data using the encoder, the spectral features of the first pre-training speech data are extracted, and the spectral features are input to the encoder for encoding, thereby obtaining the first speech features corresponding to the first pre-training speech data. In addition, in the embodiments of the present specification, the encoder and decoder are simultaneously pre-trained using various pre-training tasks. Since the spectral features introduce less acoustic details than the waveform features (speech data), when the spectral features are used as the input for the encoder, it becomes difficult for the model consisting of the encoder and the decoder to distinguish between speech data of different pre-training tasks. Therefore, different pre-training tasks are not independent of each other, but can mutually promote each other to improve the training effect.
[0082] As an example, a linear predictive cepstrum coefficient algorithm or a Mel-frequency cepstrum coefficient algorithm Alternatively, a spectral feature extraction model obtained by pre-training can be used to extract spectral features from the first pre-training speech data.
[0083] For example, the encoder may include a feature extraction layer, a speech encoding layer, and a feature encoding layer, and inputting the spectral features to the encoder for encoding may include inputting the spectral features to the feature extraction layer for speech representation extraction and downsampling to obtain a speech representation vector corresponding to first pre-training speech data, inputting the speech representation vector to the speech encoding layer to obtain initial speech features, and inputting the initial speech features to the feature encoding layer to obtain first speech features corresponding to the first pre-training speech data. The difference between the initial speech features and the first speech features is that the initial speech features are speech context features obtained by M blocks, while the first speech features are speech context features obtained after M+N block processing.
[0084] After the first phonetic units corresponding to the first pre-training speech data are determined, the first phonetic units can be used as labels for the first pre-training speech data, and the second phonetic units corresponding to the first pre-training speech data are predicted using mask prediction, and the encoder is pre-trained based on the second phonetic units and the first phonetic units.
[0085] In some embodiments of the present specification, spectral features of a first pre-training speech data are extracted, and the spectral features are input to a feature extraction layer of an encoder to determine a speech representation vector. The speech representation vector is then randomly masked, and the masked speech representation vector is input to a speech encoding layer to obtain initial speech features. The initial speech features are input to a feature encoding layer to obtain second speech features corresponding to the first pre-training speech data, thereby determining the corresponding second pronunciation unit. A loss value can be determined based on the first pronunciation unit and the second pronunciation unit. If the loss value is greater than or equal to a loss threshold, it proves that the similarity between the second pronunciation unit and the first pronunciation unit is low and the predicted second pronunciation unit is not accurate. That is, the encoder has not yet mastered the ability to predict Chinese pronunciation units based on speech data, and training needs to continue. Therefore, pre-training of the encoder continues until a pre-training stopping condition is reached.
[0086] Illustratively, the pre-training stopping condition may include a condition where the loss value is less than a loss threshold, or the pre-training stopping condition may include a condition where the number of pre-trainings is greater than or equal to a number threshold.
[0087] For example, if the loss value obtained in a certain pre-training is smaller than the loss threshold, it means that the similarity between the second pronunciation unit and the first pronunciation unit is high and the predicted second pronunciation unit is proven to be relatively accurate, i.e., the encoder can already perfectly predict the Chinese pronunciation unit based on the speech data, and there is no need to continue training any further, so the encoder pre-training is stopped.
[0088] As another example, in the pre-training process, the number of pre-trainings is recorded, and 1 is added to the number of pre-trainings each time a second phonetic unit is determined. If the number of pre-trainings is greater than the number threshold, it indicates that the number of pre-trainings of the encoder is sufficient, and even if the pre-training is continued, better results may not be achieved, so the pre-training of the encoder is stopped.
[0089] In the embodiments of the present specification, first, a first pronunciation unit of the first pre-training speech data is determined by the encoder and the pronunciation unit embedding layer, then a phonetic representation vector of the first pre-training speech data is extracted, and the phonetic representation vector is masked. After that, a second pronunciation unit is determined by the encoder and the pronunciation unit embedding layer, and the encoder is pre-trained based on the first pronunciation unit and the second pronunciation unit. This allows the encoder to pay more attention to the pronunciation information of speech words and the context between words when predicting Chinese pronunciation units, thereby improving the accuracy of the encoder's prediction of Chinese pronunciation units and improving the encoder's coding ability for speech data.
[0090] Please refer to Figure 4a. Figure 4a shows a data flow diagram of a method for determining Chinese pronunciation units provided by one embodiment of the present application. When determining Chinese pronunciation units, an encoder and a pronunciation unit embedding layer are used, and the encoder includes a feature extraction layer, a phonetic encoding layer, and a feature encoding layer. First pre-training speech data is obtained, the first pre-training speech data is input to the feature extraction layer to obtain a phonetic representation vector of the first pre-training speech data, the phonetic representation vector is input to the phonetic encoding layer for encoding, and then the encoding result is input to the feature encoding layer to obtain a first phonetic feature corresponding to the first pre-training speech data, and the first phonetic feature is input to the phonetic unit embedding layer to obtain a first Chinese pronunciation unit corresponding to the first pre-training speech data.
[0091] Please refer to Figure 4b. Figure 4b shows a data flow diagram of an encoder pre-training provided by one embodiment of the present application. When pre-training the encoder, an encoder and a phonetic unit embedding layer of a speech recognition model are used, and the encoder includes a feature extraction layer, a speech encoding layer, and a feature encoding layer. First pre-training speech data is obtained, the first pre-training speech data is input to the feature extraction layer to obtain a phonetic representation vector of the first pre-training speech data, the phonetic representation vector is masked, and the masked phonetic representation vector is input to the phonetic encoding layer for encoding, and then the encoding result is input to the feature encoding layer to obtain a second phonetic feature, and the second phonetic feature is input to the phonetic unit embedding layer to obtain a second Chinese phonetic unit, and encoder parameters are adjusted based on the first Chinese phonetic unit and the second Chinese phonetic unit.
[0092] moreover, 1st The speech mask prediction task involves predicting a second phonetic unit based on pre-training speech data. 1st Since the pre-training speech data is used as labels and the encoder learns its own output, it may cause model collapse, i.e., the encoder may output the same features regardless of the speech data input, making the speech mask prediction task meaningless. Therefore, to improve the accuracy of the encoder's prediction, the encoder can be pre-trained using the Chinese phonetic unit prediction task together with the speech mask prediction task, which allows the encoder to more accurately predict Chinese phonetic units based on the speech data.
[0093] In one or more embodiments herein, the cloud-side device 202 may: Obtaining a plurality of first pre-training pairs including second pre-training speech data and first pre-training Chinese phonetic units; Using an encoder to predict Chinese pronunciation units for the second pre-training speech data, and obtain predicted Chinese pronunciation units corresponding to the second pre-training speech data; and pre-training an encoder based on the first pre-training Chinese phonetic unit and the predicted Chinese phonetic unit.
[0094] In some embodiments, the cloud-side device 201 can obtain a first pre-training pair from an open-source pre-training dataset and pre-train an encoder based on the first pre-training pair. In addition, because the correspondence between the second pre-training speech data and the first pre-training Chinese pronunciation units contained in the first pre-training pair is highly accurate, the cloud-side device 202 can pre-train the encoder based on the pre-training pair to improve the accuracy of the encoder's prediction of Chinese pronunciation units. That is, the encoder can be pre-trained based on a supervised Chinese pronunciation unit prediction task, and the first pre-training Chinese pronunciation units are labels of the second pre-training speech data.
[0095] In some embodiments, the cloud-side device 202 using the encoder to predict Chinese pronunciation units for the second pre-training speech data includes inputting the second pre-training speech data to the encoder to obtain predicted speech features for the second pre-training speech data, and inputting the predicted speech features to a pronunciation unit embedding layer to obtain predicted Chinese pronunciation units corresponding to the second pre-training speech data.
[0096] In some other embodiments, before using the encoder to predict Chinese pronunciation units for the second pre-training speech data, the cloud-side device 202 first extracts spectral features of the second pre-training speech data, inputs the spectral features into the encoder to obtain predicted pronunciation features for the second pre-training speech data, and inputs the predicted pronunciation features into the pronunciation unit embedding layer to obtain predicted Chinese pronunciation units corresponding to the second pre-training speech data.
[0097] In some other embodiments, the encoder includes a feature extraction layer, a speech encoding layer, and a feature encoding layer, and before using the encoder to predict Chinese pronunciation units for the second pre-training speech data, the cloud-side device 202 first extracts spectral features of the second pre-training speech data, inputs the spectral features into the feature extraction layer for speech representation extraction and downsampling processing to obtain speech representation vectors corresponding to the second pre-training speech data, inputs the speech representation vectors into the speech encoding layer and the feature encoding layer to obtain predicted speech features corresponding to the second pre-training speech data, and inputs the predicted speech features into the pronunciation unit embedding layer to obtain predicted Chinese pronunciation units corresponding to the second pre-training speech data.
[0098] It should be noted that the implementation process of pre-training an encoder based on the first pre-training Chinese phonetic unit and the predicted Chinese phonetic unit is similar to the implementation process of pre-training a model with an encoder and a decoder based on the predicted Chinese text and the sample Chinese text, and the specific implementation thereof can refer to the relevant description in the above embodiment, and this embodiment will not be further described.
[0099] In the embodiments of the present specification, by pre-training an encoder using a speech mask prediction task and a Chinese pronunciation unit prediction task, the encoder can acquire the ability to predict Chinese pronunciation units based on speech data, thereby improving the prediction accuracy of the trained encoder. Furthermore, since spectral features can indicate the timbre, pitch, and Chinese pronunciation units of each word in speech, using Chinese pronunciation units as the prediction target in encoder pre-training allows the encoder to focus on capturing the pronunciation information of speech. Furthermore, since multiple tasks are simultaneously used for pre-training during the pre-training phase, the raw speech data contains more noise details, which allows the model to distinguish between data used for different pre-training tasks. This weakens the effects of inter-task promotion and constraint, leading to training instability. Therefore, using spectral features as the input for the entire model when pre-training an encoder can promote constraints between different pre-training tasks, improve the stability of model training, and avoid the problem of model collapse.
[0100] Please refer to Figure 5. Figure 5 shows another data flow diagram of pre-training an encoder provided by one embodiment of the present application. When pre-training the encoder, an encoder and a phonetic unit embedding layer are used, and the encoder includes a feature extraction layer, a speech encoding layer, and a feature encoding layer. Second pre-training speech data are input to the feature extraction layer to obtain a phonetic representation vector of the second pre-training speech data, the phonetic representation vector is input to the speech encoding layer for encoding, and the encoding result is input to the feature encoding layer to obtain predicted phonetic features corresponding to the second pre-training speech data, the predicted phonetic features are input to the phonetic unit embedding layer to obtain predicted Chinese phonetic units, and encoder parameters are adjusted based on the predicted Chinese phonetic units and the first pre-training Chinese phonetic units.
[0101] In the second part, we describe the process of pre-training the decoder.
[0102] In one or more embodiments herein, the encoder comprises a feature encoding layer, and the cloud-side device 202 Obtaining a first pre-training text set including a plurality of unsupervised first pre-training Chinese texts; Converting the first pre-training Chinese text into a second pre-training Chinese phonetic unit, and inputting the second pre-training Chinese phonetic unit into a feature encoding layer to obtain phonetic features of the second pre-training Chinese phonetic unit; Inputting the phonetic features of the second pre-training Chinese phonetic unit into a decoder to obtain predicted Chinese text corresponding to the second pre-training Chinese phonetic unit; and pre-training a decoder based on the predicted Chinese text corresponding to the second pre-training Chinese phonetic unit and the first pre-training Chinese text.
[0103] Here, the feature encoding layer is an encoding layer in the encoder that is used to pre-train the decoder using Chinese phonetic units as input. Because the second pre-training Chinese phonetic units are more abstract than the speech data, the phonetic features of the second pre-training Chinese phonetic units can be obtained by inputting the second pre-training Chinese phonetic units into the feature encoding layer without using other encoding layers of the encoder for encoding.
[0104] In the embodiments of the present specification, when pre-training the decoder, the pre-training task used is a text prediction task, i.e., predicting Chinese text based on input Chinese phonetic units, thereby adjusting the parameters of the decoder, improving the decoder's text construction ability, and making the decoder more accurate in predicting Chinese text.
[0105] In some embodiments, the correspondence between Chinese text and Chinese pronunciation units may be determined based on a dictionary to convert the first pre-training Chinese text into the second pre-training Chinese pronunciation units, or a model capable of realizing the conversion between Chinese text and Chinese pronunciation units may be pre-trained and obtained, and the model may be used to convert the first pre-training Chinese text into the second pre-training Chinese pronunciation units.
[0106] Exemplarily, in the process of converting the first pre-training Chinese text into the second pre-training Chinese pronunciation units, not only the Chinese pronunciation units of the Chinese text are obtained by conversion, but also those including the tones of pronunciation (e.g., the first tone, the second tone, the third tone, the fourth tone) need to be obtained. Furthermore, in order to avoid confusion of different words, it is also necessary to divide the Chinese pronunciation units of different words, such as dividing them into initials and finals or dividing them syllable by syllable.
[0107] For example, assuming that the first pre-training Chinese text is "今天天气真不錯,但下午可能下雨(Japanese translation: 今日はいい天気だが、午後は雨が降るかもしれません。)", the second pre-training Chinese pronunciation units obtained by conversion may be "j in1 t ian1 t ian1 qi4 zh en1 b u2 c uo4 d an4 x ia4 w u3 k e3 n eng2 x ia4 y u3". Here, the numbers represent tones, 1 represents the first tone, 2 represents the second tone, 3 represents the third tone, and 4 represents the fourth tone.
[0108] In some embodiments, for voice expression extraction and downsampling processing, the second pre-training Chinese pronunciation units may be directly input into the feature encoding layer to obtain the voice features of the second pre-training Chinese pronunciation units.
[0109] In some other embodiments, before the second pre-training Chinese phonetic units are input into the feature encoding layer, the second pre-training Chinese phonetic units may be first mapped as a feature matrix, and then the feature matrix may be input into the feature encoding layer to obtain the phonetic features of the second pre-training Chinese phonetic units. Illustratively, a Chinese phonetic unit embedding layer may be used to map the second pre-training Chinese phonetic units as a feature matrix.
[0110] In some other embodiments, before the second pre-training Chinese phonetic units are input into the feature encoding layer, the second pre-training Chinese phonetic units are first masked, and the masked second pre-training Chinese phonetic units are mapped as a feature matrix, and the feature matrix is then input into the feature encoding layer to obtain the phonetic features of the second pre-training Chinese phonetic units.
[0111] It should be noted that inputting the speech features of the second pre-training Chinese pronunciation units into a decoder to obtain predicted Chinese text corresponding to the second pre-training Chinese pronunciation units, and pre-training the decoder based on the predicted Chinese text corresponding to the second pre-training Chinese pronunciation units and the first pre-training Chinese text, is similar in implementation process to inputting speech features into a decoder to obtain predicted Chinese text and pre-training a model including an encoder and a decoder in the above embodiment, and the specific implementation can refer to the relevant description of the above embodiment, and this embodiment will not be repeated here.
[0112] In the embodiments of the present specification, after obtaining a first pre-training text set, the first pre-training Chinese text is first converted into a second pre-training Chinese phonetic unit, and then the phonetic features of the second pre-training Chinese phonetic unit are determined according to a feature encoding layer. The phonetic features of the second pre-training speech data are then input to the decoder, and the decoder performs decoding using an autoregressive method, thereby becoming a decoder that takes into account the correlation between speech contexts when predicting text. Pre-training the decoder in this way can improve the accuracy of text prediction by the pre-trained decoder, and since the input when pre-training the decoder is Chinese phonetic unit, which is closer to speech data in modality and more suitable for downstream tasks, can improve the accuracy of the task when the decoder performs downstream tasks.
[0113] Please refer to Figure 6. Figure 6 shows a data flow diagram of a decoder pre-training provided by one embodiment of the present application. The decoder includes a decoding layer and a text embedding layer. A pronunciation unit embedding layer and a feature encoding layer of the encoder are further used in pre-training the decoder. The second pre-training Chinese pronunciation units are masked, and the masked second pre-training Chinese pronunciation units are input into the pronunciation unit embedding layer to obtain a feature matrix corresponding to the second pre-training Chinese pronunciation units. The feature matrix is input into the feature encoding layer of the encoder to obtain phonetic features of the second pre-training Chinese pronunciation units. The phonetic features are input into the decoding layer to obtain predicted text features. The predicted text features are input into the text embedding layer to obtain predicted Chinese text. The decoder parameters are adjusted according to the predicted Chinese text and the first pre-training Chinese text.
[0114] Furthermore, the above method pre-trains the decoder using Chinese phonetic units and Chinese text obtained by converting Chinese text, allowing the decoder to learn grammar rules for text construction and improve the decoder's language modeling ability. However, since speech recognition tasks are closely related to speech data, the decoder can also be pre-trained using speech and pseudo-label prediction tasks to improve the decoder's positioning and encapsulation capabilities for speech data.
[0115] That is, the cloud-side device 202: Obtaining a second pre-training speech set including a plurality of third pre-training speech data carrying the target pseudo-label; Encoding third pre-training speech data using an encoder to obtain speech features of the third pre-training speech data; inputting the speech features of the third pre-training speech data into a decoder to obtain predicted pseudo labels corresponding to the third pre-training speech data; It is further used to pre-train a decoder based on the target pseudo labels and the predicted pseudo labels.
[0116] In one embodiment herein, the second pre-training speech set may be obtained from an open source pre-training database.
[0117] In another embodiment of the present specification, the specific implementation of the cloud-side device 202 acquiring the second pre-training voice set is as follows: acquiring a plurality of unsupervised third pre-training speech data; Inputting a plurality of third pre-training speech data into a pre-trained speech encoder to obtain speech features of the plurality of third pre-training speech data; and clustering the speech features of the plurality of third pre-training speech data to obtain a target pseudo label for each third pre-training speech data.
[0118] That is, the pre-trained speech encoder may be used to label third pre-training speech data with a target pseudo label. Here, the target pseudo label is a label set for multiple pre-training speech data with high speech feature similarity and has no practical meaning. If two third pre-training speech data have the same target pseudo label, it indicates that the speech features of the two third pre-training speech data are highly similar.
[0119] As an example, the pre-trained speech encoder may be a speech encoder that is pre-trained based on a plurality of pre-training speech data, and may be used to extract speech features of the speech data.
[0120] In some embodiments, clustering the audio features of the plurality of third pre-training audio data includes determining a similarity between the audio features of each third pre-training audio data and the audio features of other third pre-training audio data, and clustering the third pre-training audio data whose similarity is greater than a similarity threshold and setting a target pseudo label for the category.
[0121] In the embodiment of the present specification, the trained speech encoder is used to determine the speech features of the third pre-training speech data, and the third pre-training speech data with similar speech features are clustered to set the target pseudo label, thereby obtaining a more accurate target pseudo label.
[0122] After the third pre-training speech data and the target pseudo-labels corresponding to each speech data are obtained, the third pre-training speech data is input to an encoder to extract speech features. Note that the implementation process of using an encoder to encode the third pre-training speech data and obtain speech features of the third pre-training speech data is similar to the implementation process of using an encoder to encode sample speech data and obtain speech features of the sample speech data. For specific implementation, please refer to the relevant description in the above embodiment, and this embodiment will not be repeated here.
[0123] In some embodiments herein, the decoder includes a decoding layer, and a specific implementation of inputting audio features to the decoder to obtain predicted pseudo labels corresponding to the third pre-training audio data may include inputting the audio features to the decoding layer of the decoder to obtain predicted text features corresponding to the third pre-training audio data, inputting the predicted text features corresponding to the plurality of third pre-training audio data to a pseudo-code embedding layer to determine a probability that there is a correspondence between the third pre-training audio data and each pseudo label, and determining the pseudo label with the highest probability as the predicted pseudo label corresponding to the third pre-training audio data.
[0124] It should be noted that the implementation process of determining predicted text features corresponding to the third pre-training speech data is similar to the implementation process of determining predicted text features in the above embodiment, and for its specific implementation, please refer to the relevant description of the above embodiment, and this embodiment will not be described again here.
[0125] After the predicted pseudo label is determined, a loss value is determined based on the target pseudo label and the predicted pseudo label using a sequence-to-sequence loss function. If the loss value is equal to or greater than a loss threshold, the decoder parameters are adjusted based on the loss value. If a pre-training stopping condition is reached, the decoder pre-training is stopped.
[0126] It should be noted that the pre-training stopping condition is the same as the above pre-training stopping condition for pre-training the encoder, and the implementation process is also similar. Therefore, for the specific implementation of pre-training the decoder based on the target pseudo-label and the predicted pseudo-label, please refer to the relevant description of the above embodiment, and this embodiment will not be repeated here.
[0127] In the embodiments herein, the decoder is pre-trained using a speech and pseudo-label prediction task, and the decoder performs decoding using an autoregressive method, so that information relevant to the decoder's generation of predicted text corresponding to the next pre-training speech data can be extracted from the encoder output, improving the decoder's positioning and encapsulation capabilities for the speech data.
[0128] Please refer to Figure 7. Figure 7 shows another data flow diagram of decoder pre-training provided by one embodiment of the present disclosure. The decoder includes a decoding layer and a pseudo-code embedding layer. Pre-training the decoder uses an encoder including a feature extraction layer, a speech encoding layer, and a feature encoding layer. A plurality of third pre-training speech data are obtained, the third pre-training speech data are input to the feature extraction layer to obtain speech representation vectors of the third pre-training speech data, the speech representation vectors are input to the speech encoding layer for encoding, and the encoding results are input to the feature encoding layer to obtain speech features of the third pre-training speech data, the speech features are input to the decoding layer to obtain predicted text features, and the predicted text features are input to the pseudo-code embedding layer to determine predicted pseudo labels corresponding to the third pre-training speech data, and decoder parameters are adjusted according to the predicted pseudo labels and the target pseudo labels.
[0129] In practical applications, the encoder and decoder can be pre-trained by combining the above-mentioned speech recognition task, speech mask prediction task, Chinese phonetic unit prediction task, text prediction task, and speech and pseudo-label prediction task. The loss values for all tasks are weighted and summed, and the encoder and decoder parameters are adjusted based on the resulting loss value. In this way, the effects of multiple tasks are taken into account during parameter adjustment, and by adjusting the parameters while ensuring the effectiveness of multiple tasks as much as possible, the pre-trained speech recognition model can be applied to multiple tasks, improving training efficiency. Furthermore, by adding a speech recognition task to the pre-training tasks, all tasks are influenced by the effect of the speech recognition task during the optimization process, and parameters are updated in the direction of higher speech recognition effectiveness, thereby improving the speech recognition performance of the pre-trained speech recognition model. Furthermore, spectral features are used as input to the encoder instead of speech data because spectral features ignore some phonetic details, making it difficult for the speech recognition model to distinguish data from different tasks. As a result, the parameter adjustments of the multi-task speech recognition model can be mutually constrained during combined training, avoiding the problem of model collapse.
[0130] In the above description, five pre-training tasks are combined to pre-train a model with an encoder and a decoder (i.e., pre-training the encoder and decoder) in order to reduce the difficulty of pre-training the speech recognition model and improve the training accuracy. However, in embodiments of the present application, before performing combined pre-training, a text prediction task may be used to first pre-train the encoder and decoder until convergence is achieved, thereby obtaining a model with an encoder and a decoder. Meanwhile, the text prediction task can be understood as a simplified version of the speech recognition task (i.e., the speech-text prediction task). Using the text prediction task as a pre-training task facilitates model training, lays a good foundation for subsequent combined pre-training, and makes the combined pre-training more stable. Meanwhile, since the phonetic unit embedding layer is initialized by learning the text prediction task, the problem of model collapse due to the speech mask prediction task during the combined pre-training stage is less likely to occur.
[0131] That is, the encoder includes a feature encoding layer, and the cloud-side device 202 includes: Obtaining a plurality of second pre-training pairs including third pre-training Chinese phonetic units and second pre-training Chinese texts; Inputting the third pre-training Chinese phonetic unit into a feature encoding layer to obtain phonetic features of the third pre-training Chinese phonetic unit; Inputting the phonetic features of the third pre-training Chinese phonetic unit into a decoder to obtain predicted Chinese text corresponding to the third pre-training Chinese phonetic unit; and pre-training a feature encoding layer and a decoder based on the predicted Chinese text corresponding to the third pre-training Chinese phonetic unit and the second pre-training Chinese text to obtain a model including an encoder and a decoder.
[0132] In the embodiments herein, the model is pre-trained using a supervised text prediction task, so that a plurality of second pre-training pairs including a third pre-training Chinese phonetic unit and a second pre-training Chinese text are directly obtained, and the third pre-training Chinese phonetic unit and the second pre-training Chinese text are directly obtained. Chinese phonetic units is used as the input of the feature encoding layer, and the second pre-training Chinese text is used as the label, and the feature encoding layer and decoder are pre-trained based on the predicted Chinese text output from the decoder and the label to obtain a model.
[0133] In some embodiments herein, the cloud-side device 202 may obtain the second pre-training pairs from an open-source pre-training dataset.
[0134] In some other embodiments of the present specification, the specific implementation of the cloud-side device 202 acquiring a plurality of second pre-trained pairs is as follows: obtaining a plurality of second pre-training Chinese texts; Converting a plurality of second pre-training Chinese texts into third pre-training Chinese phonetic units respectively; determining a second pre-training pair consisting of a second pre-training Chinese text and a corresponding third pre-training Chinese phonetic unit.
[0135] That is, since the speech recognition model predicts Chinese text based on speech data, its training data is usually speech data and Chinese text. However, since the Chinese text and speech data have large differences in modality, pre-training is performed by selecting Chinese pronunciation units that are close to the speech data in modality. Therefore, first, a plurality of second pre-training Chinese texts are obtained, and the plurality of second pre-training Chinese texts are converted into third pre-training Chinese pronunciation units. Then, the second pre-training Chinese texts and the third pre-training Chinese pronunciation units are converted into third pre-training Chinese pronunciation units. Chinese phonetic unitscan be configured as the second pre-training pair, in which case the accuracy rate of the correspondence between the second pre-training Chinese text and the third pre-training Chinese phonetic unit in the pre-training pair is very high. By pre-training the model based on the second pre-training pair, the model can acquire the ability to predict Chinese text based on the Chinese phonetic unit, and the speech recognition accuracy of the model is improved.
[0136] After the second pre-training pair is obtained, the third pre-training pair is Chinese phonetic units input into the feature encoding layer for processing. It is to be noted that the third pre-training Chinese phonetic unit is input into the feature encoding layer to obtain the phonetic features of the third pre-training Chinese phonetic unit, and the third pre-training Chinese phonetic units Inputting the speech features of the first to third pre-training Chinese phonetic units into the decoder to obtain predicted Chinese text corresponding to the third pre-training Chinese phonetic unit is similar in implementation process to pre-training the decoder using a text prediction task in the above embodiment, and for specific implementation, please refer to the relevant description of the above embodiment, and this embodiment will not be repeated here.
[0137] In some embodiments of the present specification, after the predicted Chinese text corresponding to the third pre-training Chinese phonetic unit is determined, a loss value is determined based on the predicted Chinese text and the second pre-training Chinese text. If the loss value is equal to or greater than the loss threshold, it indicates that the effects of the model's speech feature prediction and text prediction are poor, and therefore the pre-training of the feature encoding layer and the decoder continues until a pre-training stopping condition is reached.
[0138] Illustratively, the pre-training stopping condition may include a condition in which the loss value is less than a loss threshold, or the pre-training stopping condition may include a condition in which the number of pre-trainings is greater than or equal to a number threshold.
[0139] For example, if the loss value obtained after a certain pre-training is smaller than the loss threshold, it indicates that the model's speech feature prediction and text prediction are effective, i.e., the model can perfectly predict Chinese text based on Chinese phonetic units, and there is no need to continue training. Therefore, the model pre-training is stopped, i.e., the parameter adjustment of the feature encoding layer and decoder is stopped.
[0140] As another example, during the pre-training process, the number of pre-training sessions is recorded. Predicted Chinese text Each time is determined, the number of pre-training iterations may be increased by 1. If the number of pre-training iterations is greater than the threshold number, it indicates that the number of pre-training iterations of the model is sufficient, and even if pre-training is continued, better results may not be achieved. Therefore, pre-training the model is stopped, i.e., the parameter adjustment of the feature encoding layer and the decoder is stopped.
[0141] In the embodiments of this specification, before the encoder and decoder are pre-trained using five types of pre-training tasks, the feature encoding layer and decoder are pre-trained using text prediction tasks to obtain a model with an encoder and decoder, so that the feature encoding layer in the encoder has the ability to predict speech features and the decoder has the ability to predict text. Compared with speech data, Chinese phonetic units are less affected by speaker emotions, noise, etc., resulting in a better training effect. In addition, the model is pre-trained using text prediction tasks and the processing rules of the phonetic unit embedding layer are initialized in advance, making the subsequent pre-training process more stable.
[0142] It should be noted that the above description describes a process of pre-training a model having an encoder and a decoder using five types of pre-training tasks, specifically, firstly, using a text prediction task to pre-train the feature encoding layer and decoder of the encoder to obtain a model having an encoder and a decoder, and then combining a speech mask prediction task, a Chinese phonetic unit prediction task, a speech recognition task, a text prediction task, and a speech pseudo-label prediction task to pre-train a model having an encoder and a decoder to obtain a speech recognition model, and sending model parameters of the speech recognition model to the terminal-side device 201, and the terminal-side device 201 using the speech recognition model to recognize the speech data to be recognized and obtain target text corresponding to the speech data to be recognized.
[0143] Because the specific downstream tasks are different, the terminal-side device 201 can fine-tune the parameters of the voice recognition model before using the voice recognition model to recognize the recognition-target voice data.
[0144] That is, the terminal side device 201: Obtaining a check set including a plurality of audio check pairs including check audio data and corresponding check Chinese texts, and a plurality of Chinese pronunciation unit check pairs including check audio data and corresponding check Chinese pronunciation units; Using an encoder of the speech recognition model, perform Chinese pronunciation unit prediction on the check speech data, and obtain speech features of the check speech data and predicted Chinese pronunciation units; inputting the speech features of the checking speech data into a decoder of a speech recognition model to obtain a predicted Chinese text corresponding to the checking speech data; The speech recognition model is fine-tuned based on the predicted Chinese pronunciation unit, the check Chinese pronunciation unit, the predicted Chinese text, and the check Chinese text, and when a fine-tuning stopping condition is reached, a speech recognition model whose fine-tuning is completed is obtained.
[0145] In one or more embodiments of the present specification, fine-tuning of the speech recognition model is realized based on a supervised fine-tuning task. Because the downstream task of the speech recognition model is a speech recognition task, during fine-tuning, one supervised fine-tuning task is a speech recognition task. In addition, since the speech recognition task is a task of determining Chinese text based on speech data, it is necessary to obtain a speech check pair including check speech data and the corresponding check Chinese text. Furthermore, the encoder of the speech recognition model can generate speech features suitable for speech recognition. In Chinese speech recognition, the Chinese pronunciation unit associates the Chinese text with the speech data, i.e., both the Chinese text and the speech data can be uniquely mapped to the same Chinese pronunciation unit sequence. Therefore, by improving the encoder's Chinese pronunciation unit prediction ability, the encoder can generate speech features more suitable for speech recognition. Therefore, another supervised fine-tuning task, i.e., a Chinese pronunciation unit prediction task, can be set. In addition, since the Chinese pronunciation unit prediction task is a task of determining Chinese pronunciation units based on speech data, it is necessary to obtain a Chinese pronunciation unit check pair including check speech data and the corresponding check Chinese pronunciation unit.
[0146] As an example, the phonetic check pairs and Chinese pronunciation unit check pairs may be obtained from an open source check database, or the phonetic check pairs and Chinese pronunciation unit check pairs may be manually generated.
[0147] In a specific implementation, to obtain Chinese pronunciation units, a pronunciation unit embedding layer can be added to the speech recognition model, which is connected to the encoder and is for mapping the speech features as Chinese pronunciation units; in this way, when the speech features output from the encoder are input to the pronunciation unit embedding layer, the predicted Chinese pronunciation units corresponding to the checking speech data can be obtained.
[0148] In some embodiments of the present specification, the encoder may include a feature extraction layer, a speech encoding layer, and a feature encoding layer, and using the encoder of the speech recognition model to predict Chinese pronunciation units for the check speech data includes first inputting the check speech data into the feature extraction layer for speech expression extraction and downsampling processes to obtain a speech expression vector of the check speech data, then inputting the speech expression vector into the speech encoding layer and the feature encoding layer to obtain speech features of the check speech data obtained by combining the context speech of the speech data, and inputting the speech features of the check speech data into the pronunciation unit embedding layer to obtain predicted Chinese pronunciation units corresponding to the check speech data.
[0149] In addition, inputting the voice features of the checking voice data into the decoder of the voice recognition model to obtain the predicted Chinese text corresponding to the checking voice data is similar in implementation process to inputting the voice features into the decoder to obtain the predicted Chinese text, and for specific implementation, please refer to the relevant descriptions in the above embodiments, and this embodiment will not be repeated here.
[0150] In some embodiments of the present specification, after the predicted Chinese pronunciation unit and the checking Chinese text are obtained, the predicted Chinese pronunciation unit and the checking Chinese text are For checking A first loss value may be determined based on the Chinese pronunciation unit, a second loss value may be determined based on the predicted Chinese text and the check Chinese text, and a third loss value may be obtained by adding the first loss value and the second loss value. If the third loss value is equal to or greater than a loss threshold, the parameters of the speech recognition model (including the parameters of the decoder and encoder) are fine-tuned based on the loss value, and the encoder of the speech recognition model is used to predict the Chinese pronunciation unit for the check speech data. The process returns to the step of obtaining the speech features of the check speech data and the predicted Chinese pronunciation unit. If a fine-tuning stop condition is reached, the fine-tuning of the parameters of the speech recognition model is stopped, and a fine-tuned speech recognition model is obtained.
[0151] In some examples herein, the fine-tuning stopping condition may be a condition where the loss value is smaller than a loss threshold, or a condition where the number of iterative fine-tuning iterations is equal to or greater than a number threshold.
[0152] As an example, if the loss value obtained after a certain fine-tuning is smaller than the loss threshold, it indicates that the speech recognition model has already achieved perfect speech recognition and there is no need to further adjust the parameters. Therefore, the fine-tuning is stopped and a fine-tuned speech recognition model is obtained.
[0153] As another example, the number of iterations and fine-tuning times is recorded to predict Chinese text is determined each time, 1 may be added to the number of iterative fine-tuning iterations. If the number of iterative fine-tuning iterations is greater than the threshold number, it indicates that the number of times the parameters of the speech recognition model have been fine-tuned is sufficient, and even if the fine-tuning is continued, it may not be possible to achieve a better effect. Therefore, the fine-tuning is stopped and a speech recognition model that has been fine-tuned is obtained.
[0154] In the embodiments of this specification, a speech recognition model is fine-tuned using a speech recognition task and a Chinese phonetic unit prediction task, i.e., two types of supervised tasks are combined to fine-tune the speech recognition model. In this way, the fine-tuning of the parameters of the speech recognition model is influenced by two tasks, which not only improves the training efficiency of the speech recognition model and the recognition accuracy of the speech recognition model, but also makes the speech recognition model suitable for more downstream tasks, thereby improving the applicability of the trained speech recognition model.
[0155] In another optional embodiment of the present specification, the specific implementation of the terminal side device 201 obtaining the check set is as follows: Obtaining a plurality of audio check pairs including audio data for checking and Chinese text for checking; Converting each Chinese text for checking into a Chinese pronunciation unit to obtain a Chinese pronunciation unit for checking corresponding to each voice data for checking; determining a Chinese pronunciation unit check pair consisting of the check speech data and the corresponding check Chinese pronunciation unit; determining a check set consisting of a plurality of phonetic check pairs and a plurality of Chinese phonetic unit check pairs.
[0156] That is, first, a plurality of audio check pairs can be obtained, and then the check Chinese text in the audio check pairs can be converted into check Chinese pronunciation units. Since there is a correspondence between the check Chinese text and the check audio data, there is also a correspondence between the check Chinese pronunciation units and the check audio data. The check Chinese pronunciation units and the corresponding check audio data constitute a Chinese pronunciation unit check pair, and a plurality of Chinese pronunciation unit check pairs and a plurality of audio check pairs constitute a check set.
[0157] For example, the correspondence between the Chinese text and the Chinese pronunciation unit may be determined based on a dictionary, and the checking Chinese text may be converted into the checking Chinese pronunciation unit accordingly; or a model capable of realizing the text-Chinese pronunciation unit conversion may be pre-trained and obtained, and the checking Chinese text may be converted into the checking Chinese pronunciation unit using the model. Chinese phonetic units It may be converted into:
[0158] For example, the Chinese pronunciation unit may be Pinyin, syllable, or Chinese phoneme. In the process of converting the Chinese text for checking into the Chinese pronunciation unit for checking, it is necessary to not only obtain the Pinyin of the Chinese text through the conversion, but also obtain the one containing the pronunciation tones (e.g., the first tone, the second tone, the third tone, and the fourth tone). Furthermore, in order to avoid confusion between different words, it is also necessary to divide the Pinyin of different words into initial consonants and final consonants.
[0159] For example, when it is assumed that the Chinese text for checking is "今天天气真不錯,但下午可能下雨 (Japanese translation: 今日はいい天気だが、午後は雨が降るかもしれません。)", the Chinese pronunciation units for checking may be "j in1 t ian1 t ian1 qi4 zh en1 b u2 c uo4 d an4 x ia4 w u3 k e3 n eng2 x ia4 y u3". Here, the numbers represent tones, where 1 represents the first tone, 2 represents the second tone, 3 represents the third tone, and 4 represents the fourth tone.
[0160] In such a case, if there is a correspondence relationship among the three: the voice data for checking, the Chinese pronunciation units for checking, and the Chinese text for checking, it is understood that there is a correspondence relationship among the check pairs used in the speech recognition task and the Chinese pronunciation unit prediction task, and thereby the training accuracy of the combination fine-tuning can be improved.
[0161] Please refer to FIG. 8. FIG. 8 shows a data flow diagram of a method for fine-tuning an audio recognition model provided by an embodiment of this specification. The audio recognition model includes an encoder, a pronunciation unit embedding layer, and a decoder. The encoder includes a feature extraction layer, an audio encoding layer, and a feature encoding layer. The decoder includes a decoding layer and a text embedding layer. Input the voice data for checking into the feature extraction layer to obtain the voice representation vector of the voice data for checking. Input the voice representation vector into the audio encoding layer to encode it, input the encoding result into the feature encoding layer to obtain the voice features, input the voice features into the pronunciation unit embedding layer to obtain the predicted Chinese pronunciation units, input the voice features into the decoding layer to obtain the predicted text features, input the predicted text features into the text embedding layer to obtain the predicted Chinese text. After obtaining the predicted Chinese pronunciation units and the predicted Chinese text, the predicted Chinese pronunciation units, the Chinese pronunciation units for checking, the predicted Chinese text, and For checkingBased on the Chinese text, the encoder and decoder parameters in the speech recognition model are fine-tuned (i.e., adjusted), and when a fine-tuning stopping condition is reached, a fine-tuned speech recognition model is obtained.
[0162] The pre-training of the speech recognition model according to the present solution includes three stages. In the first stage, the entire model is trained using a text prediction task to obtain a model with an encoder and a decoder. In the second stage, the encoder is pre-trained using a Chinese phonetic unit prediction task and a speech mask prediction task, the decoder is pre-trained using a text prediction task and a speech / pseudo-label prediction task, and the entire model is trained using a supervised speech recognition task. These five tasks may be performed individually or simultaneously to obtain a speech recognition model. In the third stage, the parameters of the speech recognition model are fine-tuned using a supervised speech recognition task and a supervised Chinese phonetic unit prediction task to obtain a fine-tuned speech recognition model. Furthermore, the first and second stages may be performed by the cloud-side device 202, and the third stage may be performed by the terminal-side device 201. Alternatively, all three stages may be performed by the cloud-side device 202.
[0163] In addition, the text prediction task is used in the first stage because Chinese phonetic units have less interference than speech data, and by training the text prediction task, the usage rules of the phonetic unit embedding layer are initialized in advance, making the pre-training in the second stage more stable. In the second stage, five tasks are used in combination for training, which improves training efficiency. In the third stage, the downstream speech recognition task is added in advance to the fine-tuning task, so that all tasks are influenced by the speech recognition task during the model parameter optimization process, and the parameters are updated in a direction that improves the speech recognition effect. The effect of downstream tasks can be evaluated in advance, which improves work efficiency.
[0164] The solution applied in the embodiments of this specification pre-trains the encoder and decoder before training a speech recognition model. This reduces the number of sample speech data and sample Chinese texts required for training to obtain a speech recognition model, thereby reducing the burden on labelers and the difficulty of obtaining labeled data. Due to the characteristic of Chinese text being an ideographic language, i.e., there are large differences between speech data and Chinese text, and the same pronunciation can correspond to approximately 100 Chinese characters, the solution adds the modality of Chinese pronunciation units to the model pre-training process. This is because Chinese pronunciation units serve as a bridge to establish the relationship between speech data and Chinese text, i.e., both speech data and Chinese text can be uniquely mapped to the same Chinese pronunciation unit sequence. During pre-training, the encoder is pre-trained by performing a speech mask prediction task and a Chinese pronunciation unit prediction task on the pre-training speech data. Both tasks map speech data into a Chinese pronunciation unit sequence, allowing the encoder to capture the pronunciation information of the speech data and contribute to speech recognition. In addition, by performing a text prediction task and a speech / pseudo-label prediction task on the pre-training Chinese pronunciation units, decoderThe decoder is pre-trained, and the decoder is capable of constructing text using speech features, thereby improving the decoder's language modeling ability. Because the encoder and decoder are pre-trained to have a certain level of speech recognition ability, retraining the pre-trained encoder and decoder can improve training efficiency and training accuracy. Furthermore, the model input used in pre-training is pre-training speech data or pre-training Chinese phonetic units, both of which are similar in modality to the speech data input when the speech recognition model is applied. This improves recognition accuracy when the speech recognition model is used to recognize target speech data. Furthermore, by using a large amount of low-cost unlabeled speech data and unlabeled Chinese text, and only using a small amount of speech-text labeled data, a speech recognition model for the Chinese language can be trained with high accuracy, thereby reducing the need for labeled data, reducing personnel costs, and improving training efficiency.
[0165] FIG. 9 shows a flowchart of a data processing method for a voice recognition model applied to a cloud-side device provided in one embodiment of this specification, where the cloud-side device is connected to multiple terminal-side devices. The data processing method for the voice recognition model specifically includes the following steps:
[0166] In step 902, a sample set including a plurality of sample pairs each including sample speech data and sample Chinese text is obtained.
[0167] In step 904, the sample speech data is encoded using an encoder that is pre-trained by performing a Chinese phonetic unit prediction task on the pre-training speech data to obtain speech features of the sample speech data.
[0168] In step 906, the phonetic features are input to a decoder that is pre-trained by performing a text prediction task on the pre-training Chinese phonetic units to obtain predicted Chinese text.
[0169] In step 908, a model including an encoder and a decoder is pre-trained based on the predicted Chinese text and the sample Chinese text, and when a pre-training stopping condition is reached, model parameters of the pre-trained speech recognition model are obtained.
[0170] In step 910, the model parameters of the speech recognition model obtained by pre-training are transmitted to a first terminal-side device, which is one of the plurality of terminal-side devices.
[0171] In one or more embodiments herein, Chinese phonetic units are used for pre-training speech data. Prediction Task A specific implementation of pre-training the encoder by running Obtaining a first pre-training speech set including a plurality of unsupervised first pre-training speech data; Using an encoder, encode first pre-training speech data to obtain first speech features corresponding to the first pre-training speech data, and determine first pronunciation units based on the first speech features; masking the first pre-training speech data; using an encoder to encode the masked first pre-training speech data, obtain second speech features corresponding to the masked first pre-training speech data, and determine second pronunciation units based on the second speech features; pre-training the encoder based on the first phonetic unit and the second phonetic unit corresponding to the first pre-training speech data.
[0172] In one or more embodiments herein, before using an encoder to encode first pre-training speech data and obtain first speech features corresponding to the first pre-training speech data, further comprising extracting spectral features of the first pre-training speech data; Encoding first pre-training speech data using an encoder to obtain first speech features corresponding to the first pre-training speech data includes: Inputting spectral features of the first pre-training speech data into an encoder to obtain first speech features corresponding to the first pre-training speech data.
[0173] In one or more embodiments herein, Chinese phonetic units are used for pre-training speech data. Prediction Task A specific implementation of pre-training the encoder by running Obtaining a plurality of first pre-training pairs including second pre-training speech data and first pre-training Chinese phonetic units; Using an encoder to predict Chinese pronunciation units for the second pre-training speech data, and obtain predicted Chinese pronunciation units corresponding to the second pre-training speech data; pre-training the encoder based on the first pre-training Chinese phonetic unit and the predicted Chinese phonetic unit.
[0174] In one or more embodiments herein, the encoder comprises a feature encoding layer, and a specific implementation of pre-training the decoder by performing a text prediction task on the pre-training Chinese phonetic units includes: Obtaining a first pre-training text set including a plurality of unsupervised first pre-training Chinese texts; Converting the first pre-training Chinese text into a second pre-training Chinese phonetic unit, and inputting the second pre-training Chinese phonetic unit into a feature encoding layer to obtain phonetic features of the second pre-training Chinese phonetic unit; Inputting the phonetic features of the second pre-training Chinese phonetic unit into a decoder to obtain predicted Chinese text corresponding to the second pre-training Chinese phonetic unit; and pre-training the decoder based on the predicted Chinese text corresponding to the second pre-training Chinese phonetic unit and the first pre-training Chinese text.
[0175] In one or more embodiments herein, a specific implementation of pre-training a decoder by performing a text prediction task on pre-training Chinese phonetic units includes: Obtaining a second pre-training speech set including a plurality of third pre-training speech data carrying the target pseudo-label; Encoding third pre-training speech data using an encoder to obtain speech features of the third pre-training speech data; inputting the speech features of the third pre-training speech data into a decoder to obtain predicted pseudo labels corresponding to the third pre-training speech data; and pre-training a decoder based on the target pseudo labels and the predicted pseudo labels.
[0176] In one or more embodiments herein, a specific implementation of obtaining the second pre-training speech set may include: acquiring a plurality of unsupervised third pre-training speech data; Inputting a plurality of third pre-training speech data into a pre-trained speech encoder to obtain speech features of the plurality of third pre-training speech data; and clustering the speech features of the plurality of third pre-training speech data to obtain a target pseudo label for each third pre-training speech data.
[0177] In one or more embodiments herein, the encoder comprises a feature encoding layer; Obtaining a plurality of second pre-training pairs including third pre-training Chinese phonetic units and second pre-training Chinese texts; Inputting the third pre-training Chinese phonetic unit into a feature encoding layer to obtain phonetic features of the third pre-training Chinese phonetic unit; Inputting the phonetic features of the third pre-training Chinese phonetic unit into a decoder to obtain predicted Chinese text corresponding to the third pre-training Chinese phonetic unit; and pre-training a feature encoding layer and a decoder based on the predicted Chinese text corresponding to the third pre-training Chinese phonetic unit and the second pre-training Chinese text to obtain a model including an encoder and a decoder.
[0178] It should be noted that the specific implementation of the data processing method of the voice recognition model applied to the cloud-side device is the same as the operation performed by the cloud-side device in the data processing system of the above voice recognition model. For the specific implementation, please refer to the relevant description of the above embodiment, and this embodiment will not be repeated here.
[0179] In the solution applied in the embodiments of the present specification, the encoder and decoder are pre-trained before obtaining a speech recognition model. This reduces the number of sample speech data and sample Chinese texts required for training to obtain a speech recognition model, thereby reducing the burden on labelers and the difficulty of obtaining labeled data. Given the characteristic that Chinese data is an ideographic language, i.e., there is a large discrepancy between speech and text, and the same pronunciation may correspond to approximately 100 Chinese characters, the modality of pronunciation units is added to the model pre-training process. This is because pronunciation units serve as a bridge between speech and text, i.e., both speech and text can be uniquely mapped to the same pronunciation unit sequence. During pre-training, a speech mask prediction task and a pronunciation unit prediction task are performed on the pre-training speech data to obtain a pre-trained encoder. Both tasks map speech data into a pronunciation unit sequence, allowing the encoder to capture pronunciation information in the speech signal and contribute to speech recognition. In addition, by performing a text prediction task on pre-training Chinese phonetic units to obtain a decoder, the decoder is able to construct text using phonetic features, thereby improving the decoder's language modeling ability. Because the encoder and decoder are pre-trained to have certain speech recognition capabilities, re-training the pre-trained encoder and decoder can improve training efficiency and training accuracy. Furthermore, the model input used in pre-training is pre-training speech data or pre-training Chinese phonetic units, both of which are similar in modality to the speech data input when the speech recognition model is applied. This improves recognition accuracy when the speech recognition model recognizes target speech data. Furthermore, by using a large amount of low-cost unlabeled speech data and unlabeled Chinese text, and only using a small amount of speech-text labeled data, a speech recognition model for the Chinese language can be trained with high accuracy, thereby reducing the need for labeled data, reducing personnel costs, and improving training efficiency.
[0180] FIG. 10 shows a flowchart of a speech recognition method applied to a terminal-side device provided in one embodiment of the present specification, where the terminal-side device is connected to a cloud-side device, and the speech recognition method specifically includes the following steps:
[0181] In step 1002, speech data to be recognized is acquired.
[0182] In step 1004, the encoder of the speech recognition model is used to encode the speech data to be recognized, and the speech features of the speech data to be recognized are obtained. The speech recognition model is pre-trained by the cloud-side device using the data processing method of the speech recognition model.
[0183] In one or more embodiments of the present disclosure, an encoder for a speech recognition model is used, and before encoding the speech data to be recognized, The method further includes obtaining a check set including a plurality of speech check pairs including the check speech data and corresponding Chinese check text, and a plurality of Chinese pronunciation unit check pairs including the check speech data and corresponding Chinese pronunciation units; using an encoder of the speech recognition model to perform Chinese pronunciation unit prediction on the check speech data, and obtaining speech features of the check speech data and predicted Chinese pronunciation units; inputting the speech features of the check speech data into a decoder of the speech recognition model to obtain predicted Chinese text corresponding to the check speech data; fine-tuning the speech recognition model based on the predicted Chinese pronunciation units, the check Chinese pronunciation units, the predicted Chinese text, and the check Chinese text, and obtaining a speech recognition model whose fine-tuning has been completed if a fine-tuning stopping condition is reached.
[0184] After the fine-tuning is completed, the speech recognition model that has been fine-tuned can be used to recognize the speech data to be recognized, thereby obtaining the target text corresponding to the speech data to be recognized.
[0185] In step 1006, the speech features are input to a decoder of the speech recognition model to obtain target text corresponding to the speech data to be recognized.
[0186] In one or more embodiments of the present specification, please refer to Figure 11. Figure 11 shows a data flow diagram of the execution of a speech recognition task by a speech recognition model provided by one embodiment of the present specification. The speech recognition model includes an encoder and a decoder, and speech data to be recognized is input to the encoder of the speech recognition model to obtain speech features of the speech data to be recognized. The decoder includes a decoding layer and a text embedding layer, and the speech features are input to the decoding layer to obtain predicted text features, and the predicted text features are input to the text embedding layer to output target text corresponding to the speech data to be recognized.
[0187] In one or more embodiments of the present specification, a specific implementation of obtaining recognition target speech data includes: receiving a speech recognition request carrying speech data to be recognized; obtaining recognition target speech data from the speech recognition request; Accordingly, after step 1006, sending the target text to the front end for display; receiving a correction text corresponding to a target text input by a user at the front end; The method further includes updating the speech recognition model based on the corrected text and the recognition target speech data to obtain an updated speech recognition model.
[0188] As an example, after feeding back the target text to the front end for display, the user can proofread the target text and receive a corrected text corresponding to the target text input by the user at the front end, and then update the speech recognition model based on the corrected text and the speech data to be recognized, thereby improving the speech recognition accuracy of the speech recognition model.
[0189] It should be noted that the specific implementation of the voice recognition method applied to the terminal-side device is the same as the operations performed by the terminal-side device in the data processing system of the above voice recognition model. For the specific implementation, please refer to the relevant descriptions in the above embodiments, and this embodiment will not be described again here.
[0190] In the solution applied in the embodiments of the present specification, the encoder and decoder are pre-trained before obtaining a speech recognition model. This reduces the number of sample speech data and sample Chinese texts required for training to obtain a speech recognition model, thereby reducing the burden on labelers and the difficulty of obtaining labeled data. Given the characteristic of Chinese data being an ideographic language, i.e., there is a large discrepancy between speech and text, and the same pronunciation can correspond to approximately 100 Chinese characters, the modality of pronunciation units is added to the model pre-training process. This is because pronunciation units serve as a bridge between speech and text, i.e., both speech and text can be uniquely mapped to the same pronunciation unit sequence. During pre-training, a speech mask prediction task and a pronunciation unit prediction task are performed on the pre-training speech data to obtain an encoder. Both tasks map speech data into a pronunciation unit sequence, allowing the encoder to capture pronunciation information in the speech signal and contribute to speech recognition. In addition, by performing a text prediction task on pre-training Chinese phonetic units to obtain a decoder, the decoder is able to construct text using phonetic features, thereby improving the decoder's language modeling ability. Because the encoder and decoder are pre-trained to have certain speech recognition capabilities, re-training the pre-trained encoder and decoder can improve training efficiency and training accuracy. Furthermore, the model input used in pre-training is pre-training speech data or pre-training Chinese phonetic units, both of which are similar in modality to the speech data input when the speech recognition model is applied. This improves recognition accuracy when the speech recognition model recognizes target speech data. Furthermore, by using a large amount of low-cost unlabeled speech data and unlabeled Chinese text, and only using a small amount of speech-text labeled data, a speech recognition model for the Chinese language can be trained with high accuracy, thereby reducing the need for labeled data, reducing personnel costs, and improving training efficiency.
[0191] Hereinafter, the data processing method for a speech recognition model provided by this specification will be further described by taking the application of the data processing method for a speech recognition model for Chinese language as an example, with reference to Figure 12. Figure 12 shows a flowchart of the processing process of the data processing method for a speech recognition model provided by an embodiment of this specification, specifically including the following steps:
[0192] In step 1202, a plurality of pre-training pairs including pre-training Chinese phonetic units and pre-training Chinese texts are obtained.
[0193] In step 1204, the pre-training Chinese phonetic units are masked to obtain masked pre-training Chinese phonetic units.
[0194] In step 1206, the masked pre-training Chinese phonetic units are input into a phonetic unit embedding layer to obtain a feature matrix corresponding to the masked pre-training Chinese phonetic units, and the feature matrix is input into a feature encoding layer to obtain phonetic features of the masked pre-training Chinese phonetic units.
[0195] In step 1208, the speech features are input into a decoding layer to obtain predicted text features, and the predicted text features are input into a text embedding layer to obtain predicted Chinese text.
[0196] In step 1210, a model is pre-trained based on the pre-training Chinese text and the predicted Chinese text to obtain a model with an encoder and a decoder.
[0197] Illustratively, steps 1202 to 1210 are a first-stage process of pre-training a model including an encoder and a decoder using a text prediction task.
[0198] In step 1212, a first Chinese phonetic unit corresponding to the pre-training speech data is obtained.
[0199] In step 1214, the pre-training speech data is input to the feature extraction layer to obtain speech representation vectors, and the speech representation vectors are masked.
[0200] In step 1216, the masked phonetic representation vector is input into a phonetic encoding layer and a feature encoding layer to obtain a second phonetic feature, and the second phonetic feature is input into a phonetic unit embedding layer to obtain a second Chinese phonetic unit.
[0201] In step 1218, a loss value is determined based on the first Chinese phonetic unit and the second Chinese phonetic unit, and an encoder is pre-trained based on the loss value.
[0202] Illustratively, steps 1212 to 1218 are a second stage process of pre-training an encoder using an audio mask prediction task.
[0203] In step 1220, pre-training Chinese phonetic units corresponding to pre-training speech data are obtained, and the pre-training speech data are input into a feature extraction layer to obtain the phonetic expression vectors of the pre-training speech data.
[0204] In step 1222, the phonetic representation vectors are input into the phonetic encoding layer and the feature encoding layer to obtain predicted Chinese phonetic units corresponding to the pre-training phonetic data.
[0205] In step 1224, the encoder is pre-trained based on the predicted Chinese phonetic units and the pre-training Chinese phonetic units.
[0206] Illustratively, steps 1220 to 1224 are a process of pre-training an encoder using a Chinese phonetic unit prediction task in the second stage.
[0207] In step 1226, a pre-training Chinese text corresponding to the pre-training Chinese phonetic unit is obtained, and the pre-training Chinese phonetic unit is masked.
[0208] In step 1228, the masked pre-training Chinese phonetic units are input into the phonetic unit embedding layer to obtain a feature matrix corresponding to the pre-training Chinese phonetic units, and the feature matrix is input into the feature encoding layer to obtain the phonetic features of the pre-training Chinese phonetic units.
[0209] In step 1230, the phonetic features are input into a decoding layer to obtain predicted text features, and the predicted text features are input into a text embedding layer to obtain predicted Chinese text.
[0210] In step 1232, a decoder is pre-trained based on the predicted Chinese text and the pre-training Chinese text.
[0211] Illustratively, steps 1226 to 1232 are a second stage process of pre-training a decoder using a text prediction task.
[0212] In step 1234, pre-training speech data is obtained, and the pre-training speech data is input into a feature extraction layer to obtain speech representation vectors of the pre-training speech data.
[0213] In step 1236, the speech representation vector is input to a speech encoding layer and a feature encoding layer to obtain speech features of the pre-training speech data.
[0214] In step 1238, the audio features are input to a decoding layer to obtain predicted text features, and the predicted text features are input to a pseudo-code embedding layer to determine predicted pseudo labels corresponding to the pre-training audio data.
[0215] In step 1240, a decoder is pre-trained according to the predicted pseudo labels and the target pseudo labels.
[0216] Illustratively, steps 1234 to 1240 are a process of pre-training a decoder using a speech and pseudo-label prediction task in the second stage.
[0217] In step 1242, sample speech data and sample Chinese text are obtained, and the sample speech data is input into a feature extraction layer to obtain a speech representation vector of the sample speech data.
[0218] In step 1244, the phonetic representation vector is input to a phonetic encoding layer and a feature encoding layer to obtain phonetic features.
[0219] In step 1246, the speech features are input into a decoding layer to obtain predicted text features, and the predicted text features are input into a text embedding layer to obtain predicted Chinese text.
[0220] In step 1248, based on the predicted Chinese text and the sample Chinese text, Model with Encoder and Decoder of advance Train.
[0221] Illustratively, steps 1242 to 1248 are a second stage process of pre-training a model including an encoder and a decoder using a speech recognition task.
[0222] In step 1250, a speech check pair including the check speech data and the corresponding check Chinese text, and a Chinese pronunciation unit check pair including the check speech data and the corresponding check Chinese pronunciation unit are obtained.
[0223] In step 1252, the checking voice data is input to the feature extraction layer to obtain the voice expression vector of the checking voice data.
[0224] In step 1254, the phonetic representation vector is input into a phonetic encoding layer and a feature encoding layer to obtain phonetic features, and the phonetic features are input into a Chinese phonetic unit embedding layer to obtain predicted Chinese phonetic units.
[0225] In step 1256, the audio features can be input into a decoding layer to obtain predicted text features, and the predicted text features can be input into a text embedding layer to obtain predicted Chinese text.
[0226] In step 1258, the predicted Chinese phonetic unit, the check Chinese phonetic unit, the predicted Chinese text, and For checking Fine-tune the parameters of the speech recognition model based on Chinese text.
[0227] Illustratively, steps 1250 to 1258 are a process in the third stage for fine-tuning the parameters of a speech recognition model using a speech recognition task and a Chinese phonetic unit prediction task.
[0228] Please refer to Figure 13. Figure 13 shows a data flow diagram for the combined training of a speech recognition model provided by an embodiment of the present application. In the diagram, line 1 represents the data flow for the speech mask prediction task, line 2 represents the data flow for the Chinese phonetic unit prediction task, line 3 represents the data flow for the speech and pseudo-label prediction task, line 4 represents the data flow for the speech recognition task, and line 5 represents the data flow for the text prediction task. For the specific process of data flow in each task, please refer to the relevant descriptions of Figures 3, 4a, 4b, 5, 6, and 7 above, and this embodiment will not be repeated here.
[0229] In the solution applied in the embodiments of the present specification, the encoder and decoder are pre-trained before obtaining a speech recognition model. This reduces the number of sample speech data and sample Chinese texts required for training to obtain a speech recognition model, thereby reducing the burden on labelers and the difficulty of obtaining labeled data. Given the characteristic that Chinese data is an ideographic language, i.e., there is a large discrepancy between speech and text, and the same pronunciation may correspond to approximately 100 Chinese characters, the modality of pronunciation units is added to the model pre-training process. This is because pronunciation units serve as a bridge between speech and text, i.e., both speech and text can be uniquely mapped to the same pronunciation unit sequence. During pre-training, a speech mask prediction task and a pronunciation unit prediction task are performed on the pre-training speech data to obtain a pre-trained encoder. Both tasks map speech data into a pronunciation unit sequence, allowing the encoder to capture pronunciation information in the speech signal and contribute to speech recognition. In addition, by performing a text prediction task on pre-training Chinese phonetic units to obtain a decoder, the decoder is able to construct text using phonetic features, thereby improving the decoder's language modeling ability. Because the encoder and decoder are pre-trained to have certain speech recognition capabilities, re-training the pre-trained encoder and decoder can improve training efficiency and training accuracy. Furthermore, the model input used in pre-training is pre-training speech data or pre-training Chinese phonetic units, both of which are similar in modality to the speech data input when the speech recognition model is applied. This improves recognition accuracy when the speech recognition model recognizes target speech data. Furthermore, by using a large amount of low-cost unlabeled speech data and unlabeled Chinese text, and only using a small amount of speech-text labeled data, a speech recognition model for the Chinese language can be trained with high accuracy, thereby reducing the need for labeled data, reducing personnel costs, and improving training efficiency.
[0230] Corresponding to the embodiment of the data processing method for the voice recognition model applied to the cloud-side device, this specification further provides an embodiment of a data processing device for the voice recognition model applied to the cloud-side device, and Figure 14 shows a structural schematic diagram of a data processing device for the voice recognition model applied to the cloud-side device provided by one embodiment of this specification. As shown in Figure 14, the device: a first acquiring module 1402 configured to acquire a sample set including a plurality of sample pairs including sample speech data and sample Chinese text; a first encoding module 1404 configured to encode sample speech data using an encoder that is pre-trained by performing a Chinese phonetic unit prediction task on pre-training speech data to obtain speech features of the sample speech data; a first decoding module 1406 configured to input the speech features into a decoder that is pre-trained by performing a text prediction task on pre-training Chinese phonetic units to obtain predicted Chinese text; a pre-training module 1408 configured to pre-train a model including an encoder and a decoder based on the predicted Chinese text and the sample Chinese text, and obtain model parameters of the pre-trained speech recognition model when a pre-training stopping condition is reached; and a first transmitting module 1410 configured to transmit model parameters of the pre-trained speech recognition model to a first terminal-side device, which is one of the plurality of terminal-side devices.
[0231] In one or more embodiments herein, the apparatus further comprises an encoder pre-training module, the encoder pre-training module comprising: obtaining a first pre-training speech set including a plurality of unsupervised first pre-training speech data; Encoding first pre-training speech data using an encoder to obtain first speech features corresponding to the first pre-training speech data; and determining first pronunciation units based on the first speech features; masking the first pre-training speech data; Using an encoder to encode the masked first pre-training speech data, obtain second speech features corresponding to the masked first pre-training speech data, and determine second pronunciation units based on the second speech features; The encoder is configured to pre-train based on the first phonetic unit and the second phonetic unit corresponding to the first pre-training speech data.
[0232] In one or more embodiments herein, the encoder pre-training module comprises: Extracting spectral features from the first pre-training speech data; The apparatus is further configured to input the spectral features of the first pre-training speech data to an encoder to obtain first speech features corresponding to the first pre-training speech data.
[0233] In one or more embodiments herein, the encoder pre-training module comprises: Obtain a plurality of first pre-training pairs including second pre-training speech data and first pre-training Chinese phonetic units; Using an encoder to predict Chinese phonetic units for the second pre-training speech data, and obtain predicted Chinese phonetic units corresponding to the second pre-training speech data; The encoder is configured to pre-train based on the first pre-training Chinese phonetic unit and the predicted Chinese phonetic unit.
[0234] In one or more embodiments herein, the encoder comprises a feature encoding layer; The apparatus further comprises a decoder pre-training module, the decoder pre-training module comprising: Obtain a first pre-training text set including a plurality of unsupervised first pre-training Chinese texts; Converting the first pre-training Chinese text into a second pre-training Chinese phonetic unit, and inputting the second pre-training Chinese phonetic unit into a feature encoding layer to obtain phonetic features of the second pre-training Chinese phonetic unit; Inputting the phonetic features of the second pre-training Chinese phonetic unit into a decoder to obtain predicted Chinese text corresponding to the second pre-training Chinese phonetic unit; The decoder is configured to pre-train based on the predicted Chinese text corresponding to the second pre-training Chinese phonetic unit and the first pre-training Chinese text.
[0235] In one or more embodiments herein, the decoder pre-training module comprises: Obtain a second pre-training speech set including a plurality of third pre-training speech data carrying the target pseudo-label; Encoding third pre-training speech data using an encoder to obtain speech features of the third pre-training speech data; inputting the speech features of the third pre-training speech data into a decoder to obtain predicted pseudo-labels corresponding to the third pre-training speech data; The decoder is configured to pre-train based on the target pseudo labels and the predicted pseudo labels.
[0236] In one or more embodiments herein, the decoder pre-training module comprises: Obtain multiple unsupervised third-stage pre-training speech data sets, inputting a plurality of third pre-training speech data into a pre-trained speech encoder to obtain speech features of the plurality of third pre-training speech data; The method is further configured to cluster the speech features of the plurality of third pre-training speech data to obtain a target pseudo label for each third pre-training speech data.
[0237] In one or more embodiments herein, the encoder comprises a feature encoding layer; The first acquisition module includes: Obtaining a plurality of second pre-training pairs including third pre-training Chinese phonetic units and second pre-training Chinese texts; Input the third pre-trained Chinese phonetic unit into a feature encoding layer to obtain phonetic features of the third pre-trained Chinese phonetic unit; Inputting the phonetic features of the third pre-training Chinese phonetic unit into a decoder to obtain predicted Chinese text corresponding to the third pre-training Chinese phonetic unit; It is further configured to pre-train a feature encoding layer and a decoder based on the predicted Chinese text corresponding to the third pre-training Chinese phonetic unit and the second pre-training Chinese text to obtain a model including an encoder and a decoder.
[0238] In the solution applied in the embodiments of this specification, the encoder and decoder are pre-trained before obtaining a speech recognition model. This reduces the number of sample speech data and sample Chinese texts required for training to obtain a speech recognition model, thereby reducing the burden on labelers and the difficulty of obtaining labeled data. Given the characteristic of Chinese data being an ideographic language, i.e., there is a large discrepancy between speech and text, and the same pronunciation can correspond to approximately 100 Chinese characters, the modality of pronunciation units is added to the model pre-training process. This is because pronunciation units serve as a bridge between speech and text, i.e., both speech and text can be uniquely mapped to the same pronunciation unit sequence. During pre-training, a speech mask prediction task and a pronunciation unit prediction task are performed on the pre-training speech data to obtain an encoder. Both tasks map speech data into a pronunciation unit sequence, allowing the encoder to capture pronunciation information in the speech signal and contribute to speech recognition. In addition, by performing a text prediction task on pre-training Chinese phonetic units to obtain a decoder, the decoder is able to construct text using phonetic features, thereby improving the decoder's language modeling ability. Because the encoder and decoder are pre-trained to have certain speech recognition capabilities, re-training the pre-trained encoder and decoder can improve training efficiency and training accuracy. Furthermore, the model input used in pre-training is pre-training speech data or pre-training Chinese phonetic units, both of which are similar in modality to the speech data input when the speech recognition model is applied. This improves recognition accuracy when the speech recognition model recognizes target speech data. Furthermore, by using a large amount of low-cost unlabeled speech data and unlabeled Chinese text, and only using a small amount of speech-text labeled data, a speech recognition model for the Chinese language can be trained with high accuracy, thereby reducing the need for labeled data, reducing personnel costs, and improving training efficiency.
[0239] The above is an exemplary solution for a data processing device of a voice recognition model applied to a cloud-side device according to this embodiment. Note that the technical solution for the data processing device of a voice recognition model applied to the cloud-side device is based on the same idea as the technical solution for the data processing method of a voice recognition model applied to the cloud-side device, and for any details not described in detail in the technical solution for the data processing device of a voice recognition model applied to the cloud-side device, reference can be made to the description of the technical solution for the data processing method of a voice recognition model applied to the cloud-side device.
[0240] Corresponding to the above embodiment of the speech recognition method, this specification further provides an embodiment of a speech recognition device, and Figure 15 shows a structural schematic diagram of a speech recognition device provided by one embodiment of this specification. As shown in Figure 15, the device includes: a second acquisition module 1502 configured to acquire speech data to be recognized; a second encoding module 1504 configured to use an encoder of the speech recognition model pre-trained by the cloud-side device using the data processing method for the speech recognition model to encode the speech data to be recognized and obtain speech features of the speech data to be recognized; and a second decoding module 1506 configured to input the speech features into a decoder of the speech recognition model to obtain target text corresponding to the speech data to be recognized.
[0241] In one or more embodiments herein, the device comprises: Obtain a check set including a plurality of audio check pairs each including a check audio data and a corresponding check Chinese text, and a plurality of Chinese pronunciation unit check pairs each including a check audio data and a corresponding check Chinese pronunciation unit; Using the encoder of the speech recognition model, Chinese pronunciation unit prediction is performed on the check speech data, and the speech features of the check speech data and the predicted Chinese pronunciation unit are obtained. inputting the speech features of the checking speech data into a decoder of a speech recognition model to obtain a predicted Chinese text corresponding to the checking speech data; The apparatus further includes a fine-tuning module configured to fine-tune the speech recognition model based on the predicted Chinese pronunciation unit, the check Chinese pronunciation unit, the predicted Chinese text, and the check Chinese text, and obtain a speech recognition model that has been fine-tuned when a fine-tuning stopping condition is reached.
[0242] In one or more embodiments herein, the device further comprises a display module, an input module, and an update module; the display module is configured to send the target text to the front end for display; the input module is configured to receive correction text corresponding to target text input by a user at the front end; The update module is configured to update the speech recognition model based on the modified text and the recognition target speech data to obtain an updated speech recognition model.
[0243] In the solution applied in the embodiments of this specification, the encoder and decoder are pre-trained before obtaining a speech recognition model. This reduces the number of sample speech data and sample Chinese texts required for training to obtain a speech recognition model, thereby reducing the burden on labelers and the difficulty of obtaining labeled data. Given the characteristic of Chinese data being an ideographic language, i.e., there is a large discrepancy between speech and text, and the same pronunciation can correspond to approximately 100 Chinese characters, the modality of pronunciation units is added to the model pre-training process. This is because pronunciation units serve as a bridge between speech and text, i.e., both speech and text can be uniquely mapped to the same pronunciation unit sequence. During pre-training, a speech mask prediction task and a pronunciation unit prediction task are performed on the pre-training speech data to obtain an encoder. Both tasks map speech data into a pronunciation unit sequence, allowing the encoder to capture pronunciation information in the speech signal and contribute to speech recognition. In addition, by performing a text prediction task on pre-training Chinese phonetic units to obtain a decoder, the decoder is able to construct text using phonetic features, thereby improving the decoder's language modeling ability. Because the encoder and decoder are pre-trained to have certain speech recognition capabilities, re-training the pre-trained encoder and decoder can improve training efficiency and training accuracy. Furthermore, the model input used in pre-training is pre-training speech data or pre-training Chinese phonetic units, both of which are similar in modality to the speech data input when the speech recognition model is applied. This improves recognition accuracy when the speech recognition model recognizes target speech data. Furthermore, by using a large amount of low-cost unlabeled speech data and unlabeled Chinese text, and only using a small amount of speech-text labeled data, a speech recognition model for the Chinese language can be trained with high accuracy, thereby reducing the need for labeled data, reducing personnel costs, and improving training efficiency.
[0244] The above is an exemplary solution of the speech recognition device according to this embodiment. Note that the technical solution of the speech recognition device is based on the same idea as the technical solution of the speech recognition method described above, and for details not described in detail in the technical solution of the speech recognition device, please refer to the description of the technical solution of the speech recognition method described above.
[0245] 16 shows a structural block diagram of a computing device 1600 provided in accordance with an embodiment of the present disclosure. Components of the computing device 1600 include, but are not limited to, a memory 1610 and a processor 1620. The processor 1620 is connected to the memory 1610 via a bus 1630, and a database 1650 is used to store data.
[0246] Computing device 1600 further includes access device 1640, which enables computing device 1600 to communicate over one or more networks 1660. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1640 may include one or more of any type of network interface (e.g., network interface card (NIC)), wired or wireless, including an IEEE 802.11 Wireless Local Area Network (WLAN) radio interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, etc.
[0247] In one embodiment of the present specification, the above components of the computing device 1600, as well as other components not shown in Figure 16, may be connected to each other, for example, via a bus. It should be understood that the structural block diagram of the computing device shown in Figure 16 is for illustrative purposes only and is not intended to limit the scope of the present specification. Those skilled in the art can add or substitute other components as needed.
[0248] Computing device 1600 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., tablet computer, personal digital assistant, laptop computer, notebook computer, netbook, etc.), a mobile phone (e.g., smartphone), a wearable computing device (e.g., smart watch, smart glasses, etc.), or other type of mobile device, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1600 may also be a mobile or stationary server.
[0249] The processor 1620 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the data processing method for the speech recognition model or the steps of the speech recognition method.
[0250] The above is an exemplary solution for a computing device according to this embodiment. The technical solution for the computing device is based on the same concept as the technical solution for the data processing method for a speech recognition model or the speech recognition method described above. For details not described in detail in the technical solution for the computing device, please refer to the description of the technical solution for the data processing method for a speech recognition model or the speech recognition method described above.
[0251] An embodiment of the present specification further provides a computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, implements steps of the data processing method for the speech recognition model or steps of the speech recognition method.
[0252] The above is an exemplary solution of a computer-readable storage medium according to this embodiment. The technical solution of the storage medium is based on the same concept as the technical solution of the data processing method for a speech recognition model or the speech recognition method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the data processing method for a speech recognition model or the speech recognition method described above.
[0253] An embodiment of the present specification further provides a computer program, which, when executed on a computer, causes the computer to perform steps of the data processing method for the speech recognition model or steps of the speech recognition method.
[0254] The above is an exemplary solution of a computer program according to this embodiment. The technical solution of this computer program is based on the same idea as the technical solution of the data processing method for a speech recognition model or the speech recognition method described above. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the data processing method for a speech recognition model or the speech recognition method described above.
[0255] The foregoing describes specific examples of the present specification. Other examples are within the scope of the following claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the examples and still achieve desirable results. Also, the processes depicted in the accompanying figures do not necessarily require the particular order or sequential order shown to achieve desirable results. In some embodiments, multitasking and parallel processing may also be possible or advantageous.
[0256] The computer instructions include computer program code, which may be in the form of source code, object code, an executable file, or any intermediate form, etc. The computer-readable medium may comprise any entity or device capable of carrying the computer program code, a recording medium, a USB memory, a removable hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier wave signal, an electrical communication signal, a software distribution medium, etc.
[0257] Although the above-described embodiments of the methods are all expressed as a combination of a series of operations for ease of explanation, those skilled in the art should recognize that the embodiments of the present specification are not limited by the order of the operations described, since some steps may be performed in other orders or simultaneously. Furthermore, those skilled in the art should recognize that the embodiments described herein are preferred embodiments, and that the related operations and modules are not necessarily required for the embodiments of the present specification.
[0258] In the above embodiments, the description of each embodiment has its own focus, and the parts of the embodiments that are not described in detail can refer to the relevant descriptions of other embodiments.
[0259] The preferred examples disclosed herein above are intended only to assist in the description of the present specification. The selected examples do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described above. Obviously, various modifications and variations can be made based on the content of the examples of the present specification. These examples have been chosen and specifically described in the present specification to better explain the principles and practical applications of the examples of the present specification, and to enable those skilled in the art to easily understand and utilize the present specification. The present specification is limited only by the claims, their full scope, and equivalents.
[0260] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to a Chinese patent application bearing application number 202211329674.7 and entitled "Speech recognition model data processing system and method, speech recognition method," filed with the State Intellectual Property Office of China on October 27, 2022, the entire contents of which are incorporated herein by reference.
Claims
1. A data processing system for a speech recognition model, comprising a cloud-side device and a terminal-side device, the cloud-side device is used for: acquiring a sample set including a plurality of sample pairs each including sample speech data and sample Chinese text; using an encoder for pre-training by executing a Chinese phonetic unit prediction task on pre-training speech data to encode the sample speech data and obtain speech features of the sample speech data; inputting the speech features into a decoder for pre-training by executing a text prediction task on the pre-training Chinese phonetic units to obtain predicted Chinese text; pre-training a model including the encoder and the decoder based on the predicted Chinese text and the sample Chinese text; and obtaining model parameters of the speech recognition model obtained by pre-training when a pre-training stopping condition is reached; the cloud-side device is further used to transmit model parameters of the speech recognition model obtained by the pre-training to the terminal-side device; The terminal device is a data processing system for a speech recognition model, which is used to perform speech recognition on speech data to be recognized using the speech recognition model and obtain target text corresponding to the speech data to be recognized.
2. The cloud-side device Obtaining a first pre-training speech set including a plurality of unsupervised first pre-training speech data; using an encoder to encode the first pre-training speech data, obtain first speech features corresponding to the first pre-training speech data, and determine a first phonetic unit based on the first speech features; masking the first pre-training speech data; using the encoder to encode the masked first pre-training speech data, obtain second speech features corresponding to the masked first pre-training speech data, and determine second phonetic units based on the second speech features; 2. The data processing system of claim 1, further adapted to: pre-training the encoder based on first phonetic units and second phonetic units corresponding to the first pre-training speech data.
3. Specifically, the cloud-side device is extracting spectral features of the first pre-training speech data; and inputting the spectral features of the first pre-training speech data into an encoder to obtain first speech features corresponding to the first pre-training speech data.
4. The cloud-side device Obtaining a plurality of first pre-training pairs including second pre-training speech data and first pre-training Chinese phonetic units; Using the encoder, perform Chinese pronunciation unit prediction on the second pre-training speech data to obtain predicted Chinese pronunciation units corresponding to the second pre-training speech data; 3. The data processing system of claim 2, further used for: pre-training the encoder based on the first pre-training Chinese phonetic unit and a predicted Chinese phonetic unit.
5. the encoder comprises a feature encoding layer, and the cloud-side device Obtaining a first pre-training text set including a plurality of unsupervised first pre-training Chinese texts; Converting the first pre-training Chinese text into second pre-training Chinese pronunciation units, and inputting the second pre-training Chinese pronunciation units into the feature encoding layer to obtain phonetic features of the second pre-training Chinese pronunciation units; inputting the phonetic features of the second pre-training Chinese phonetic unit into a decoder to obtain a predicted Chinese text corresponding to the second pre-training Chinese phonetic unit; and pre-training the decoder based on predicted Chinese text corresponding to the second pre-training Chinese phonetic unit and the first pre-training Chinese text.
6. The cloud-side device acquiring a second pre-training speech set including a plurality of third pre-training speech data carrying the target pseudo-label; using the encoder to encode the third pre-training speech data to obtain speech features of the third pre-training speech data; inputting speech features of the third pre-training speech data into the decoder to obtain predicted pseudo labels corresponding to the third pre-training speech data; and pre-training the decoder based on the target pseudo labels and predicted pseudo labels.
7. Specifically, the cloud-side device is acquiring a plurality of unsupervised third pre-training speech data; inputting the plurality of third pre-training speech data into a pre-trained speech encoder to obtain speech features of the plurality of third pre-training speech data; and clustering the speech features of the plurality of third pre-training speech data to obtain a target pseudo label for each third pre-training speech data.
8. the encoder comprises a feature encoding layer, and the cloud-side device Obtaining a plurality of second pre-training pairs including third pre-training Chinese phonetic units and second pre-training Chinese texts; inputting the third pre-training Chinese phonetic unit into the feature encoding layer to obtain phonetic features of the third pre-training Chinese phonetic unit; inputting the phonetic features of the third pre-training Chinese phonetic unit into the decoder to obtain a predicted Chinese text corresponding to the third pre-training Chinese phonetic unit; and pre-training the feature encoding layer and the decoder based on a predicted Chinese text corresponding to the third pre-training Chinese phonetic unit and the second pre-training Chinese text to obtain a model comprising an encoder and a decoder.
9. 1. A data processing method for a speech recognition model applied to a cloud-side device, the cloud-side device being connected to a plurality of terminal-side devices, the method comprising: obtaining a sample set including a plurality of sample pairs including sample speech data and sample Chinese text; using an encoder that is pre-trained by performing a Chinese phonetic unit prediction task on pre-training speech data to encode the sample speech data and obtain speech features of the sample speech data; inputting the speech features into a decoder that is pre-trained by performing a text prediction task on pre-training Chinese phonetic units to obtain predicted Chinese text; pre-training a model including the encoder and the decoder based on the predicted Chinese text and the sample Chinese text, and obtaining model parameters of the pre-trained speech recognition model when a pre-training stopping condition is reached; and transmitting model parameters of the speech recognition model obtained by pre-training to a first terminal-side device, which is one of the plurality of terminal-side devices.
10. A speech recognition method applied to a terminal-side device, the terminal-side device being connected to a cloud-side device, the method comprising: acquiring speech data to be recognized; a step of using an encoder of a speech recognition model pre-trained by the cloud-side device using the data processing method for a speech recognition model according to claim 9 to encode the speech data to be recognized and acquire speech features of the speech data to be recognized; inputting the speech features into a decoder of the speech recognition model to obtain target text corresponding to the speech data to be recognized.
11. The method comprises: Obtaining a check set including a plurality of speech check pairs including check speech data and corresponding check Chinese texts, and a plurality of Chinese pronunciation unit check pairs including check speech data and corresponding check Chinese pronunciation units; using an encoder of the speech recognition model to perform Chinese pronunciation unit prediction on the check speech data, and obtaining speech features of the check speech data and predicted Chinese pronunciation units; inputting the speech features of the checking speech data into a decoder of the speech recognition model to obtain a predicted Chinese text corresponding to the checking speech data; 11. The speech recognition method of claim 10, further comprising: fine-tuning the speech recognition model based on the predicted Chinese pronunciation unit, the check Chinese pronunciation unit, the predicted Chinese text, and the check Chinese text; and obtaining a speech recognition model that has been fine-tuned when a fine-tuning stopping condition is reached.
12. After inputting the speech features into a decoder of the speech recognition model to obtain target text corresponding to the speech data to be recognized, sending the target text to a front end for display; receiving a correction text corresponding to the target text entered by a user at the front end; The speech recognition method according to claim 10 , further comprising: updating the speech recognition model based on the corrected text and the recognition-target speech data to obtain an updated speech recognition model.
13. A computing device comprising a memory and a processor, A computing device, wherein the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, such that, when the computer-executable instructions are executed by the processor, steps of the speech recognition model data processing method of claim 9 or steps of the speech recognition method of any one of claims 10 to 12 are implemented.
14. A computer-readable storage medium having stored thereon computer-executable instructions, which, when executed by a processor, implement the steps of the data processing method for a speech recognition model according to claim 9 or the steps of the speech recognition method according to any one of claims 10 to 12.
Citation Information
Patent Citations
Speech recognition model training method and device, electronic equipment and storage medium
CN115116443A
Transcription device and program
JP2014134640A
End-to-End Streaming Keyword Spotting
JP2021524615A
Information processing device, information processing method, information processing program, terminal device, inference method, and inference program
JP2022049569A