Speech recognition model training method, device, recording medium, and electronic device
The method trains a speech recognition model using self-supervised and association loss functions to reduce data labeling costs and enhance accuracy, addressing the limitations of existing models by aligning with real-world speech feature dependencies.
Patent Information
- Application Number
- JP2024573793
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-07-14
- Filing Date
- 2023-02-13
- Publication Date
- 2025-10-09
- Estimated Expiration
- 2043-02-13
AI Technical Summary
Existing speech recognition models, particularly those based on end-to-end deep neural networks, rely heavily on large amounts of labeled data and are limited by the Connectionist Temporal Classification (CTC) framework, which assumes independent speech features between frames, leading to suboptimal performance when labeled data is scarce.
A method involving self-supervised training of a speech recognition model using an initial model with two networks, where one network is pre-trained with an unlabeled dataset through a control learning loss function, and the other is trained with a labeled dataset using an association loss function, followed by fine-tuning with a joint loss function to achieve a target model.
This approach reduces the reliance on labeled data, enhances model performance, and improves recognition accuracy by aligning with actual speech feature dependencies, overcoming the limitations of the CTC framework.
Smart Images

Figure 0007752264000009 
Figure 0007752264000010 
Figure 0007752264000011
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority based on a Chinese patent application filed with the China Patent Office on July 14, 2022, bearing application number 202210833610.4 and entitled "Method, apparatus, recording medium and electronic device for training a voice recognition model," and the entire contents of the Chinese patent application are incorporated herein by reference.
[0002] The present invention relates to the field of speech recognition, and more particularly to a method for training a speech recognition model, a training device for a speech recognition model, a recording medium, and an electronic device. [Background technology]
[0003] In recent years, with the rapid development of deep learning technology, automatic speech recognition (ASR) based on end-to-end deep neural networks has increasingly become the mainstream technology in the current speech recognition field.
[0004] Because the parameter amount of end-to-end ASR models is large, model performance often depends on a large amount of labeled data. Furthermore, self-supervised ASR methods are typically implemented within the Connectionist Temporal Classification (CTC) framework, but the CTC framework assumes that speech features are independent of each other between frames, which is different from the real situation, limiting its performance. It is necessary to further improve the recognition performance of speech recognition models under conditions where labeled data is scarce.
[0005] It should be noted that the information disclosed in the above background art is intended only to enhance understanding of the background of the present invention, and may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention [Means for solving the problem]
[0006] According to a first aspect of an embodiment of the present invention, there is provided a method for training a speech recognition model, the method including: constructing an initial speech recognition model including a first network having first initial parameters and a second network having second initial parameters; adjusting the first initial parameters to first intermediate parameters by fixing the second initial parameters, calculating a control learning loss function based on an unlabeled dataset, and performing self-supervised training on the first network based on the control learning loss function; adjusting the second initial parameters to second intermediate parameters by fixing the first intermediate parameters, calculating a first association loss function based on a labeled dataset, and training the second network based on the first association loss function; and adjusting the first intermediate parameters and the second intermediate parameters to obtain a target speech recognition model by calculating a second association loss function based on the labeled dataset and training the first network and the second network based on the second association loss function.
[0007] According to a second aspect of an embodiment of the present invention, there is provided a speech recognition model training apparatus, the speech recognition model training apparatus including: a model construction module for constructing an initial speech recognition model including a first network having first initial parameters and a second network having second initial parameters; a first training module for adjusting the first initial parameters to first intermediate parameters by fixing the second initial parameters, calculating a control learning loss function based on an unlabeled dataset, and performing self-supervised training on the first network based on the control learning loss function; a second training module for adjusting the second initial parameters to second intermediate parameters by fixing the first intermediate parameters, calculating a first association loss function based on a labeled dataset, and training the second network based on the first association loss function; and a model adjustment module for adjusting the first intermediate parameters and the second intermediate parameters to obtain a target speech recognition model by calculating a second association loss function based on the labeled dataset and training the first network and the second network based on the second association loss function.
[0008] According to a third aspect of an embodiment of the present invention, there is provided a computer-readable recording medium having a computer program stored thereon, which, when executed by a processor, realizes the method for training a speech recognition model in the above embodiment.
[0009] According to a fourth aspect of an embodiment of the present invention, there is provided an electronic device, the electronic device including one or more processors and a storage device for storing one or more programs, the one or more programs, when executed by the one or more processors, causing the one or more processors to realize the speech recognition model training method in the above embodiment.
[0010] It should be noted that the above general description and the following detailed description are merely exemplary and explanatory and are not intended to limit the present invention.
[0011] BRIEF DESCRIPTION OF THE DRAWINGS The drawings herein are incorporated into the specification and constitute a part of this specification, illustrate embodiments suitable for the present application, and, together with the specification, serve to explain the principles of the present application. Note that the drawings in the following description are merely some embodiments of the present application, and those skilled in the art can obtain other drawings from these drawings without any creative effort. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 2 is a schematic diagram illustrating a flow of a method for training a speech recognition model in an exemplary embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram illustrating the flow of a method for preparing a training dataset in an exemplary embodiment of the present invention. [Figure 3] FIG. 1 is a schematic diagram illustrating a flow of a method for calculating a contrastive learning loss function in an exemplary embodiment of the present invention. [Figure 4] 1 is a schematic diagram showing a flow of a mask processing method in an exemplary embodiment of the present invention; [Figure 5] FIG. 10 is a schematic diagram illustrating the flow of another method for calculating a contrastive learning loss function in an exemplary embodiment of the present invention. [Figure 6] 1 is a schematic diagram illustrating the configuration of a training device for a speech recognition model in an exemplary embodiment of the present invention; [Figure 7] 1 is a schematic diagram illustrating a computer-readable recording medium according to an exemplary embodiment of the present invention; [Figure 8] FIG. 1 is a schematic diagram illustrating the structure of a computer system of an electronic device in an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0013]
[0023] Exemplary embodiments will be described more fully below with reference to the accompanying drawings. However, the exemplary embodiments may be embodied in many different forms and are not limited to the examples set forth herein. Rather, these embodiments are provided so as to fully complete this application and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0014] It should be noted that the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to provide a thorough understanding of the embodiments of the present invention. However, it should be understood by those skilled in the art that the technical solutions of the present invention may be realized without one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other instances, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0015] The block diagrams shown in the drawings are merely functional entities that do not necessarily correspond to physically separate entities, i.e., these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller devices.
[0016] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily have to be performed in the order described. For example, some operations / steps may be separated, while others may be merged or partially merged, so the order of actual execution may be changed according to actual circumstances.
[0017] The implementation details of the technical solutions of the embodiments of the present invention are described in detail below. 1 is a schematic diagram showing a flow of a method for training a speech recognition model in an exemplary embodiment of the present invention. As shown in FIG. 1, the method for training a speech recognition model includes steps S101 to S104.
[0018] In step S101, an initial speech recognition model is constructed, which includes a first network having a first initial parameter and a second network having a second initial parameter.
[0019] In step S102, the second initial parameters are fixed, a control learning loss function is calculated based on an unlabeled dataset, and the first initial parameters are adjusted to first intermediate parameters by performing self-supervised training on the first network based on the control learning loss function.
[0020] In step S103, the first intermediate parameters are fixed, a first joint loss function is calculated based on a labeled dataset, and the second network is trained based on the first joint loss function, thereby adjusting the second initial parameters to second intermediate parameters.
[0021] In step S104, a second joint loss function is calculated based on the labeled dataset, and the first network and the second network are trained based on the second joint loss function, thereby adjusting the first intermediate parameters and the second intermediate parameters to obtain a target speech recognition model.
[0022] In some embodiments of the present invention, a technical solution is provided in which, based on an initial speech recognition model, a first network of the model is first pre-trained by calculating a contrastive learning loss function using an unlabeled data set, the parameters of the first network are fixed, and a second network of the model is trained by calculating an association loss function using a labeled data set, and finally, the parameters of the first network and the second network are fine-tuned by using the labeled data to calculate the association loss function and train the speech recognition model until convergence is achieved to obtain a final speech recognition model. Since the speech recognition model training method of the present invention does not rely on a large amount of labeled data during training, it reduces the cost of labeling data in automatic speech recognition (ASR) and improves the progress of speech recognition model development and optimization. Meanwhile, since the model training process is not restricted by the connectionist time series classification (CTC) framework, it avoids showing that speech features are independent from each other between frames, is more consistent with actual situations, and further improves the recognition accuracy of the speech recognition model.
[0023] Each step of the method for training a speech recognition model in this exemplary embodiment will now be described in more detail with reference to the accompanying drawings and examples.
[0024] In step S101, an initial speech recognition model is constructed, which includes a first network having a first initial parameter and a second network having a second initial parameter.
[0025] In one embodiment of the present invention, a randomly initialized speech recognition model is first constructed. The network structure of the speech recognition model may include an embedding layer, a transformer layer, and an output layer. Here, the transformer layer is composed of a first network and a second network, where the first network is an encoder network and the second network is a decoder network.
[0026] For the initial speech recognition model after being randomly initialized, the first network and the second network both have their own initial parameters, and the trained speech recognition model is obtained by adjusting the network model parameters in subsequent model training.
[0027] In one embodiment of the present invention, before performing the training in steps S102 to S104, it is also necessary to prepare a training dataset. Figure 2 is a schematic diagram showing a flow of a method for preparing a training dataset in an exemplary embodiment of the present invention. As shown in Figure 2, this method for preparing a training dataset includes the following steps:
[0028] In step S201, audio sample data is obtained based on a preset audio sampling rate, and the audio sample data is divided into a first audio sample and a second audio sample.
[0029] In step S202, the unlabeled dataset is obtained by calculating an audio feature matrix of the first audio sample.
[0030] In step S203, the labeled dataset is obtained based on the calculated audio feature matrix of the second audio sample and the obtained text labeling result of the second audio sample.
[0031] In step S201, audio is sampled according to a preset audio sampling rate to obtain audio sample data, the sampled audio may be Chinese audio or audio of other languages, for example, sampled at an audio sampling rate of 16 kHz to obtain audio samples for a certain period.
[0032] Then, to arrange the unlabeled and labeled datasets, the sampled audio sample data can be split into two parts: one part is used to generate an unlabeled dataset, a total of i pieces, and the other part is used to generate a labeled dataset, a total of j pieces.
[0033] It should be noted that in the process of division, some audio samples may be the first audio samples and some may be the second audio samples, that is, there may be overlapping parts in the content.
[0034] In step S202, an unlabeled dataset is generated. Since the unlabeled dataset does not require labeling of the audio, the audio feature matrix of the first audio sample is calculated as U={x i |i∈[1,N u ]}, where x i is the audio feature matrix of the i-th first audio sample, and Nu is the number of unlabeled first audio samples in the unlabeled dataset.
[0035] In step S203, a labeled dataset is generated. Each audio sample in the labeled dataset has a corresponding text labeling result, so that an audio feature matrix of the second audio sample is calculated and the second audio sample is labeled to obtain the text labeling result, where L={x j ,y j |j∈[1,N l ]}, where x j is the audio feature matrix of the j-th second audio sample, and y j is the audio feature matrix x j is the text labeling result corresponding to N l is the number of unlabeled second audio samples in the unlabeled dataset.
[0036]
number
[0037] In steps S202 and S203, when calculating the audio feature matrix of the audio sample, the audio feature matrix may be an 80-dimensional mel-spectrogram feature, where the time length of each frame of the spectrogram is 25 ms and the step size is 10 ms.
[0038] In step S102, the second initial parameters are fixed, a control learning loss function is calculated based on an unlabeled dataset, and the first initial parameters are adjusted to first intermediate parameters by performing self-supervised training on the first network based on the control learning loss function.
[0039] In one embodiment of the present invention, in step S102, a first network undergoes self-supervised training, where the first network includes a convolutional neural network module and a convolutional reinforcement module.
[0040] Here, the first network may be an encoder network, including a convolutional neural network (CNN) module and a conformer module, which is a convolutional neural network module. For example, the encoder network may include a 5-layer CNN module and 12 conformer modules connected in sequence.
[0041] 3 is a schematic diagram showing a flow of a method for calculating a contrast learning loss function in an exemplary embodiment of the present invention. As shown in FIG. 3, the method for calculating a contrast learning loss function includes steps S301 to S304.
[0042] In step S301, a shallow representation result of audio sample data in the unlabeled dataset is calculated based on the convolutional neural network module.
[0043] In step S302, a masking process is performed on the shallow representation result to obtain a masked representation result, and a deep representation result is calculated for the masked representation result based on the convolution enhancement module.
[0044] In step S303, the shallow representation result is linearly transformed to obtain the target representation result.
[0045] In step S304, the contrastive learning loss function is calculated based on the deep representation result and the target representation result.
[0046] Next, steps S301 to S304 will be described in detail. In step S301, a shallow representation result of audio sample data in the unlabeled dataset is calculated based on the convolutional neural network module.
[0047] Specifically, given audio sample data xi ∈ U in the unlabeled dataset, x i We perform multi-layer CNN calculations on the vectors to obtain a shallow representation result denoted as e.
[0048] Then, the shallow representation result e is subjected to two processes, that is, the two processes in step S302 and step S301, and the results of these processes are compared.
[0049] In step S302, a masking process is performed on the shallow representation result to obtain a masked representation result, and a deep representation result is calculated for the masked representation result based on the convolution enhancement module.
[0050] Specifically, Fig. 4 is a schematic diagram showing the flow of a mask processing method in an exemplary embodiment of the present invention. As shown in Fig. 4, the mask processing method includes the following steps:
[0051] In step S401, a seed sample frame is randomly selected from the shallow representation result based on a random mask probability.
[0052] In step S402, the feature vectors of K consecutive frames after the seed sample frame in the shallow representation result are replaced with trainable vectors to obtain the mask representation result, where K is a positive integer.
[0053]
number
[0054] Here, p is a random mask probability and is a preset value, for example, p=6.5, and K is a mask parameter of successive frames, and K is also a preset value and is a positive integer, for example, K=10. Of course, the embodiments of the present invention are merely illustrative, and the values of the random mask probability and the mask parameter of successive frames can be adaptively adjusted according to practical needs.
[0055]
number
[0056] In step S303, the shallow representation result is linearly transformed to obtain the target representation result.
[0057] Specifically, a linear transformation is a linear mapping from one vector space V to another vector space W, which preserves addition and multiplication. The shallow representation result e is linearly transformed to obtain a target representation result denoted as q.
[0058] In step S304, the contrastive learning loss function is calculated based on the deep representation result and the target representation result.
[0059] 5 is a schematic diagram illustrating the flow of another method for calculating a contrast learning loss function in an exemplary embodiment of the present invention. As shown in FIG. 5, the method for calculating a contrast learning loss function includes the following steps:
[0060] In step S501, an anchor sample of M frames is selected as a first sample from the mask portion of the deep representation result, where M is a positive integer.
[0061] In step S502, an anchor sample of an M frame that corresponds one-to-one with the anchor sample of the M frame in the first sample is selected as the second sample from the target representation result, and a negative sample of an S frame is selected as the third sample, where S is a positive integer.
[0062] In step S503, the contrastive learning loss function is calculated based on the similarity between the first sample and the second sample and the similarity between the first sample and the third sample.
[0063] Specifically, M anchor samples are selected from the masked portion of the deep representation result h, and each frame sample, i.e., the first sample, is calculated as h m where M is the number of frames of anchor samples, is a preset value, and is a positive integer. For example, the number of frames of anchor samples is M=10.
[0064]
number
[0065] Then, as shown in equation (1), x i Contrastive learning loss function for audio samples i Calculate.
[0066]
number
[0067]
number
[0068] Specifically, sim() is a similarity function, and the calculation formula is as shown in formula (2).
[0069]
number
[0070]
number
[0071] x i A contrastive learning loss function loss for each audio sample i If we can calculate the loss function for each audio sample, we need to integrate the loss function for each audio sample, for example by taking the average, over the total control learning loss function loss for all unlabeled datasets U.
[0072] Based on the above method, a contrastive learning task is designed, and self-supervised training is performed on the first encoder network in the speech recognition model using the unlabeled dataset U. After the training is completed, the first initial parameters of the encoder network are adjusted to the first intermediate parameters. Since this does not rely on a large amount of labeled data, it can reduce the cost of labeling data in automatic speech recognition (ASR) and accelerate the progress of speech recognition model development and optimization.
[0073] In step S103, the first intermediate parameters are fixed, a first joint loss function is calculated based on a labeled dataset, and the second network is trained based on the first joint loss function, thereby adjusting the second initial parameters to second intermediate parameters.
[0074] In one embodiment of the present invention, step S103 trains a second network including a feature transformation module.
[0075] Here, the second network may be a decoder network, which includes one or more feature transformation modules, i.e., transform modules, for example, the decoder network is composed of six transform modules.
[0076] After step S102, the training of the encoder network has been completed, but the decoder is still in a randomly initialized state. In order to avoid the training states of the decoder and the encoder becoming unbalanced, in this step, the decoder network part is trained using a joint loss function, so as to achieve the purpose of initially training the decoder network.
[0077] In one embodiment of the present invention, the decoder network is trained with an association loss function, which is a CTC-attention association loss function.
[0078] Specifically, the loss functions used in the training process of current end-to-end ASR models mainly include (1) a loss function based on Connectionist Temporal Classification (CTC), (2) an encoder-decoder loss function based on an attention mechanism, and (3) a CTC-attention joint loss function. Here, the CTC-attention joint loss function combines the advantages of both the CTC and attention mechanisms, so the present invention uses the CTC-attention joint loss function for model training.
[0079] When training the model, the labeled dataset L is used to fix the encoder network, i.e., fix the first intermediate parameters, and then complete model training for the decoder network using the CTC-attention joint loss function until the decoder network converges, and then adjust the decoder network from the second initial parameters to the second intermediate parameters.
[0080] In step S104, a second joint loss function is calculated based on the labeled dataset, and the first network and the second network are trained based on the second joint loss function, thereby adjusting the first intermediate parameters and the second intermediate parameters to obtain a target speech recognition model.
[0081] In one embodiment of the present invention, step S104 fine-tunes the parameters of the two networks in the speech recognition model, and the loss function still uses the CTC-attention joint loss function.
[0082] Specifically, using the labeled dataset L, the encoder network and the decoder network are opened, and the CTC-attention joint loss function is optimized to perform fine-tuning training on the encoder network and the decoder network until the model converges, thereby adjusting the first and second intermediate parameters to obtain the final speech recognition model.
[0083] According to the speech recognition model training method provided in the embodiment, the model training process is not restricted by the connectionist temporal classification (CTC) framework, which avoids showing that speech features are independent of each other between frames, which is more consistent with actual situations, and further improves the recognition accuracy of the speech recognition model.
[0084] FIG. 6 is a schematic diagram illustrating the configuration of a training device for a speech recognition model in an exemplary embodiment of the present invention. As shown in FIG. 6, the training device 600 for a speech recognition model may include a model construction module 601, a first training module 602, a second training module 603, and a model adjustment module 604.
[0085] Here, the model construction module 601 is configured to construct an initial speech recognition model including a first network having first initial parameters and a second network having second initial parameters.
[0086] The first training module 602 is configured to fix the second initial parameters, calculate a control learning loss function based on an unlabeled dataset, and perform self-supervised training on the first network based on the control learning loss function, thereby adjusting the first initial parameters to first intermediate parameters.
[0087] A second training module 603 is configured to adjust the second initial parameters to second intermediate parameters by fixing the first intermediate parameters, calculating a first joint loss function based on a labeled dataset, and training the second network based on the first joint loss function.
[0088] The model adjustment module 607 is configured to calculate a second joint loss function based on the labeled dataset, and to adjust the first intermediate parameters and the second intermediate parameters to obtain a target speech recognition model by training the first network and the second network based on the second joint loss function.
[0089] According to an exemplary embodiment of the present invention, the first network includes a convolutional neural network module and a convolutional reinforcement module.
[0090] According to an exemplary embodiment of the present invention, the first training module 602 includes a shallow layer unit, a mask unit, a target unit, and a comparison unit, where the shallow layer unit is configured to calculate shallow representations of audio sample data in the unlabeled dataset based on the convolutional neural network module. The mask unit is configured to perform a masking process on the shallow representations to obtain masked representations, and to calculate deep representations of the masked representations based on the convolutional enhancement module. The target unit is configured to linearly transform the shallow representations to obtain target representations. The comparison unit is configured to calculate the contrastive learning loss function based on the deep representations and the target representations.
[0091] According to an exemplary embodiment of the present invention, the mask unit is further configured to: randomly select from the shallow representation result based on a random mask probability to obtain a seed sample frame; and replace feature vectors of K consecutive frames after the seed sample frame in the shallow representation result with trainable vectors to obtain the mask representation result, where K is a positive integer.
[0092] According to an exemplary embodiment of the present invention, the comparison unit is further configured to: select an anchor sample of an M frame from a mask portion in the deep representation result as a first sample; select an anchor sample of an M frame from the target representation result that has a one-to-one correspondence with the anchor sample of the M frame in the first sample as a second sample; select a negative sample of an S frame as a third sample; and calculate the contrast learning loss function based on the similarity between the first sample and the second sample and the similarity between the first sample and the third sample, where M is a positive integer and S is a positive integer.
[0093] According to an exemplary embodiment of the present invention, the second network includes a feature transformation module.
[0094] According to an exemplary embodiment of the present invention, the speech recognition model training apparatus 600 further includes a data preparation module, configured to obtain audio sample data based on a preset audio sampling rate, divide the audio sample data into first audio samples and second audio samples, and calculate an audio feature matrix of the first audio samples to obtain the unlabeled dataset, and obtain the labeled dataset based on the calculated audio feature matrix of the second audio samples and the obtained text labeling result of the second audio samples.
[0095] The specific details of each module in the above-mentioned speech recognition model training device 600 are explained in detail in the corresponding speech recognition model training method, and therefore will not be explained here.
[0096] Although the above detailed description describes several modules and units of the device for performing operations, such division is not mandatory. In fact, according to an embodiment of the present invention, the features and functions of two or more of the modules and units described above may be embodied in one module and unit. Conversely, the features and functions of one of the modules and units described above may be further embodied in multiple modules and units.
[0097] In an exemplary embodiment of the present invention, a computer-readable recording medium capable of implementing the above method is further provided. FIG. 7 is a schematic diagram illustrating a computer-readable recording medium in an exemplary embodiment of the present invention. As shown in FIG. 7, a program product 700 for implementing the above method according to an embodiment of the present invention is shown. This program product can be stored on a portable compact disc read-only memory (CD-ROM), contain program code, and can be executed on a terminal device such as a mobile phone. However, the program product of the present invention is not limited thereto. In this specification, a computer-readable recording medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0098] In an exemplary embodiment of the present invention, an electronic device capable of implementing the above method is further provided. Figure 8 is a schematic diagram showing the structure of a computer system of the electronic device in an exemplary embodiment of the present invention.
[0099] It should be noted that the computer system 800 of the electronic device shown in FIG. 8 is merely an example and does not impose any limitations on the functions and scope of use of the embodiments of the present invention.
[0100] 8, a computer system 800 includes a central processing unit (CPU) 801, and can execute various appropriate operations and processes based on programs stored in a read-only memory (ROM) 802 or programs loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data necessary for the operation of the system. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0101] The I / O interface 805 is connected to an input unit 806 including a keyboard, a mouse, etc., an output unit 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, a storage unit 808 including a hard disk, etc., and a communication unit 809 including, for example, a network interface card such as a LAN (Local Area Network) card or a modem. The communication unit 809 performs communication processing via a network such as the Internet. A driver 810 is also connected to the I / O interface 805 as needed. A removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory is installed in the driver 810 as needed, and a computer program read from the medium is installed in the storage unit 808 as needed.
[0102] In particular, according to an embodiment of the present invention, the processes described below with reference to the flowcharts may be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product including a computer program carried on a computer-readable medium, the computer program including program code for executing the methods illustrated in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network via the communication unit 809 and / or installed from a removable medium 811. When the computer program is executed by the central processing unit (CPU) 801, it performs various functions specific to the system of the present invention.
[0103] Note that the computer-readable medium according to the embodiments of the present invention may be a computer-readable signal medium, a computer-readable recording medium, or any combination thereof. The computer-readable recording medium may be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable recording medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable recording medium may be any tangible medium that contains or stores a program, and the program may be used by or in connection with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, in which computer-readable program code is carried. Such propagated data signals may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable recording medium that may transmit, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium may be transmitted over any suitable medium, including, but not limited to, radio waves, wires, etc., or any suitable combination thereof.
[0104] The flowcharts and block diagrams in the figures illustrate possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, program segment, or portion of code, which includes executable instructions for implementing one or more predetermined logical functions. It should also be noted that in some alternative embodiments, the functions depicted in the blocks may be executed in a different order than depicted in the figures. For example, two blocks shown in succession may actually be executed substantially in parallel, or may be executed in the reverse order depending on the functionality involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented in a dedicated hardware-based system for performing a predetermined function or operation, or in a combination of dedicated hardware and computer instructions.
[0105] The units according to the embodiments of the present invention may be implemented in software or hardware, and the described units may be provided in a processor, where the names of the units do not necessarily limit the units themselves.
[0106] According to another aspect, the present invention further provides a computer-readable medium. The computer-readable medium may be included in the electronic device according to the above embodiment, or may exist independently but not integrated into the electronic device. One or more programs are stored on the computer-readable medium, and when the one or more programs are executed by the electronic device, the electronic device implements the method according to the above embodiment.
[0107] Although the above detailed description describes several modules and units of the device for performing operations, such division is not mandatory. In fact, according to an embodiment of the present invention, the features and functions of two or more of the modules and units described above may be embodied in one module and unit. Conversely, the features and functions of one of the modules and units described above may be further embodied in multiple modules and units.
[0108] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be realized by hardware or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be expressed in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash memory, a mobile hard disk, etc.) or a network, and includes commands for causing a computer device (such as a personal computer, a server, a touch terminal, or a network device) to execute the method according to the embodiments of the present invention.
[0109] Other embodiments of the present invention will be readily apparent to those skilled in the art from consideration of this specification and practice of the teachings disclosed herein. The present invention includes any modification, use, or adaptation to the present invention, which modification, use, or adaptation follows the general principles of the present invention and includes techniques known or ordinary in the art that are not disclosed herein.
[0110] The present invention is not limited to the specific constructions described above and illustrated in the drawings, and various modifications and variations may be made without departing from the scope thereof, which is limited only by the appended claims.
Claims
1. constructing an initial speech recognition model including a first network having first initial parameters and a second network having second initial parameters; adjusting the first initial parameters to first intermediate parameters by fixing the second initial parameters, calculating a control learning loss function based on an unlabeled dataset, and performing self-supervised training on the first network based on the control learning loss function; adjusting the second initial parameters to second intermediate parameters by fixing the first intermediate parameters, calculating a first joint loss function based on a labeled dataset, and training the second network based on the first joint loss function; calculating a second joint loss function based on the labeled dataset, and training the first network and the second network based on the second joint loss function to adjust the first intermediate parameters and the second intermediate parameters to obtain a target speech recognition model. How to train a speech recognition model.
2. The first network includes a convolutional neural network module and a convolutional reinforcement module. The method of training a speech recognition model according to claim 1 .
3. The step of calculating a control learning loss function based on an unlabeled dataset includes: calculating a shallow representation of audio sample data in the unlabeled dataset based on the convolutional neural network module; performing a masking process on the shallow representation result to obtain a masked representation result, and calculating a deep representation result of the masked representation result based on the convolution enhancement module; linearly transforming the shallow representation result to obtain a target representation result; and calculating the contrastive learning loss function based on the deep representation result and the target representation result. The method of training a speech recognition model according to claim 2.
4. The step of performing a mask process on the shallow representation result to obtain a mask representation result includes: randomly selecting a seed sample frame from the shallow representation result based on a random mask probability; and replacing feature vectors of K consecutive frames after the seed sample frame in the shallow representation result with trainable vectors to obtain the mask representation result; K is a positive integer The method of training a speech recognition model according to claim 3.
5. The step of calculating the contrastive learning loss function based on the deep representation result and the target representation result includes: Selecting an anchor sample of M frames as a first sample from a mask portion in the deep representation result; Selecting an anchor sample of an M frame that corresponds one-to-one with the anchor sample of the M frame in the first sample from the target representation result as a second sample, and selecting a negative sample of an S frame as a third sample; calculating the contrast learning loss function based on a similarity between the first sample and the second sample and a similarity between the first sample and the third sample; Including, M is a positive integer, S is a positive integer The method of training a speech recognition model according to claim 3.
6. The second network includes a feature transformation module. The method of training a speech recognition model according to claim 1 .
7. obtaining audio sample data based on a preset audio sampling rate, and dividing the audio sample data into first audio samples and second audio samples; obtaining the unlabeled dataset by calculating an audio feature matrix of the first audio samples; and obtaining the labeled dataset based on the calculated audio feature matrix of the second audio sample and the obtained text labeling result of the second audio sample. The method of training a speech recognition model according to claim 1 .
8. a model construction module for constructing an initial speech recognition model including a first network having first initial parameters and a second network having second initial parameters; a first training module for adjusting the first initial parameters to first intermediate parameters by fixing the second initial parameters, calculating a control learning loss function based on an unlabeled dataset, and performing self-supervised training on the first network based on the control learning loss function; and a second training module for adjusting the second initial parameters to second intermediate parameters by fixing the first intermediate parameters, calculating a first joint loss function based on a labeled dataset, and training the second network based on the first joint loss function; and a model adjustment module for calculating a second joint loss function based on the labeled dataset and for adjusting the first intermediate parameters and the second intermediate parameters by training the first network and the second network based on the second joint loss function to obtain a target speech recognition model. A training device for speech recognition models.
9. A computer-readable recording medium storing a computer program, When the program is executed by a processor, it implements the method for training a speech recognition model according to any one of claims 1 to 7. A computer-readable recording medium.
10. one or more processors; a storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more programs cause the one or more processors to implement the speech recognition model training method according to any one of claims 1 to 7. electronic equipment.
Citation Information
Patent Citations
Voice recognition model training method and device, electronic equipment and storage medium
CN111916067A
Model training method and device and electronic equipment
CN112509563A
Pronunciation bias error detection method and device and storage medium
CN113327595A
Model training method and system, terminal equipment and storage medium
CN113744727A
Heterogeneous language model training method and device, equipment and storage medium
CN114416955A