Method, apparatus, recording medium, and electronic device for training an acoustic recognition model
The method addresses the data dependency and frame independence limitations of end-to-end ASR models by using a two-network training approach with contrastive and association loss functions, reducing labeling costs and improving recognition accuracy.
Patent Information
- Application Number
- JP2024573793
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-14
- Filing Date
- 2023-02-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-02-13
AI Technical Summary
End-to-end deep neural network-based automatic speech recognition models require large amounts of labeled data and are limited by the assumption of independent speech features between frames in the CTC framework, leading to suboptimal recognition performance.
A method involving constructing an initial acoustic recognition model with two networks, performing self-supervised training using a contrastive learning loss function on an unlabeled dataset, followed by association loss function training on a labeled dataset to adjust network parameters, and finally fine-tuning with a combined loss function to obtain a target model.
Reduces the reliance on labeled data, improves recognition accuracy by aligning with actual speech feature dependencies, and enhances the development and optimization of speech recognition models.
Smart Images

Figure 2025521290000001_ABST
Abstract
Description
Technical Field
[0001] Cross-reference to Related Applications This application claims priority based on a Chinese patent application filed with the China National Intellectual Property Administration on July 14, 2022, with an application number of 202210833610.4 and an invention title of "Method, Apparatus, Recording Medium, and Electronic Device for Training an Automatic Speech Recognition Model", and incorporates all the content of the Chinese patent application herein by reference.
[0002] The present invention relates to the field of automatic speech recognition, and in particular, to a method for training an automatic speech recognition model, an apparatus for training an automatic speech recognition model, a recording medium, and an electronic device.
Background Art
[0003] In recent years, with the rapid development of deep learning technology, automatic speech recognition (ASR) based on an end-to-end deep neural network has increasingly become the mainstream technology in the current field of automatic speech recognition.
[0004] Since the end-to-end ASR model has a large amount of parameters, the performance of the model often depends on a large amount of labeled data. Also, usually, the self-supervised ASR method is mainly performed in a CTC (Connectionist Temporal Classification) framework. However, in the CTC framework, it is assumed that the speech features are independent of each other between frames, which is different from the actual situation, so its performance is limited. It is necessary to further improve the recognition performance of the automatic speech recognition model under the condition of insufficient labeled data.
[0005] It should be noted that the information disclosed in the above background art is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those skilled in the art.
Summary of the Invention
Means for Solving the Problem
[0006] According to a first aspect of an embodiment of the present invention, there is provided a method for training an acoustic recognition model. The method for training the acoustic recognition model includes: constructing an initial acoustic recognition model including a first network having first initial parameters and a second network having second initial parameters; fixing the second initial parameters, calculating a contrastive learning loss function based on an unlabeled dataset, and performing self-supervised training on the first network based on the contrastive learning loss function to adjust the first initial parameters to first intermediate parameters; fixing the first intermediate parameters, calculating a first association loss function based on a labeled dataset, and training the second network based on the first association loss function to adjust the second initial parameters to second intermediate parameters; calculating a second association loss function based on the labeled dataset, and training the first network and the second network based on the second association loss function to adjust the first intermediate parameters and the second intermediate parameters to obtain a target acoustic recognition model.
[0007] According to a second aspect of an embodiment of the present invention, there is provided an apparatus for training an acoustic recognition model, the apparatus for training the acoustic recognition model including: a model construction module for constructing an initial acoustic recognition model including a first network having first initial parameters and a second network having second initial parameters; a first training module for fixing the second initial parameters, calculating a contrastive learning loss function based on an unlabeled dataset, and performing self-supervised training on the first network based on the contrastive learning loss function to adjust the first initial parameters to first intermediate parameters; a second training module for fixing the first intermediate parameters, calculating a first association loss function based on a labeled dataset, and training the second network based on the first association loss function to adjust the second initial parameters to second intermediate parameters; and a model adjustment module for calculating a second association loss function based on the labeled dataset, and training the first network and the second network based on the second association loss function to adjust the first intermediate parameters and the second intermediate parameters to obtain a target acoustic recognition model.
[0008] According to a third aspect of an embodiment of the present invention, there is provided a computer-readable recording medium storing a computer program, and when the program is executed by a processor, the method for training an acoustic recognition model in the above embodiment is realized.
[0009] According to a fourth aspect of an embodiment of the present invention, there is provided an electronic device including one or more processors and a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the method for training an acoustic recognition model in the above embodiment is realized by the one or more processors.
[0010] Note that the above general description and the detailed description below are merely illustrative and explanatory descriptions and do not limit the present invention.
[0011] Brief Description of the Drawings The drawings here are incorporated into the specification and constitute a part of the specification, exemplify embodiments suitable for the present application, and are for interpreting the principles of the present application together with the specification. Note that the drawings in the following description are merely some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings on the premise of not investing creative labor.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0013] Hereinafter, exemplary embodiments will be described more comprehensively with reference to the drawings. However, the exemplary embodiments can be implemented in multiple forms and are not limited to the examples described in this specification. On the contrary, these embodiments are provided to make the present application complete and comprehensive, and to convey the idea of the exemplary embodiments to those skilled in the art in an all-round way.
[0014] In addition, the features, configurations or characteristics described can be combined with one or more embodiments in any suitable manner. In the following description, many specific details are provided to enable a complete understanding of the embodiments according to the present invention. However, those skilled in the art should understand that the technical solution according to the present invention can be realized even without one or more of these specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring the aspects of the present application.
[0015] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be realized in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0016] The flowcharts shown in the drawings are only exemplary explanations and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, an operation / step may be decomposed, or an operation / step may be combined or partially combined, so the actual execution order may be changed according to the actual situation.
[0017] Hereinafter, the details of the realization of the technical solution of the embodiments of the present invention will be described in detail. FIG. 1 is a schematic diagram schematically showing the flow of a method for training an acoustic recognition model in an exemplary embodiment of the present invention. As shown in FIG. 1, this method for training an acoustic recognition model includes steps S101 to S104.
[0018] In step S101, an initial acoustic recognition model including a first network having first initial parameters and a second network having second initial parameters is constructed.
[0019] In step S102, the second initial parameters are fixed, a contrastive learning loss function is calculated based on an unlabeled dataset, and self-supervised training is performed on the first network based on the contrastive learning loss function, thereby adjusting the first initial parameters to first intermediate parameters.
[0020] In step S103, the first intermediate parameters are fixed, a first association loss function is calculated based on a labeled dataset, and the second network is trained based on the first association loss function, thereby adjusting the second initial parameters to second intermediate parameters.
[0021] In step S104, a second association loss function is calculated based on the labeled dataset, and the first network and the second network are trained based on the second association loss function, thereby adjusting the first intermediate parameters and the second intermediate parameters to obtain a target acoustic recognition model.
[0022] In the technical solutions provided by some embodiments of the present invention, first, based on an initial speech recognition model, a contrastive learning loss function is calculated using an unlabeled dataset to pre-train the first network of the model. Then, the parameters of the first network are fixed, and an association loss function is calculated using a labeled dataset to train the second network of the model. Finally, the parameters of the first network and the second network are fine-tuned by calculating the association loss function using labeled data to train the speech recognition model, and the model is trained until convergence to obtain the final speech recognition model. The training method of the speech recognition model of the present invention does not depend on a large amount of labeled data during training, so it reduces the labeling cost of data in automatic speech recognition (ASR), improves the progress of the development and optimization of the speech recognition model, and at the same time, the training process of the model is not restricted by the Connectionist Temporal Classification (CTC) framework, so it avoids indicating that speech features are independent of each other between frames, makes it more consistent with the actual situation, and further improves the recognition accuracy of the speech recognition model.
[0023] Hereinafter, each step of the training method of the speech recognition model in this exemplary embodiment will be described in more detail in combination with the drawings and embodiments.
[0024] In step S101, an initial speech recognition model including a first network having first initial parameters and a second network having second initial parameters is constructed.
[0025] In one embodiment of the present invention, first, a randomly initialized speech recognition model is constructed. The network structure of the speech recognition model may include an embedding layer (i.e., Embedding layer), a transformation layer (i.e., Transformer layer), and an output layer. Here, the Transformer layer is composed of a first network and a second network. The first network is an encoder network, and the second network is a decoder network.
[0026] For the initial speech recognition model after being randomly initialized, both the first network and the second network have their respective initial parameters, and in subsequent model training, the trained speech recognition model is obtained by adjusting the network model parameters.
[0027] In one embodiment of the present invention, before performing the training of steps S102 to S104, it is also necessary to prepare a training dataset. FIG. 2 is a schematic diagram schematically showing the flow of a method for preparing a training dataset in an exemplary embodiment of the present invention. As shown in FIG. 2, this method for preparing the training dataset includes the following steps.
[0028] In step S201, audio sample data is acquired based on a preset audio sampling rate, and the audio sample data is divided into a first audio sample and a second audio sample.
[0029] In step S202, a label-free dataset is obtained by calculating the audio feature matrix of the first audio sample.
[0030] In step S203, a labeled dataset is obtained based on the calculated audio feature matrix of the second audio sample and the text labeling result of the acquired second audio sample.
[0031] In step S201, audio is sampled according to a preset audio sampling rate to obtain audio sample data, and the sampled audio may be Chinese audio or audio in other languages. For example, audio samples for a certain period are obtained by sampling at an audio sampling rate of 16 kHz.
[0032] Afterwards, in order to arrange the unlabeled dataset and the labeled dataset, the sampled audio sample data can be split into two parts. A part is for generating the unlabeled dataset, with a total of i samples, and the other part is for generating the labeled dataset, with a total of j samples.
[0033] In addition, in the process of splitting, some audio samples can be used as the first audio sample or the second audio sample, that is, there may be overlapping parts in content.
[0034] In step S202, an unlabeled dataset is generated. Since the unlabeled dataset does not require audio labeling, the audio feature matrix of the first audio sample is directly calculated to obtain an unlabeled dataset denoted as U = {x i |i ∈ [1, N u}, where x i is the audio feature matrix of the i-th first audio sample, and Nu is the number of unlabeled first audio samples in the unlabeled dataset.
[0035] In step S203, a labeled dataset is generated. Since each audio sample in the labeled dataset has a corresponding text labeling result, the audio feature matrix of the second audio sample is calculated, and the second audio sample is labeled to obtain a text labeling result, so as to obtain a labeled dataset denoted as L = {x j , y j |j ∈ [1, N l}, where x j is the audio feature matrix of the j-th second audio sample, y j is the text labeling result corresponding to the audio feature matrix x j , and N l is the number of unlabeled second audio samples in the unlabeled dataset.
[0036]
Number
[0037] In steps S202 and S203, when calculating the audio feature matrix of the audio sample, the audio feature matrix may be the features of a 80-dimensional Mel spectrogram. Here, the time length of each frame of the spectrogram is 25 ms, and the step size is 10 ms.
[0038] In step S102, the second initial parameter is fixed, a contrastive learning loss function is calculated based on the unlabeled dataset, and self-supervised training is performed on the first network based on the contrastive learning loss function, thereby adjusting the first initial parameter to a first intermediate parameter.
[0039] In an embodiment of the present invention, in step S102, self-supervised training is performed on the first network, and the first network includes a convolutional neural network module and a convolutional enhancement module.
[0040] Here, the first network may be an encoder network, and includes a CNN (Convolutional Neural Network) module which is a convolutional neural network module and a Conformer module which is a convolutional enhancement module. For example, the encoder network is a sequential connection of a 5-layer CNN module and 12 Conformer modules.
[0041] FIG. 3 is a schematic diagram schematically showing the flow of the method for calculating the contrastive learning loss function in an exemplary embodiment of the present invention. As shown in FIG. 3, this method for calculating the contrastive learning loss function includes steps S301 to S304.
[0042] In step S301, based on the convolutional neural network module, a shallow representation result of audio sample data in the unlabeled dataset is calculated.
[0043] In step S302, mask processing is performed on the shallow representation result to obtain a mask representation result, and based on the convolutional enhancement module, a deep representation result of the mask representation result is calculated.
[0044] In step S303, the shallow representation result is linearly transformed to obtain a target representation result.
[0045] In step S304, based on the deep representation result and the target representation result, the contrastive learning loss function is calculated.
[0046] Next, steps S301 to S304 will be described in detail. In step S301, based on the convolutional neural network module, a shallow representation result of audio sample data in the unlabeled dataset is calculated.
[0047] Specifically, audio sample data xi ∈ U in the unlabeled dataset is given, and multilayer CNN calculation is performed on x i to obtain a shallow representation result denoted as e.
[0048] Then, two types of processing, namely the processing of steps S302 and S301, are respectively performed on the shallow representation result e, and the results of such processing are further compared.
[0049] In step S302, mask processing is performed on the shallow representation result to obtain a mask representation result, and based on the convolutional enhancement module, a deep representation result of the mask representation result is calculated.
[0050] Specifically, FIG. 4 is a schematic diagram schematically showing the flow of the mask processing method in an exemplary embodiment of the present invention. As shown in FIG. 4, this mask processing method includes the following steps.
[0051] In step S401, a seed sample frame is randomly selected from the shallow layer representation result based on a random mask probability.
[0052] In step S402, the feature vectors of the K consecutive frames after the seed sample frame in the shallow layer representation result are replaced with learnable vectors to obtain the mask representation result, where K is a positive integer.
[0053]
Number
[0054] Here, p is a random mask probability, which is a preset value. For example, when p = 6.5, K is a mask parameter for consecutive frames, and K is also a preset value and a positive integer. For example, K = 10. Of course, the embodiments of the present invention are merely illustrative explanations, and the values of the random mask probability and the mask parameter for consecutive frames can be adaptively adjusted according to actual needs.
[0055]
Number
[0056] In step S303, the shallow layer representation result is linearly transformed to obtain a target representation result.
[0057] Specifically, the linear transformation is a linear mapping, which is a mapping from a vector space V to another vector space W and maintains addition and multiplication by a number. The shallow layer representation result e is linearly transformed to obtain a target representation result denoted as q.
[0058] In step S304, the contrastive learning loss function is calculated based on the deep representation result and the target representation result.
[0059] FIG. 5 is a schematic diagram schematically showing a flow of another method for calculating a contrastive learning loss function in an exemplary embodiment of the present invention. As shown in FIG. 5, this method for calculating the contrastive learning loss function includes the following steps.
[0060] In step S501, M-frame anchor samples are selected from the masked portion in the deep representation result as the first sample, where M is a positive integer.
[0061] In step S502, M-frame anchor samples that correspond one-to-one to the M-frame anchor samples in the first sample are selected from the target representation result as the second sample, and S-frame negative samples are selected as the third sample, where S is a positive integer.
[0062] In step S503, the contrastive learning loss function is calculated based on the similarity between the first sample and the second sample and the similarity between the first sample and the third sample.
[0063] Specifically, M-frame anchor samples are selected from the masked portion in the deep representation result h, and each frame sample, i.e., the first sample, is denoted as h m where M is the number of frames of the anchor sample, which is a preset value and a positive integer. For example, let the number of frames of the anchor sample M = 10.
[0064]
Equation
[0065] And, as shown in Equation (1), x i The contrastive learning loss function loss of the audio sample i is calculated.
[0066] [Number]
[0067] [Number]
[0068] Specifically, sim() is a similarity function, and the calculation formula is as shown in Formula (2).
[0069] [Number]
[0070] [Number]
[0071] x i For each audio sample, for the contrastive learning loss function loss i if it can be calculated, for the total contrastive learning loss function loss of all unlabeled dataset U, it is necessary to integrate the loss functions of each audio sample, for example, by calculating the average value.
[0072] Based on the above method, a contrastive learning task is designed, and self-supervised training is performed on the first network encoder network in the speech recognition model using the unlabeled dataset U. After training is completed, the first initial parameters of the encoder network are adjusted to the first intermediate parameters. Since it does not depend on a large amount of labeled data, it can reduce the labeling cost of data in automatic speech recognition (ASR) and improve the progress of the development and optimization of the speech recognition model.
[0073] In step S103, the first intermediate parameter is fixed, a first association loss function is calculated based on the labeled dataset, and the second network is trained based on the first association loss function, thereby adjusting the second initial parameter to a second intermediate parameter.
[0074] In one embodiment of the present invention, step S103 trains a second network including a feature transformation module.
[0075] Here, the second network may be a decoder network and includes one or more feature transformation modules, i.e., transform modules. For example, the decoder network is composed of six transform modules.
[0076] After step S102, the training of the encoder network has been completed, but the decoder is still in a randomly initialized state. In order to avoid the unbalanced training state of the decoder and the encoder, in this step, the decoder network part is trained by the association loss function to achieve the purpose of initially training the decoder network.
[0077] In one embodiment of the present invention, the decoder network is trained by the association loss function, and the association loss function is a CTC-attention (attention) association loss function.
[0078] Specifically, the loss functions used in the training process of the current end-to-end ASR model mainly include: (1) a loss function based on Connectionist Temporal Classification (CTC); (2) an encoder-decoder loss function based on an attention mechanism; and (3) a CTC-attention combined loss function. Here, since the CTC-attention combined loss function combines the advantages of CTC and the attention mechanism, the present invention performs model training using the CTC-attention combined loss function.
[0079] When training the model, using a labeled dataset L, the encoder network is fixed, that is, the first intermediate parameters are fixed, and the model training for the decoder network is completed using the CTC-attention combined loss function until the decoder network converges. Furthermore, the decoder network is adjusted from the second initial parameters to the second intermediate parameters.
[0080] In step S104, a second combined loss function is calculated based on the labeled dataset, and the first network and the second network are trained based on the second combined loss function to adjust the first intermediate parameters and the second intermediate parameters to obtain a target speech recognition model.
[0081] In an embodiment of the present invention, step S104 finely adjusts the parameters of the two networks in the speech recognition model. The loss function still uses the CTC-attention combined loss function.
[0082] Specifically, by using the labeled dataset L, opening the encoder network and the decoder network, and optimizing the CTC-attention joint loss function, fine-tuning training is performed on the encoder network and the decoder network until the model converges, thereby adjusting the first intermediate parameter and the second intermediate parameter to obtain a final speech recognition model.
[0083] According to the method for training a speech recognition model provided in the embodiment, since the training process of the model is not restricted by the Connectionist Temporal Classification (CTC) framework, it is avoided to show that speech features are independent of each other between frames, making it more consistent with the actual situation, and further improving the recognition accuracy of the speech recognition model.
[0084] FIG. 6 is a schematic diagram schematically showing the configuration of a training apparatus for a speech recognition model according to an exemplary embodiment of the present invention. As shown in FIG. 6, the training apparatus 600 for the speech recognition model may include a model construction module 601, a first training module 602, a second training module 603, and a model adjustment module 604.
[0085] Here, the model construction module 601 is configured to construct an initial speech recognition model including a first network having a first initial parameter and a second network having a second initial parameter.
[0086] The first training module 602 is configured to fix the second initial parameter, calculate a contrastive learning loss function based on an unlabeled dataset, and perform self-supervised training on the first network based on the contrastive learning loss function, thereby adjusting the first initial parameter to a first intermediate parameter.
[0087] The second training module 603 is configured to adjust the second initial parameters to second intermediate parameters by fixing the first intermediate parameters, calculating a first association loss function based on the labeled dataset, and training the second network based on the first association loss function.
[0088] The model adjustment module 607 is configured to obtain a target speech recognition model by calculating a second association loss function based on the labeled dataset and training the first network and the second network based on the second association loss function to adjust the first intermediate parameters and the second intermediate parameters.
[0089] According to an exemplary embodiment of the present invention, the first network includes a convolutional neural network module and a convolutional enhancement module.
[0090] According to an exemplary embodiment of the present invention, the first training module 602 includes a shallow unit, a mask unit, a target unit, and a comparison unit. The shallow unit is configured to calculate a shallow representation result of audio sample data in the unlabeled dataset based on the convolutional neural network module. The mask unit is configured to perform a mask process on the shallow representation result to obtain a mask representation result and calculate a deep representation result of the mask representation result based on the convolutional enhancement module. The target unit is configured to linearly transform the shallow representation result to obtain a target representation result. The comparison unit is configured to calculate the contrastive learning loss function based on the deep representation result and the target representation result.
[0091] According to an exemplary embodiment of the present invention, the mask unit randomly selects from the shallow representation result based on a random mask probability to obtain a seed sample frame, and further configured to replace the feature vectors of K consecutive frames after the seed sample frame in the shallow representation result with learnable vectors to obtain the mask representation result, where K is a positive integer.
[0092] According to an exemplary embodiment of the present invention, the comparison unit selects M-frame anchor samples from the mask part in the deep representation result as the first sample, selects M-frame anchor samples corresponding one-to-one to the M-frame anchor samples in the first sample from the target representation result as the second sample, selects S-frame negative samples as the third sample, and further configured to calculate the contrastive learning loss function based on the similarity between the first sample and the second sample and the similarity between the first sample and the third sample, where M is a positive integer and S is a positive integer.
[0093] According to an exemplary embodiment of the present invention, the second network includes a feature deformation module.
[0094] According to an exemplary embodiment of the present invention, the training device 600 of the speech recognition model further includes a data preparation module, which acquires audio sample data based on a preset audio sampling rate, divides the audio sample data into a first audio sample and a second audio sample, calculates the audio feature matrix of the first audio sample to obtain the unlabeled dataset, and configured to obtain the labeled dataset based on the calculated audio feature matrix of the second audio sample and the text labeling result of the acquired second audio sample.
[0095] Specific details of each module in the above-described voice recognition model training apparatus 600 are described in detail in the corresponding voice recognition model training method, and thus the description thereof is omitted here.
[0096] In addition, in the above detailed description, some modules and units of the devices for executing operations have been described, but such a classification is not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the above-described two or more modules and units may be embodied in one module and unit. Conversely, the features and functions of the above-described one module and unit may be further embodied by a plurality of modules and units.
[0097] In an exemplary embodiment of the present invention, a recording medium capable of realizing the above method is further provided. FIG. 7 is a schematic diagram schematically showing a computer-readable recording medium in an exemplary embodiment of the present invention. As shown in FIG. 7, a program product 700 for implementing the above method according to an embodiment of the present invention is shown. This program product can use a portable compact disc read-only memory (CD-ROM), includes program code, and can be executed on a terminal device such as a mobile phone. However, the program product of the present invention is not limited thereto. In this specification, the readable recording medium can be any tangible medium that includes or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0098] In an exemplary embodiment of the present invention, an electronic device capable of realizing the above method is further provided. FIG. 8 is a schematic diagram schematically showing the structure of a computer system of an electronic device in an exemplary embodiment of the present invention.
[0099] Note that the computer system 800 of the electronic device shown in FIG. 8 is merely an example and does not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0100] As shown in FIG. 8, computer system 800 includes a Central Processing Unit (CPU) 801 and can execute various appropriate operations and processes based on a program stored in a Read-Only Memory (ROM) 802 or a program loaded from a storage unit 808 into a Random Access Memory (RAM) 803. Various programs and data necessary for the operation of the system are further stored in RAM 803. The CPU 801, ROM 802, and RAM 803 are interconnected via a bus 804. An Input / Output (I / O) interface 805 is also connected to the bus 804.
[0101] Connected to the I / O interface 805 are an input unit 806 including a keyboard, a mouse, etc., an output unit 807 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc., and a speaker, etc., a storage unit 808 including a hard disk, etc., and a communication unit 809 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication unit 809 performs communication processing via a network such as the Internet. A driver 810 is also connected to the I / O interface 805 as needed. A removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed in the driver 810 as needed, and the computer program read therefrom is installed in the storage unit 808 as needed.
[0102] In particular, according to an embodiment of the present invention, the process described below with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product including a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication unit 809 and / or installed from the removable medium 811. When the computer program is executed by the central processing unit (CPU) 801, various functions limited to the system of the present invention are executed.
[0103] Note that the computer-readable medium according to the embodiments of the present invention may be a computer-readable signal medium, a computer-readable recording medium, or any combination thereof. The computer-readable recording medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable recording medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable recording medium is any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable recording medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code included in the computer-readable medium may be transmitted via any suitable medium, including, but not limited to, wireless, wire, or any suitable combination of the above.
[0104] Flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of systems, methods, and computer program products that can be implemented in various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, a program segment, or a portion of code, and the above-mentioned module, program segment, or portion of code includes executable instructions for implementing one or more specified logical functions. Also, it should be noted that in some alternative embodiments, the functions described in the blocks may be executed in an order different from the order shown in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or depending on the related functions, may be executed in the reverse order. In addition, it should be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing a specified function or operation, or by a combination of dedicated hardware and computer instructions.
[0105] The units according to the embodiments of the present invention may be implemented in software, may be implemented in hardware, and the described units may be provided in a processor. Here, in some cases, the names of these units do not limit the unit itself.
[0106] According to another aspect, the present invention further provides a computer-readable medium. The computer-readable medium may be included in the electronic device according to the above embodiments, or may exist alone but not incorporated into the electronic device. One or more programs are arranged in the above computer-readable medium, and when the above one or more programs are executed by the electronic device, the electronic device realizes the method described in the above embodiments.
[0107] Note that in the above detailed description, some modules and units of the device for executing operations have been described, but such classification is not mandatory. In fact, according to the embodiments of the present invention, the features and functions of two or more of the above-described modules and units may be embodied in one module and unit. Conversely, the features and functions of one module and unit described above may be further embodied by a plurality of modules and units.
[0108] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein may be implemented in hardware or in a manner combining software with the necessary hardware. For this reason, although the technical solution according to the embodiments of the present invention may be represented in the form of a software product, the software product may be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash memory, a mobile hard disk, etc.) or on a network, and includes several commands for causing a computer facility (which may be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method described in the embodiments of the present invention.
[0109] Those skilled in the art will easily conceive of other embodiments of the present invention by considering this specification and implementing the content disclosed herein. The present invention includes any modifications, uses, or adaptive changes to the present invention, and such modifications, uses, or adaptive changes follow the general principles of the present invention and include known techniques or ordinary technical means in the technical field not disclosed in the present invention.
[0110] The present invention is not limited to the specific configurations described above and illustrated in the drawings, and various modifications and changes may be made without departing from its scope. The scope of the present invention is limited only by the appended claims.
Claims
1. Constructing an initial speech recognition model including a first network with first initial parameters and a second network with second initial parameters; fixing the second initial parameters, calculating a contrastive learning loss function based on an unlabeled dataset, and performing self-supervised training on the first network based on the contrastive learning loss function to adjust the first initial parameters to first intermediate parameters; fixing the first intermediate parameters, calculating a first association loss function based on a labeled dataset, and training the second network based on the first association loss function to adjust the second initial parameters to second intermediate parameters; calculating a second association loss function based on the labeled dataset, and training the first network and the second network based on the second association loss function to adjust the first intermediate parameters and the second intermediate parameters to obtain a target speech recognition model, including A method for training a speech recognition model.
2. The first network includes a convolutional neural network module and a convolutional enhancement module The method for training a speech recognition model according to claim 1.
3. The step of calculating a contrastive learning loss function based on an unlabeled dataset includes Calculating a shallow representation result of audio sample data in the unlabeled dataset based on the convolutional neural network module; Performing a masking process on the shallow representation result to obtain a masked representation result, and calculating a deep representation result of the masked representation result based on the convolutional enhancement module; Linearly transforming the shallow representation result to obtain a target representation result; Calculating the contrastive learning loss function based on the deep representation result and the target representation result, including The method for training a speech recognition model according to claim 2.
4. The step of performing a masking process on the shallow representation result to obtain a masked representation result includes Randomly selecting from the shallow representation result based on a random masking probability to obtain a seed sample frame; Replacing the feature vectors of the K consecutive frames after the seed sample frame in the shallow layer representation result with learnable vectors to obtain the mask representation result, and K is a positive integer The method for training an acoustic recognition model according to claim 3
5. The step of calculating the contrastive learning loss function based on the deep layer representation result and the target representation result is Selecting M frame anchor samples from the mask part in the deep layer representation result as the first sample, Selecting, as the second sample, M frame anchor samples that correspond one-to-one to the M frame anchor samples in the first sample from the target representation result, and selecting S frame negative samples as the third sample, Calculating the contrastive learning loss function based on the similarity between the first sample and the second sample and the similarity between the first sample and the third sample, and including M is a positive integer, S is a positive integer The method for training an acoustic recognition model according to claim 3
6. The second network includes a feature transformation module The method for training an acoustic recognition model according to claim 1
7. Obtaining audio sample data based on a preset audio sampling rate and dividing the audio sample data into a first audio sample and a second audio sample, Obtaining the label-free dataset by calculating the audio feature matrix of the first audio sample, Further including obtaining the labeled dataset based on the calculated audio feature matrix of the second audio sample and the text labeling result of the obtained second audio sample The method for training an acoustic recognition model according to claim 1
8. A model construction module for constructing an initial acoustic recognition model including a first network having first initial parameters and a second network having second initial parameters Fix the second initial parameter, calculate a contrastive learning loss function based on an unlabeled dataset, and perform self-supervised training on the first network based on the contrastive learning loss function to adjust the first initial parameter to a first intermediate parameter, a first training module; Fix the first intermediate parameter, calculate a first association loss function based on a labeled dataset, and train the second network based on the first association loss function to adjust the second initial parameter to a second intermediate parameter, a second training module; Calculate a second association loss function based on the labeled dataset, and train the first network and the second network based on the second association loss function to adjust the first intermediate parameter and the second intermediate parameter to obtain a target speech recognition model, a model adjustment module, comprising A training device for a speech recognition model.
9. A computer-readable recording medium storing a computer program, When the program is executed by a processor, the training method of the speech recognition model according to any one of claims 1 to 7 is realized A computer-readable recording medium.
10. One or more processors, A storage device for storing one or more programs, comprising When the one or more programs are executed by the one or more processors, the training method of the speech recognition model according to any one of claims 1 to 7 is realized by the one or more processors An electronic device.
Citation Information
Patent Citations
Voice recognition model training method and device, electronic equipment and storage medium
CN111916067A
Model training method and device and electronic equipment
CN112509563A
Pronunciation bias error detection method and device and storage medium
CN113327595A
Model training method and system, terminal equipment and storage medium
CN113744727A
Heterogeneous language model training method and device, equipment and storage medium
CN114416955A