Training Method, Device, Equipment and Computer Readable Storage Medium for Speech Model
By using pseudo-labels from unlabeled samples in conjunction with labeled samples for joint training, the method addresses the challenge of limited annotated data in ASR systems, enhancing recognition accuracy and reducing annotation costs.
Patent Information
- Application Number
- CN202210067196.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-01-20
AI Technical Summary
Existing voice recognition technologies rely on a large amount of labeled data, resulting in limited recognition accuracy and it is difficult to maintain high accuracy in the absence of labeled data.
By obtaining the initial speech model, the initial model is trained using a small number of first speech samples carrying tags, and the second speech sample without tags is recognized to generate a pseudo-label, combined with the first speech sample carrying tags is trained, and the pseudo-label and tag data are used for semi-supervised learning and comparative learning are optimized for model parameters.
While reducing the annotation cost, the recognition accuracy of the speech model is improved, and the adaptability and recognition performance of the model in different scenarios is enhanced.
Smart Images

Figure CN114399995B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to speech recognition technology, and in particular, to a method, apparatus, device, computer program product, and computer-readable storage medium for training a speech model. Background Art
[0002] Automatic Speech Recognition (ASR) technology is a technology that converts a speech signal into corresponding text information. This technology can provide multiple applications such as automatic customer service, automatic speech translation, command control, and voice verification codes.
[0003] In the process of speech recognition processing, for the same speech data, multiple speech recognition results are usually parsed, and it is necessary to select the speech recognition result that best matches the speech data among the multiple speech recognition results. A reasonable and accurate selection determines the accuracy of the speech recognition processing.
[0004] The training scheme of the speech recognition system provided by the related technology has to rely on a large amount of labeled data in order to improve the recognition accuracy, which forms a contradiction with the reality that a large amount of pre-labeled data is difficult to obtain, affecting the accuracy of speech recognition. Summary of the Invention
[0005] Embodiments of the present application provide a method, apparatus, device, computer program product, and computer-readable storage medium for training a speech model, which can improve the recognition accuracy of the speech model.
[0006] The technical solution of the embodiments of the present application is implemented as follows:
[0007] Embodiments of the present application provide a method for training a speech model, including:
[0008] Obtain an initial speech model for speech recognition, where the initial speech model is trained based on a first speech sample carrying sample labels;
[0009] Perform speech recognition on a second speech sample through the initial speech model to obtain a recognition result, and use the recognition result as the pseudo-label of the second speech sample;
[0010] Jointly train the initial speech model based on the second speech sample carrying the pseudo-label and the first speech sample carrying the sample label to obtain a target speech model.
[0011] Embodiments of the present application provide a training apparatus for a speech model, including:
[0012] An acquisition module, configured to acquire an initial speech model for speech recognition, where the initial speech model is trained based on a first speech sample carrying a sample label;
[0013] A recognition module, configured to perform speech recognition on a second speech sample through the initial speech model to obtain a recognition result, and use the recognition result as a pseudo-label of the second speech sample;
[0014] A training module, configured to jointly train the initial speech model based on the second speech sample carrying the pseudo-label and the first speech sample carrying the sample label to obtain a target speech model.
[0015] In the above solution, the training module is further configured to perform prediction on a training sample in a jointly trained sample set through the initial speech model to obtain a prediction result; where the jointly trained sample set includes: the first speech sample carrying the sample label and the second speech sample carrying the pseudo-label;
[0016] Obtain the difference between the prediction result and the label of the training sample, and determine the value of the target loss function based on the difference;
[0017] Perform contrastive learning based on the training sample and the prediction result to determine the value of the contrastive loss function corresponding to the training sample;
[0018] Combine the value of the target loss function and the value of the contrastive loss function to update the model parameters of the initial speech model to obtain a target speech model.
[0019] In the above solution, the training module is further configured to construct corresponding positive samples and negative samples based on the training sample;
[0020] Determine the value of the contrastive loss function as the first loss between the positive sample and the training sample based on the positive sample and the prediction result;
[0021] Determine the value of the contrastive loss function as the second loss between the negative sample and the training sample based on the negative sample and the prediction result;
[0022] Determine the value of the contrastive loss function corresponding to the training sample based on the first loss and the second loss.
[0023] In the above solution, the training module is further configured to extract features from the training sample to obtain an initial speech feature corresponding to the training sample;
[0024] Perform discretization processing on the initial speech feature to obtain a target speech feature as the positive sample corresponding to the training sample;
[0025] Generate a noise feature corresponding to the initial voice feature as a negative sample corresponding to the training sample.
[0026] In the above solution, the training module is further configured to obtain a feature exchange ratio;
[0027] Based on the feature exchange ratio, exchange some features in the prediction result with some features in the training sample to obtain an exchanged prediction result and an exchanged training sample;
[0028] Among them, the exchanged prediction result is used to determine the value of the target loss function in combination with the label of the training sample;
[0029] The exchanged training sample is used to perform contrastive learning in combination with the prediction result to determine the value of the contrastive loss function corresponding to the training sample.
[0030] In the above solution, the training module is further configured to respectively obtain the weight of the target loss function and the weight of the contrastive loss function;
[0031] Based on the weight of the target loss function and the weight of the contrastive loss function, perform weighted summation on the value of the target loss function and the value of the contrastive loss function to obtain a weighted summation result;
[0032] Based on the weighted summation result, update the model parameters of the initial voice model to obtain a target voice model.
[0033] In the above solution, the training module is further configured to obtain a third voice sample carrying a voice label;
[0034] Based on the third voice sample, train the target voice model to update the model parameters of the target voice model.
[0035] In the above solution, the trained target voice model is further configured to obtain voice data to be recognized;
[0036] Input the voice data to be recognized into the target voice model and output the recognized voice result.
[0037] An embodiment of the present application provides an electronic device, including:
[0038] A memory for storing executable instructions;
[0039] A processor, when executing the executable instructions stored in the memory, implements the voice model training method provided by the embodiment of the present application.
[0040] An embodiment of the present application provides a computer-readable storage medium storing executable instructions that, when executed by a processor, implement the method for training a speech model provided by the embodiment of the present application.
[0041] The embodiment of the present application has the following beneficial effects:
[0042] Applying the above embodiment of the present application, an initial speech model is trained based on the first speech samples carrying labels, and then the initial speech model is used to identify the second speech samples without labels. The obtained recognition results are used as the pseudo-labels of the second speech samples, and the initial speech model is jointly trained based on the first speech samples and the second speech samples. In this way, the training of the speech model can be realized through a small number of labeled speech and unlabeled speech samples, reducing the annotation cost for speech samples and improving the recognition accuracy of the speech model. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1A-1B is a schematic structural diagram of a training system 100 for a speech model provided by an embodiment of the present application;
[0044] Figure 2 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0045] Figure 3 is a schematic flowchart of a method for training a speech model provided by an embodiment of the present application;
[0046] Figure 4 is a schematic diagram of a speech sample provided by an embodiment of the present application;
[0047] Figure 5 is a flowchart of training a speech model provided by an embodiment of the present application;
[0048] Figure 6 is a schematic flowchart of a method for jointly training a speech model provided by an embodiment of the present application;
[0049] Figure 7 is an example diagram of continuous time series classification provided by an embodiment of the present application;
[0050] Figure 8 is a schematic flowchart for determining a contrast loss function value provided by an embodiment of the present application;
[0051] Figure 9 is a schematic flowchart for constructing positive and negative samples provided by an embodiment of the present application;
[0052] Figure 10 is a flowchart of a method for jointly training a speech model provided by an embodiment of the present application;
[0053] Figure 11It is a flowchart of a fine-tuning method for a speech model provided by an embodiment of the present application;
[0054] Figure 12 It is a flowchart of a joint training method for a speech model provided by an embodiment of the present application;
[0055] Figure 13 It is a schematic diagram of an unsupervised training process of a speech model provided by the related art;
[0056] Figure 14 It is a flowchart of a contrastive learning for a speech model provided by the related art;
[0057] Figure 15 It is a schematic diagram of a pre-training process of a speech model provided by an embodiment of the present application;
[0058] Figure 16 It is a schematic diagram of a joint training of a speech model provided by an embodiment of the present application. Detailed implementation manners
[0059] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0060] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0061] If similar descriptions such as "first / second" appear in the application documents, the following explanation is added. In the following description, the terms "first \ second \ third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first \ second \ third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0063] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0064] 1) The acoustic model (AM) of the speech model, which represents the differentiated knowledge of acoustics, phonetics, environmental variables, speaker gender, and accent, including acoustic models based on the hidden Markov model (HMM), such as the mixed Gaussian-hidden Markov model (GMM-HMM) and deep neural network-hidden Markov model (DNN-HMM). The hidden Markov model is a weighted finite state automaton in the discrete time domain; of course, it can also include end-to-end acoustic models, such as the connectionist temporal classification-long short-term memory (CTC-LSTM) model and the attention model.
[0065] Each state of the acoustic model represents the probability distribution of the speech features of a speech unit (such as a word, syllable, or phoneme) in that state, and is connected into an ordered sequence of states through transitions between states, thus obtaining a sequence of speech units represented by a speech signal.
[0066] 2) The language model (LM) of the speech model is a simple representation of the language structure. The language structure here may include the rules between words and sentences, such as grammar, common word collocations, etc. The language model may include N-gram Model, Recurrent Neural Network (RNN), etc.
[0067] For a text sequence, the task of the language model is to calculate the probability distribution of the sequence, which can be generally explained as determining whether a language sequence is a normal sentence.
[0068] 3) The pronunciation dictionary records the correspondence between words and phonemes and is the hub connecting the acoustic model and the language model.
[0069] 4) Word Error Rate (WER) or Character Error Rate (CER), which describes the degree of match between the recognized word sequence and the true word sequence in the speech recognition task, and is an evaluation indicator of the speech recognition system; specifically: in order to make the recognized word sequence consistent with the standard word sequence, some words need to be replaced, deleted or inserted. The total number of these inserted, replaced or deleted words is divided by the percentage of the total number of words in the standard word sequence. Usually, English speech recognition is described by WER, and Chinese speech recognition is described by CER.
[0070] 5) Unsupervised learning: Unsupervised learning is a method of machine learning. Without given pre-labeled training examples, it automatically classifies or clusters the input data. The main applications of unsupervised learning include: cluster analysis, association rule, and dimensionality reduction.
[0071] 6) Contrastive Learning: Contrastive learning is a commonly used unsupervised training algorithm. This method constructs positive and negative samples and contrasts them in the feature space to learn the latent feature representation of the model. This method aims to learn the latent speech representation by maximizing the mutual information between positive and negative samples through contrastive learning.
[0072] The embodiments of the present application provide a method, device, equipment, and computer-readable storage medium for training a speech model, which can realize the training of the speech model based on a small amount of labeled data and a large amount of unlabeled data, and at the same time can reduce the acquisition cost of labeled training data while improving the speech recognition accuracy.
[0073] Based on the above explanations of the nouns and terms involved in the embodiments of the present application, the speech model training system provided by the embodiments of the present application will be described below. Refer to Figure 1A , Figure 1A FIG. is a schematic architecture diagram of the speech model training system 100 provided by the embodiments of the present application. To support an exemplary application, the terminals (exemplarily shown as terminals 400-1 and 400-2) are connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, and uses wireless or wired links to implement data transmission.
[0074] The terminals (such as terminals 400-1 and 400-2) are used to receive a trigger operation for performing speech recognition based on the speech recognition client (such as clients 410-1 and 410-2), and send a speech recognition request carrying speech data to the server 200;
[0075] The server 200 is used to receive the speech recognition request sent by the terminal, and in response to the acquisition request, return the speech recognition result for the speech data to be recognized to the terminal through the trained speech model;
[0076] A server 200 is configured to obtain an initial speech model for speech recognition. The initial speech model is trained based on first speech samples carrying sample labels. The second speech samples are subjected to speech recognition through the initial speech model to obtain recognition results, and the recognition results are used as pseudo-labels for the second speech samples. The initial speech model is jointly trained based on the second speech samples carrying pseudo-labels and the first speech samples carrying sample labels to obtain a target speech model.
[0077] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminals (such as terminals 400-1 and 400-2) may include, but are not limited to, smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart TVs, smart watches, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, etc., but are not limited thereto. The terminals (such as terminals 400-1 and 400-2) and the server 200 may be directly or indirectly connected through wired or wireless communication means, which is not limited in this application.
[0078] In some embodiments, the terminals (including terminals 400-1 and 400-2) are installed with and run a speech recognition client. The terminals (including terminals 400-1 and 400-2) send a speech recognition request carrying speech data to the server 200 based on the speech recognition client. After receiving the speech recognition request, the server 200 returns corresponding speech recognition results to the terminals in response to the recognition request. The terminals display or play the speech recognition results.
[0079] In some embodiments, the server 200 may be a server cluster or distributed system composed of multiple servers. Taking the distributed system as a blockchain system as an example, multiple servers can form a blockchain network, and the server 200 is a node on the blockchain network.
[0080] Next, an exemplary application of the blockchain network will be described by taking multiple servers accessing the blockchain network to implement the training of the speech model as an example.
[0081] In some embodiments, refer to Figure 1B , Figure 1BSchematic diagram of the architecture of the training system 100 for the speech model provided by the embodiments of the present application. Multiple servers involved in the speech model participate in the training of the speech model, such as the terminal 600 and the terminal 700. After obtaining the authorization of the blockchain management platform 900, the client 610 of the terminal 600 and the client 710 of the terminal 700 can both access the blockchain network 800.
[0082] The terminal 600 sends a speech recognition request to the blockchain management platform 900 (the terminal 700 sends a speech model acquisition request to the blockchain management platform 900). The blockchain management platform 900 generates a corresponding update operation according to the speech model acquisition request. The update operation specifies the smart contract to be called to implement the update operation / query operation and the parameters to be passed to the smart contract. The transaction also carries the digital signature signed by the web page and sends the update operation to the blockchain network 800.
[0083] When the nodes 210-1, 210-2, and 210-3 in the blockchain network 800 receive the update operation, they verify the digital signature of the update operation. After the digital signature verification is successful, according to the identity of the client 610 carried in the update operation, it is confirmed whether the client 610 has the acquisition permission. Any verification judgment in the digital signature and permission verification will result in acquisition failure. After the verification is successful, the signing node 210 signs its own digital signature (for example, encrypts the digest of the transaction using the private key of the node 210-1) and continues to broadcast in the blockchain network 800.
[0084] The nodes 210-1, 210-2, 210-3, etc. with sorting functions in the blockchain network 800, after receiving the successfully verified acquisition, fill the acquisition request into a new block and broadcast it to the nodes providing consensus services in the blockchain network 800.
[0085] The nodes in the blockchain network 800 that provide consensus services perform a consensus process on the new block to reach an agreement. The nodes that provide the ledger function append the new block to the tail of the blockchain and execute the fetch requests in the new block: for the submitted voice model requests, update the key-value pairs corresponding to the voice models in the status database; for the fetch requests of the voice models, query the key-value pairs corresponding to the voice models from the status database and send the corresponding voice models to the terminal. After the terminals 600 and 700 receive the voice models returned by the blockchain network 800, the terminals 600 and 700 train the voice models to obtain the trained voice models and display a prompt message indicating successful training in the graphical interfaces 610-1 and 710-1. The terminals 600 and 700 send the trained voice models to the blockchain network 800, and the blockchain network 800 performs voice recognition processing on the to-be-recognized voice data by invoking the trained voice models to obtain the voice results of the to-be-recognized voice.
[0086] The embodiments of the present application can also be implemented with the aid of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing.
[0087] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources.
[0088] Next, the electronic device for implementing the above-mentioned voice model training method provided by the embodiments of the present application will be described. Refer to Figure 2 , Figure 2 which is a schematic structural diagram of the electronic device provided by the embodiments of the present application. In actual applications, the electronic device 500 can be implemented as the server in FIG. 1. Taking the electronic device as the server 200 shown in FIG. 1 as an example, the electronic device for implementing the voice model training method of the embodiments of the present application will be described. Figure 2 The electronic device 500 shown includes: at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. Each component in the electronic device 500 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to realize the connection and communication between these components. The bus system 540 includes, in addition to the data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 540.
[0089] The processor 510 may be an integrated circuit chip with the ability to process signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0090] The user interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.
[0091] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 550 optionally includes one or more storage devices that are physically remote from the processor 510.
[0092] The memory 550 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0093] In some embodiments, the memory 550 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are described below by way of example.
[0094] The operating system 551 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0095] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, wireless fidelity (WiFi), and universal serial bus (USB), etc.;
[0096] A presentation module 553 for enabling the presentation of information (e.g., a user interface for operating a peripheral device and displaying content and information) via one or more output devices 531 associated with the user interface 530 (e.g., a display screen, a speaker, etc.);
[0097] An input processing module 554 for detecting one or more user inputs or interactions from one of one or more input devices 532 and translating the detected inputs or interactions.
[0098] In some embodiments, the training device of the speech model provided by the embodiments of the present application can be implemented in software. Figure 2 FIG. shows a training device 555 of a speech model stored in the memory 550, which can be software in the form of a program and a plug-in, etc., including the following software modules: an acquisition module 5551, an identification module 5552, and a training module 5553. These modules are logical, so they can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.
[0099] In other embodiments, the training device of the speech model provided by the embodiments of the present application can be implemented in hardware. As an example, the training device of the speech model provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the speech model provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0100] Next, the training method of the speech model provided by the embodiments of the present application will be described. In some embodiments, the training method of the speech model provided by the embodiments of the present application can be implemented independently by a terminal or a server, or jointly implemented by a terminal and a server. Taking the implementation by the server as an example, refer to Figure 3 , Figure 3 FIG. is a schematic flowchart of the training method of the speech model provided by the embodiments of the present application, which will be described in conjunction with the steps shown in Figure 3 FIG.
[0101] In step 101, the server obtains an initial speech model for speech recognition, where the initial speech model is trained based on a first speech sample carrying sample labels.
[0102] In practical applications, the first speech sample carrying labels can be the labeled speech data obtained by the server from other devices. The speech data is usually a string of continuous speech information signals. For example, the speech data can be received by the server from a terminal device. The collected speech data is usually a digital speech signal. The speech data can be from a speech assistant plugin or application, by collecting the speech of the user when using the intelligent assistant; the speech data can also be from instant chat communications of devices such as smartphones and tablets, by collecting the speech input by the user through the microphone; the speech data can also come from the sound collection in occasions such as work meeting recordings and artificial customer service calls. The embodiments of the present application do not limit the acquisition sources and acquisition methods of the speech data.
[0103] After the server obtains the speech data, it preprocesses the speech data. The preprocessing process includes pre-filtering, pre-emphasis, windowing and framing, and endpoint detection to obtain an initial speech signal.
[0104] For example, first, it is necessary to pre-filter and sample the speech signal. Usually, a band-pass filter is used for filtering, and then the original discrete signal is quantized to exclude the interference of signals with frequencies other than human vocalizations and the 50 Hertz (Hz) power frequency; the pre-emphasis technique is to smooth the connection segment between the high-frequency and low-frequency parts of the signal to make the spectrum of the speech signal smooth; the windowing and framing operation is to divide the continuous signal into independent parts with stable frequency domains using collection windows of different lengths; finally, endpoint detection is performed to correctly judge the start and end points of the input speech signal; the purpose of preprocessing the speech data is to eliminate the influence of factors such as aliasing, high-order harmonic distortion, and high frequency caused by the human vocal organs themselves and the devices for collecting speech signals on the quality of the speech signal.
[0105] Among them, after the speech signal is windowed and framed, speech feature extraction is performed on the speech signal. After speech feature extraction, a speech feature vector is obtained. In actual implementation, the server can divide the speech signal into multiple frames with a preset duration. For example, the speech signal is divided into multiple frames with a length of 25 milliseconds; after the speech signal is framed, each frame of speech waveform is converted into a multi-dimensional vector; converting each frame into a multi-dimensional vector can be: extracting feature parameters on each frame of the speech segment to form a speech feature sequence, and processing the speech feature sequence to obtain a speech feature vector; among them, the feature parameters can be linear predictive cepstral coefficients or simulate the human ear's auditory model and extract Mel frequency cepstral coefficients through Fourier transform, and can also be other types of speech features extracted from the speech data. The embodiments of the present application do not limit this.
[0106] In some embodiments, the preprocessed speech data can be converted into speech feature vectors for each frame. The speech feature vectors are converted into corresponding phonemes through a speech model, and words corresponding to each phoneme are obtained according to the pronunciation dictionary of phonemes and words, that is, the speech feature vectors for each frame are correspondingly converted into multiple possible phonemes, and probabilities of the multiple phonemes are given. Combining the mapping relationship between phonemes and the pronunciation dictionary, multiple possible words corresponding to the speech feature vectors for each frame and probabilities of each word are obtained. Then, grammatical recognition is performed on the obtained words, that is, the words are arranged and combined according to the possibility of consecutive occurrence, and the path of the word sequence is searched in the decoding network through the relevance between words, and multiple word sequences and probabilities of the word sequences are obtained.
[0107] Exemplarily, referring to Figure 4 , Figure 4 is a schematic diagram of a speech sample provided by an embodiment of the present application. In the figure, a speech signal "nihao" is input. Each frame represents a frame of data. The 1st, 2nd, 3rd, and 4th frames correspond to the pronunciation of n, the 5th, 6th, and 7th frames correspond to the phoneme of i, the 8th and 9th frames correspond to the phoneme of h, the 10th and 11th frames correspond to the phoneme of a, and the 12th frame corresponds to the phoneme of o. (Here, each letter is temporarily regarded as a pronunciation phoneme).
[0108] In step 102, through an initial speech model, speech recognition is performed on the second speech sample to obtain a recognition result, and the obtained recognition result is used as the pseudo-label of the second speech sample.
[0109] In practical applications, during the process of training a speech model by a server, first, an initial speech model is trained using a small amount of labeled data (i.e., the first speech sample), and then based on semi-supervised learning, the initial model is used to recognize a large amount of unlabeled data (i.e., the second speech sample without a label) to obtain corresponding recognition results, and the recognition results are used as the pseudo-labels of the second speech sample. In this way, in the case where the amount of data of the first speech sample is small, the initial speech model can be trained in a semi-supervised learning manner based on the pseudo-labels of the second speech sample, and at the same time, the initial speech model can be trained in a contrastive learning manner based on the second speech sample. At this time, during the pre-training process of the speech model, the loss function corresponding to the pseudo-label information and the loss function corresponding to unsupervised contrastive learning can be used to jointly optimize the model parameters of the speech model.
[0110] A description is given of semi-supervised learning. In a real speech model training scenario, there are usually labeled speech samples (the first speech samples) and a large number of unlabeled speech samples (the second speech samples). If the label data information and the hidden information in the unlabeled data can be used simultaneously for joint optimization of the model, a large amount of unlabeled data can be used to further improve the performance of the speech model. This model training method that uses both unlabeled and labeled data simultaneously can be called semi-supervised learning. By jointly optimizing through semi-supervised learning and unsupervised contrast learning, all data information can be fully utilized during the speech model training process to learn a more robust and downstream task-matching speech representation, thereby making the recognition accuracy of the speech model higher.
[0111] In actual implementation, refer to Figure 5 , Figure 5 which is the flowchart of speech model training provided by an embodiment of the present application. In the figure, the server obtains an initial speech model, and the initial speech model is trained based on a small number of labeled first speech samples; then a large number of unlabeled second speech samples are input into the initial speech model for speech recognition, and the recognition results corresponding to the second speech samples are obtained and used as the pseudo-labels of the second speech samples. In this way, semi-supervised learning and contrast learning based on pseudo-labels can be performed on the second speech samples simultaneously.
[0112] In step 103, the initial speech model is jointly trained based on the second speech samples with pseudo-labels and the first speech samples with sample labels to obtain a target speech model.
[0113] In practical applications, the initial speech model is trained with the labeled first speech samples, and then based on the initial speech model, pseudo-labels are attached to the second speech samples. Supervised learning is performed on the labels of the first speech samples and the pseudo-labels of the second speech samples, and contrast learning is performed on the first speech samples and the second speech samples. The loss function corresponding to the supervised learning based on the labels and pseudo-labels and the loss function corresponding to the contrast learning are used to jointly train the initial speech model to obtain a converged target speech model for speech recognition.
[0114] In some embodiments, refer to Figure 6 , Figure 6 which is a schematic flowchart of the joint training method of the speech model provided by an embodiment of the present application. Figure 3 The step 103 shown can be implemented through steps 1031 to 1034, and each step will be described in combination.
[0115] In step 1031, the server uses the initial speech model to predict the training samples in the joint training sample set, and obtains the prediction results. Among them, the joint training sample set includes: the first speech sample carrying the sample label and the second speech sample carrying the pseudo-label.
[0116] In actual implementation, after the server uses the initial speech model to label the pseudo-labels for the second speech samples, the first speech samples carrying the labels and the second speech samples carrying the pseudo-labels are used as the joint training sample set to jointly train the initial speech model. It can be understood that at this time, the second speech sample can be regarded as the first speech sample carrying the label (the label at this time is the pseudo-label and there may be incorrect labels). During the training process of the speech model, based on the label of the first speech sample and the pseudo-label of the second speech sample, supervised learning can be performed on the speech model, that is, the set target loss function is used to represent the correlation between the recognition result of the speech model and the label. The server iteratively updates the model parameters of the speech model according to the target loss value determined by the target loss function. It should be noted that when the joint training sample is the first speech sample, the label is the label carried by the first speech sample; when the joint training sample is the second speech sample, the label is the pseudo-label of the second speech sample. At the same time, unsupervised contrast learning can also be performed based on the first speech sample and the second speech sample, and the model parameters of the speech model are iteratively updated based on the contrast loss value determined by the set contrast loss function.
[0117] In step 1032, obtain the difference between the prediction result and the label of the training sample, and determine the value of the target loss function based on the difference.
[0118] In actual implementation, since the training sample can be the first speech sample carrying the label, that is, the first speech sample is the labeled speech sample, or it can be the second speech sample carrying the pseudo-label. The pseudo-label carried by the second speech sample is obtained by predicting with the initial speech model, that is, the label of the second speech sample is marked by the initial speech model. Determine the prediction result of the training sample through the trained initial speech model, and calculate the difference between the prediction result and the label (when the training sample is the second speech sample, the label refers to the pseudo-label), and determine the value of the target loss function. It can be understood that the target loss function is determined based on a small amount of labeled speech data and a large amount of unlabeled speech data.
[0119] In actual implementation, the determination of the target loss function can be continuous time series classification (CTC, Connectionist Temporal Classificatio). As a loss function, CTC can be used to measure how much the input sequence data differs from the true output after passing through the neural network.
[0120] Continue to explain CTC. As a sequence-to-sequence speech model training method, CTC does not require pre-aligning the data in advance. It only needs an input sequence and an output sequence to train. In this way, there is no need to align and label the data one by one, and CTC directly outputs the probability of sequence prediction without external post-processing.
[0121] During the training process of CTC, a speech input sequence X of length T = [x1, x2, x3, ……, x T , and the corresponding output label sequence Y = [y1, y2, y3, ……, y U . CTC gives all possible output distributions p(π|C1, C2, C3, ……, C T ), where C1, C2, C3, ……, C T are the computational outputs of the speech model, and π represents one of the possible output distributions. According to this distribution, the most likely result can be output or the probability of a certain output can be given. The loss function of CTC can be defined as: for a given input X, a model can be trained to maximize the probability P(Y|X) of the correct output sequence;
[0122]
[0123] In the above formula, π represents all the sequence paths that make up Y output by the speech model; p(π|C1, C2, …, C T ) represents the probability of a sequence path (the probability of the output label sequence Y);
[0124] means that the probability of the output label sequence Y is the sum of the probabilities of multiple sequence paths.
[0125] Exemplarily, input 200 frames of audio data, and the true output is the 5 ordered phonemes of "nihao". After being processed by the speech model, the output is still data with a sequence length of 200. Suppose two people both say the sentence "nihao", and their true output results are both the 5 ordered phonemes of "nihao". However, because each person has different pronunciation characteristics, for example, some people speak fast and some people speak slow. After the original speech data is calculated by the speech model, the result obtained by the first person may be: nnnniiiiii…hhhhhaaaaaooo (with a length of 200), and the result obtained by the second person's speech may be: niiiii…hhhhhaaaaaooo (with a length of 200). Both of these results belong to correct prediction results. It can be imagined that for data with a length of 200, there are many results that can finally correspond to the pronunciation order of "nihao". CTC is a method used in such cases where the sequence has multiple possibilities to calculate the loss value with the final true sequence value. See Figure 7 , Figure 7 Figure Figure 7 is an example diagram of continuous time series classification provided by an embodiment of the present application. In the figure, after the speech sample "nihao" is subjected to feature extraction, 30 frame segments are generated, and the figure shows two sequence paths (shown as q and r in the figure) of the speech model output result being "nihao". Through the CTC loss, the target sequence for "nihao" can be determined.
[0126] In step 1033, based on the training sample and the prediction result, contrastive learning is performed to determine the value of the contrastive loss function corresponding to the training sample.
[0127] In some embodiments, see Figure 8 , Figure 8 Figure Figure 8 is a schematic diagram of the process for determining the value of the contrastive loss function provided by an embodiment of the present application. Figure 6 The shown step 1033 can be implemented through steps 201 to 204, and will be described in combination with each step.
[0128] Step 201, the server constructs corresponding positive samples and negative samples based on the training sample.
[0129] In some embodiments, see Figure 9 , Figure 9 Figure Figure 9 is a schematic diagram of the process for constructing positive and negative samples provided by an embodiment of the present application. Figure 8 The shown step 201 can be implemented through steps 2011 to 2013, and will be described in combination with each step.
[0130] In step 2021, the server extracts features from the training samples to obtain the initial speech features corresponding to the training samples; in step 2022, the initial speech features are discretized to obtain the target speech features as the positive samples corresponding to the training samples; in step 2023, noise features corresponding to the initial speech features are generated as the negative samples corresponding to the training samples.
[0131] In actual implementation, the noise features corresponding to the initial speech features can be used as the negative samples of the speech model trained based on contrastive learning.
[0132] In step 202, based on the positive samples and the prediction results, the value of the contrastive loss function is determined as the first loss between the positive samples and the training samples; in step 203, based on the negative samples and the prediction results, the value of the contrastive loss function is determined as the second loss between the negative samples and the training samples; in step 204, based on the first loss and the second loss, the value of the contrastive loss function corresponding to the training samples is determined.
[0133] In actual implementation, the loss function for training the speech model based on unsupervised contrastive learning is as follows:
[0134]
[0135] In the above formula, where q t is the positive sample, ~q t is the negative sample, is used to calculate the mutual information between samples, and k is the temperature coefficient.
[0136] In step 1034, by combining the value of the target loss function and the value of the contrastive loss function, the model parameters of the initial speech model are updated to obtain the target speech model.
[0137] In actual implementation, when the server performs joint training on the speech model, based on the target loss function and the weight corresponding to the target loss function, the contrastive loss function and the weight corresponding to the contrastive loss function, the joint loss function is determined. Among them, the joint loss function can be defined as:
[0138] L = ∑αLctc+(1 - α)Lc Formula (3)
[0139] In the above formula, Lctc is the CTC as the target loss function, and Lc is the contrastive loss function corresponding to contrastive learning. 1 - α is the weight of the contrastive loss function, and α is the weight of the target loss function.
[0140] In some embodiments, refer to Figure 10 , Figure 10 which is the schematic flowchart of the joint training method of the speech model provided by the embodiments of the present application, Figure 6The shown step 1034 can be implemented through steps 301 to 303, which will be described in combination with each step.
[0141] Step 301, the server respectively obtains the weights of the target loss function and the weights of the contrast loss function.
[0142] In actual implementation, the server obtains, in the above formula (3), the weights for the target loss function and the weights for the contrast loss function.
[0143] Step 302, based on the weights of the target loss function and the weights of the contrast loss function, perform a weighted sum of the values of the target loss function and the values of the contrast loss function to obtain a weighted sum result.
[0144] In actual implementation, when the server performs joint training on the speech model, the server substitutes the obtained weights of the target loss function and the weights of the contrast loss function into the above formula (3) to obtain a weighted sum result.
[0145] Step 303, based on the weighted sum result, update the model parameters of the initial speech model to obtain a target speech model.
[0146] In actual implementation, the server updates the model parameters of the initial speech model according to the weighted sum combination obtained in step 302 until the model converges to obtain a trained target speech model.
[0147] In some embodiments, feature exchange can also be achieved in the following manner to achieve the exchange of prediction results and the exchange of training samples: The server obtains a feature exchange ratio; based on the feature exchange ratio, exchange some features in the prediction results with some features in the training samples to obtain the exchanged prediction results and the exchanged training samples. It should be noted that the exchanged prediction results are used to determine the value of the target loss function in combination with the labels of the training samples; the exchanged training samples are used for contrast learning in combination with the prediction results to determine the value of the contrast loss function corresponding to the training samples.
[0148] In actual implementation, during the process of the server performing joint training on the speech model, in order to enable the target loss function and the contrast loss function to optimize each other, so as to finally learn relatively consistent speech representations, when determining the value of the target loss function, the input corresponding to the target loss function can be randomly replaced according to the feature exchange ratio, that is, some features in the prediction results are exchanged with some features in the training samples, so that the label information learned based on the target loss function can also guide the contrast learning training at the same time, so that the two loss functions are optimized towards a consistent training goal during the learning process.
[0149] Exemplarily, taking the CTC loss as the target loss function, the CTC loss is as shown in formula (2), then the swapped CTC loss can be changed to the following form:
[0150]
[0151] In the above formula, C' T is the output of the speech model or the output of the replaced quantizer, and the feature exchange ratio can be set to 0.5 according to the actual situation. It should be noted that when exchanging according to the feature exchange ratio, for the convenience of calculation, the speech features at the same position of the prediction result and the training sample can be exchanged.
[0152] In some embodiments, referring to Figure 11 , Figure 11 is the flowchart of the fine-tuning method for the speech model provided by the embodiments of the present application. Based on Figure 3 , after step 103, the server can also execute steps 401 to 402.
[0153] Step 401, the server obtains a third speech sample carrying a speech label.
[0154] In actual implementation, after the server obtains the trained target speech model, in order to improve the adaptability of the target speech model to various speech recognition scenarios, it can obtain the labeled speech samples of the target scenario and fine-tune the target speech model to obtain the target speech model adapted to each target scenario.
[0155] Exemplarily, for speech recognition in the navigation scenario, the server can fine-tune the trained target speech model based on a small number of labeled navigation speech samples to obtain a speech recognition model suitable for the navigation scenario.
[0156] Step 402, the server trains the target speech model based on the third speech sample to update the model parameters of the target speech model.
[0157] In actual implementation, after the server jointly trains the initial speech model based on the second speech sample carrying pseudo-labels and the first speech sample carrying labels to obtain the target speech model, it can also fine-tune the target speech model based on the third speech sample of the target speech recognition scenario and carrying the target scenario label, and finally obtain the target speech model adapted to the target scenario. In this way, the training efficiency of the speech model in the target scenario and the accuracy of speech recognition of the target speech model in the target scenario can be improved, and the universality of the speech model can be increased.
[0158] In some embodiments, referring to Figure 12 , Figure 12is a flow chart of the joint training method of the speech model provided in the embodiment of the present application, based on Figure 3 After step 103, the server may further execute step 501, where the server obtains speech data to be recognized; and step 502, where the server inputs the speech data to be recognized into a target speech model and outputs the recognized speech result.
[0159] In actual implementation, the server receives the speech recognition request sent by the speech recognition device, parses the speech recognition request, and obtains the speech data to be recognized; then the speech data to be recognized is input into the trained target speech model to obtain the recognition result, and sends the recognition result to the speech recognition device.
[0160] Exemplarily, the server receives the "nihao" voice input by the user in the voice recognition client, and outputs the Chinese text "你好" to the user terminal through the trained voice model for Chinese voice recognition.
[0161] In the embodiment of the present application, an initial speech model is obtained by training based on a first speech sample with a label, and then a second speech sample without a label is recognized through the initial speech model, and the recognition result is used as a pseudo label of the second speech sample. Then, the initial speech model is jointly trained based on the first speech sample and the second speech sample. In this way, the speech model can be trained through a small amount of labeled speech and unlabeled speech samples, reducing the annotation cost for speech samples and improving the recognition accuracy of the speech model; and in the speech recognition task in the target field, a small amount of manually annotated data is used to fine-tune the speech model, which can train a speech model that is adapted to the target field and has a high accuracy on the basis of reducing manual workload and data annotation costs, thereby improving the performance of downstream speech recognition training tasks.
[0162] The following is an explanation of an exemplary application of the embodiments of the present application in a practical application scenario.
[0163] Speech recognition model training usually requires a large amount of labeled audio data to achieve good performance. In the absence of labeled data, pre-training of neural networks has become an effective technique. The key idea is to first perform unsupervised pre-training on a large amount of labeled or unlabeled data, and then fine-tune the training on the target data with limited data to improve the performance of downstream tasks. This pre-training method is particularly effective for tasks that require a lot of work to obtain labeled data (such as speech recognition).
[0164] In the related art, see Figure 13 , Figure 13It is a schematic diagram of the unsupervised training process of the speech model provided by the related technology. Contrastive learning is a commonly used unsupervised training algorithm. Through constructing positive samples and negative samples and comparing them in the feature space, contrastive learning learns the latent feature representation of the model. This method aims to learn the latent speech representation by maximizing the mutual information between positive and negative samples through contrastive learning.
[0165] Exemplarily, refer to Figure 14 , Figure 14 It is a flowchart of contrastive learning for the speech model provided by the related technology. First, the continuous speech signal X is framed to obtain a speech sequence {x1, x2, x3, x4, x5, x6, …… xN} containing N (N≥1 and N is an integer) elements. Feature extraction (feature downsampling) is performed on the speech sequence to obtain the corresponding speech features. Then, a masking (mask) operation is performed on the speech features. The masked signal is passed through a quantizer to convert the continuous speech signal into a discrete representation, and these representations are used as positive samples. At the same time, a large number of negative samples are constructed to carry out the contrastive learning process.
[0166] In the related technology, contrastive learning learns the latent speech representation through an unsupervised learning method and is used for downstream supervised training tasks. However, this completely unsupervised training process usually requires careful design and improvement, and the information learned may not match the true label information, thus affecting the performance of downstream speech recognition tasks. In the real speech model training scenario, there are usually labeled training data and a large amount of unlabeled data. If the label data information and the hidden information in the unlabeled data (the data distribution information corresponding to the unlabeled data) can be used simultaneously for joint optimization of the model, the performance of the model will be further improved. This model training method that uses both unlabeled and labeled data can be called semi-supervised learning.
[0167] Continue to illustrate semi-supervised learning. Semi-supervised learning is a model training process between unsupervised and supervised learning. This model training method can improve the performance of the model trained with labeled data through a large amount of unlabeled data. In practical applications, the method based on pseudo-labels is a semi-supervised learning method. This method does not require manual annotation of a large amount of unlabeled data. Instead, on the model obtained through supervised training (i.e., the model trained with labeled data), the unlabeled speech signal is input to obtain approximate pseudo-labels, which form a new set of pseudo-label data. The final training process can combine these pseudo-label data and the label data to train a new speech model together.
[0168] Based on this, the embodiments of the present application propose a method for training a speech model that combines semi-supervised learning and contrastive learning. During the training process of the speech model, the embodiments of the present application can optimize the speech model by combining semi-supervised learning and unsupervised contrastive learning, helping the speech model make full use of all data (labeled data and unlabeled data) information to learn more robust and downstream task-matching speech representations.
[0169] In some embodiments, referring to Figure 15 , Figure 15 is a schematic diagram of the pre-training process of the speech model provided by the embodiments of the present application. During the process of training the speech model by the server or terminal, first, a small amount of labeled data is used to train an initial speech model (also called a seed model), and then based on semi-supervised learning, the initial speech model is used to predict a large amount of unlabeled data (unlabeled data), obtaining corresponding prediction results, and using these prediction results as pseudo-labels for the unlabeled data. In this way, in the step 1 stage shown in the figure, a large amount of unlabeled data can be pre-trained using approximate label information (pseudo-labels) and contrastive learning. At this time, the speech model can jointly optimize the model parameters using the loss function corresponding to the pseudo-label information and the loss function of unsupervised contrastive learning during the pre-training process. This joint optimization method can be defined as multi-task contrastive learning. Eventually, the model will learn more robust latent speech representations beyond a single training process. After joint pre-training, in the step 2 stage shown in the figure, the model can be fine-tuned using labeled data to obtain the final speech model.
[0170] Next, taking the loss function corresponding to the pseudo-label as the Connectionist Temporal Classification (CTC) loss as an example, the multi-task pre-training method based on CTC and contrastive learning will be described. Referring to Figure 16 , Figure 16 is a schematic diagram of the joint training of the speech model provided by the embodiments of the present application. Based on semi-supervised learning, after obtaining the pseudo-label information corresponding to the unlabeled data, the data can be further subjected to multi-task joint training. Among them, for the pseudo-label information, the CTC loss can be used for training.
[0171] During the process of training the speech model in combination with CTC, a speech input sequence X of length T = [x1, x2, x3, ……, x T , and the corresponding output label sequence Y = [y1, y2, y3, ……, y U , CTC gives all possible output distributions p(π|C1, C2, C3, ……, C T ) of the input sequence X, where C1, C2, C3, ……, CT The computational output π of the speech model represents one set of possible output distributions. According to this distribution, the most likely result can be output or the probability of a certain output can be given. The loss function of CTC can be defined as follows: for a given input X, a model can be trained to maximize the probability P(Y|X) of the correct output sequence.
[0172] In the above formula (1), π represents all the sequence paths that make up Y output by the neural network model; p(π|C1, C2, …, C T ) represents the probability of a sequence path (the output label sequence Y). It means that the probability of the output label sequence Y is the sum of the probabilities of multiple sequence paths.
[0173] Secondly, in the loss function for training the speech model based on unsupervised contrastive learning as shown in the above formula (2), where q t is the positive sample, ~q t is the negative sample, is used to calculate the mutual information between samples, and k is the temperature coefficient.
[0174] Based on the above CTC supervised learning, the speech model can directly use the label sequence information to clearly direct the input speech sample to a speech unit. At the same time, the potential speech representation information is also obtained through unsupervised contrastive learning. Finally, these two representation information are combined. The loss function for multi-task training of the speech model is as shown in the above formula (3), where Lctc is the CTC loss function and Lc is the contrastive learning loss function.
[0175] In the process of jointly training the speech model, real data (data with labels, also called labeled data) and pseudo-label data (data without labels, also called unlabeled data) can be jointly trained at the same time. Among them, the labeled data and the unlabeled data learn the supervised speech representation information through the CTC loss, while the unsupervised contrastive learning process learns the unsupervised speech representation information through the quantizer.
[0176] However, in the specific implementation process, the representation information learned by the two loss functions is independent of each other, and directly performing multi-task training has limited improvement in the final learned speech representation information.
[0177] Therefore, in order to ensure that the two learning processes can optimize each other and finally learn relatively consistent and better speech representations, the calculation input of each CTC loss can be randomly replaced. Then the exchanged ctc loss is as shown in the above formula (4), where C' T is Figure 16The output of the shown speech model or the output of the replaced quantizer, and the replacement probability can be set to 0.5 according to the actual situation. The learning objective is to enable the label information learned by CTC to also guide the contrastive learning training at the same time, so that the two loss functions are optimized towards a consistent training objective during the learning process, and finally a more robust and consistent pre-trained speech representation is learned.
[0178] To verify the effectiveness of the speech model training method proposed in this application embodiment, experiments are carried out on a mixed Chinese dataset with both labeled data and unlabeled data, and the task of the experiment is speech recognition. For unsupervised training, the experiment will first perform contrastive learning pre-training using all data, and then fine-tune using the labeled data. For semi-supervised pseudo-label learning and semi-supervised contrastive learning, the experiment will first train an initial seed model using the labeled data, and then perform fine-tuning training using the pseudo-labels generated by the seed model. Refer to Table 1, which is the experimental result table of the speech model training provided in this application embodiment.
[0179] Pre-training method Fine-tuning training result (word error rate) Initial model baseline 16.54 Contrastive learning (wav2vec2) 15.98 Pseudo-label training 15.97 Semi-supervised contrastive learning 15.53
[0180] Table 1
[0181] The experimental results shown in Table 1 indicate that the combined semi-supervised contrastive learning method has significantly better fine-tuning training results for the speech model than other training methods. Compared with directly using contrastive learning or pseudo-label semi-supervised learning, the speech model training method proposed in this application embodiment can learn a speech pre-training model with better performance and further improve the performance of downstream speech recognition tasks.
[0182] The speech model training method provided in this application embodiment can be used in the pre-training process of the speech model. Through the semi-supervised learning idea, pseudo-label information is added in the pre-training process of the speech model, and the model is jointly pre-trained by combining contrastive learning and CTC learning based on pseudo-labels. Finally, the model will learn more robust potential speech representation information and significantly improve the performance of downstream speech tasks.
[0183] Next, continue to describe the exemplary structure of the software module implementation of the speech model training device 555 provided in this application embodiment. In some embodiments, as Figure 2 shown, the software module in the speech model training device 555 stored in the memory 550 may include:
[0184] An acquisition module 5551, configured to acquire an initial speech model for speech recognition, where the initial speech model is trained based on a first speech sample carrying sample labels;
[0185] The recognition module 5552 is used to perform speech recognition on the second speech sample through the initial speech model to obtain a recognition result, and use the recognition result as the pseudo-label of the second speech sample;
[0186] The training module 5553 is used to jointly train the initial speech model based on the second speech sample carrying the pseudo-label and the first speech sample carrying the sample label to obtain a target speech model.
[0187] In some embodiments, the training module is further used to predict the training samples in the joint training sample set through the initial speech model to obtain a prediction result; wherein, the joint training sample set includes: the first speech sample carrying the sample label and the second speech sample carrying the pseudo-label; obtain the difference between the prediction result and the label of the training sample, and determine the value of the target loss function based on the difference; perform contrast learning based on the training sample and the prediction result to determine the value of the contrast loss function corresponding to the training sample; combine the value of the target loss function and the value of the contrast loss function to update the model parameters of the initial speech model to obtain a target speech model.
[0188] In some embodiments, the training module is further used to construct corresponding positive samples and negative samples based on the training sample; determine the value of the contrast loss function based on the positive sample and the prediction result as the first loss between the positive sample and the training sample; determine the value of the contrast loss function based on the negative sample and the prediction result as the second loss between the negative sample and the training sample; determine the value of the contrast loss function corresponding to the training sample based on the first loss and the second loss.
[0189] In some embodiments, the training module is further used to extract features from the training sample to obtain the initial speech features corresponding to the training sample; perform discretization processing on the initial speech features to obtain target speech features as the positive sample corresponding to the training sample; generate noise features corresponding to the initial speech features as the negative sample corresponding to the training sample.
[0190] In some embodiments, the training module is further used to obtain a feature exchange ratio; based on the feature exchange ratio, exchange some features in the prediction result with some features in the training sample to obtain an exchanged prediction result and an exchanged training sample; wherein, the exchanged prediction result is used to determine the value of the target loss function in combination with the label of the training sample; the exchanged training sample is used to perform contrast learning in combination with the prediction result to determine the value of the contrast loss function corresponding to the training sample.
[0191] In some embodiments, the training module is further configured to respectively obtain the weights of the target loss function and the weights of the contrastive loss function; based on the weights of the target loss function and the weights of the contrastive loss function, perform a weighted sum of the values of the target loss function and the values of the contrastive loss function to obtain a weighted sum result; and based on the weighted sum result, update the model parameters of the initial speech model to obtain a target speech model.
[0192] In some embodiments, the training module is further configured to obtain a third speech sample carrying a speech label; and based on the third speech sample, train the target speech model to update the model parameters of the target speech model.
[0193] In some embodiments, the trained target speech model is further configured to obtain speech data to be recognized;
[0194] Input the speech data to be recognized into the target speech model, and output the recognized speech result.
[0195] An embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned speech model training method of the embodiment of the present application.
[0196] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions are stored, and when the executable instructions are executed by a processor, the processor will be caused to execute the method provided by the embodiment of the present application, for example, as Figure 3 shown in the method.
[0197] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.
[0198] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0199] As an example, the executable instructions may or may not correspond to files in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts stored in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code).
[0200] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.
[0201] In summary, through the embodiments of the present application, based on contrastive learning pre-training, while utilizing the pseudo-labeling technique in semi-supervised learning, with a small amount of labeled data, jointly pre-training the speech model using CTC and contrastive learning can help the speech model learn better and more robust latent information representations of speech, thereby improving the performance of downstream speech recognition model tasks.
[0202] The above description is only for the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A training method for a speech model, characterized in that, The method includes: Obtaining an initial speech model for speech recognition, where the initial speech model is trained based on a first speech sample carrying a sample label; Performing speech recognition on a second speech sample through the initial speech model to obtain a recognition result, and using the recognition result as the pseudo-label of the second speech sample; Predicting a training sample in a joint training sample set through the initial speech model to obtain a prediction result; wherein, the joint training sample set includes: the first speech sample carrying the sample label and the second speech sample carrying the pseudo-label; Randomly exchanging the prediction result with the training sample to obtain an exchanged prediction result and an exchanged training sample; the exchanged prediction result is used to determine the value of a target loss function in combination with the label of the training sample; the exchanged training sample is used to perform contrastive learning in combination with the prediction result to determine the value of a contrastive loss function corresponding to the training sample; Combining the value of the target loss function and the value of the contrastive loss function to update the model parameters of the initial speech model to obtain a target speech model.
2. The method according to claim 1, characterized in that, The method further includes: Constructing corresponding positive samples and negative samples based on the training sample; Determining the value of the contrastive loss function as the first loss between the positive sample and the training sample based on the positive sample and the prediction result; Determining the value of the contrastive loss function as the second loss between the negative sample and the training sample based on the negative sample and the prediction result; Determining the value of the contrastive loss function corresponding to the training sample based on the first loss and the second loss.
3. The method according to claim 2, wherein The constructing corresponding positive samples and negative samples based on the training sample includes: Performing feature extraction on the training sample to obtain an initial speech feature corresponding to the training sample; Performing discretization processing on the initial speech feature to obtain a target speech feature as the positive sample corresponding to the training sample; Generating a noise feature corresponding to the initial speech feature as the negative sample corresponding to the training sample.
4. The method according to claim 1, characterized in that For the randomly exchanging the prediction result with the training sample to obtain an exchanged prediction result and an exchanged training sample, the method further includes: Obtaining a feature exchange ratio; Based on the feature exchange ratio, exchanging some features in the prediction result with some features in the training sample to obtain an exchanged prediction result and an exchanged training sample.
5. The method according to claim 1, characterized in that, The combining the value of the target loss function and the value of the contrastive loss function to update the model parameters of the initial speech model to obtain a target speech model includes: Respectively obtaining the weight of the target loss function and the weight of the contrastive loss function; Performing weighted summation on the value of the target loss function and the value of the contrastive loss function based on the weight of the target loss function and the weight of the contrastive loss function to obtain a weighted summation result; Updating the model parameters of the initial speech model based on the weighted summation result to obtain a target speech model.
6. The method according to claim 1, wherein After obtaining the target speech model, the method further includes: Obtain a third speech sample carrying a speech label; Based on the third speech sample, train the target speech model to update the model parameters of the target speech model.
7. The method according to claim 1 or 6, characterized in that, The method further includes: Obtain speech data to be recognized; Input the speech data to be recognized into the target speech model, and output the recognized speech result.
8. A training device for a speech model, characterized in that The device includes: An acquisition module, configured to obtain an initial speech model for speech recognition, where the initial speech model is trained based on a first speech sample carrying a sample label; A recognition module, configured to perform speech recognition on a second speech sample through the initial speech model, obtain a recognition result, and use the recognition result as the pseudo-label of the second speech sample; A training module, configured to predict a training sample in a joint training sample set through the initial speech model to obtain a prediction result; where the joint training sample set includes: the first speech sample carrying the sample label and the second speech sample carrying the pseudo-label; randomly exchange the prediction result and the training sample to obtain an exchanged prediction result and an exchanged training sample; the exchanged prediction result is used to determine the value of a target loss function in combination with the label of the training sample; the exchanged training sample is used to perform contrastive learning in combination with the prediction result to determine the value of a contrastive loss function corresponding to the training sample; combine the value of the target loss function and the value of the contrastive loss function, and update the model parameters of the initial speech model to obtain a target speech model.
9. An electronic device, characterized in that, The electronic device includes: A memory, configured to store executable instructions; A processor, configured to implement the training method of the speech model according to any one of claims 1 to 7 when executing the executable instructions stored in the memory.
10. A computer-readable storage medium storing executable instructions, characterized in that, The executable instructions, when executed by the processor, implement the training method of the speech model according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instructions, when executed by the processor, implement the training method of the speech model according to any one of claims 1 to 7.
Citation Information
Patent Citations
Model training method and language recognition method, device and equipment
CN110853617A
Training data screening method, system and device, and medium
CN113901992A