A speech recognition method and device
Through the speech recognition method combined with feature extraction and mapping networks, the problem of high manual labeling costs and inability to utilize labeled data in the prior art is solved, and the speech recognition performance and labeling accuracy are improved.
Patent Information
- Application Number
- CN202210312961.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-03-28
AI Technical Summary
Existing speech recognition technologies have problems such as high cost of manual labeling, inability to effectively utilize unlabeled audio data, ignoring voice phase information, and insufficient ability to model complex speech characteristics.
The pre-trained features of labelless audio data are extracted through the feature extraction network, and the normalized weight vector of phonemes is obtained by using the feature mapping network, and the speech recognition model is trained in combination with labeled data, effectively using labelless data to improve recognition performance.
It reduces the cost of manual labeling, improves the accuracy of labeling, and is suitable for ultra-large-scale speech recognition training, solving the problem of speech phase information being ignored and the lack of modeling capabilities of complex speech characteristics.
Smart Images

Figure CN114550702B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a speech recognition method and device. Background Art
[0002] Speech recognition technology aims to solve the problem of converting speech audio signals into speech text. Based on the speech recognition results, the purpose of human-computer interaction can be achieved by integrating natural language understanding, multimodal fusion and other technologies. The current speech recognition system usually adopts a supervised training scheme, that is, based on manual annotation of the collected audio data, the classifier for speech recognition is trained based on the original audio data and features, with text annotation as the ultimate goal. Currently, the commonly used speech recognition technologies are divided into two categories. One is a hybrid framework based on hidden Markov deep neural network (HMM-DNN), which is divided into two modules: acoustic model and language model. The decoding algorithm uses Viterbi search in the recognition process to obtain the optimal sequence and generate the decoding output. The other type of speech recognition algorithm is based on end-to-end neural network design, and the optimization target is designed through the CTC (Connectionist Temporal Classification) criterion, so that the neural network directly outputs the recognition text results according to the original audio features.
[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0004] The cost of manual labeling is usually high, time-consuming, and the labeling quality needs to be checked, which makes it unsuitable for ultra-large-scale speech recognition training. It is also impossible to effectively utilize the large amount of unlabeled audio data generated every day in existing speech recognition products. Traditional speech recognition systems based on features such as MFCC (Mel-frequency cepstral coefficients) and FBANK (filter bank features) ignore speech phase information, and based on simplified filter theory extraction, their ability to model complex speech characteristics still has certain defects. Summary of the invention
[0005] In view of this, the embodiments of the present invention provide a speech recognition method and device, which can solve the data dependence and speech representation problems of speech recognition in various business fields and application scenarios, and can effectively utilize a large amount of unlabeled audio data in existing speech recognition products to improve the performance of speech recognition, reduce the cost of manual labeling, reduce the time of labeling, and improve the accuracy of labeling. It is suitable for the training of ultra-large-scale speech recognition and solves the problems in the prior art of ignoring speech phase information and having defects in the ability to model complex speech characteristics.
[0006] To achieve the above objective, according to one aspect of an embodiment of the present invention, a speech recognition method is provided.
[0007] A speech recognition method comprises: extracting pre-trained features corresponding to an unlabeled first audio data sample through a feature extraction network, obtaining a normalized weight vector of the phonemes of the first audio data sample through a feature mapping network based on the pre-trained features corresponding to the first audio data sample, wherein the normalized weight vector represents the category of the phonemes of the first audio data sample; using the normalized weight vector as a training target corresponding to the first audio data sample, and using the label of the labeled second audio data sample as the training target corresponding to the second audio data sample, training a speech recognition model using the first audio data sample and the second audio data sample, and performing speech recognition using the trained speech recognition model, wherein the label of the second audio data sample represents the category of the phonemes of the second audio data sample.
[0008] Optionally, before extracting the pre-trained features corresponding to the unlabeled first audio data sample through the feature extraction network, it includes: extracting the pre-trained features corresponding to the labeled third audio data sample through the feature extraction network; using the pre-trained features corresponding to the third audio data sample as the input of the feature mapping network, and using the label of the third audio data sample as the training target, training the feature mapping network, the label of the third audio data sample represents the category of the phoneme of the third audio data sample.
[0009] Optionally, before extracting the pre-trained features corresponding to the labeled third audio data samples through the feature extraction network, it includes: constructing training samples of the feature extraction network using unlabeled fourth audio data samples, wherein each combination of multiple training samples obtains a training sample subset; inputting the training sample subset into the feature extraction network to obtain a network output result corresponding to each training sample in the training sample subset; clustering the network output results corresponding to each training sample to obtain a pairing combination of training samples and cluster centers, and updating the cluster centers according to the pairing combination; using a clustering criterion function as a loss function during training of the feature extraction network, and updating the network parameters of the feature extraction network through back propagation, wherein the clustering criterion function is constructed based on the network output results and the cluster centers.
[0010] Optionally, the loss function is constructed in the following manner: the network output result of the i-th training sample is used as the i-th target sample y i , with c k Indicates the distance from the target sample y i The nearest cluster center, constructing the first relation: the i-th target sample y i and the distance target sample y i The nearest cluster center c kThe square of the absolute value of the difference between the two is used to construct the second relationship: within the range of i from 1 to M, the first relationship corresponding to each value of i is calculated and summed, where M is the number of training samples in a single subset of the training samples; the second relationship is used as the loss function.
[0011] Optionally, the method of constructing the training samples of the feature extraction network using the unlabeled fourth audio data samples includes: performing time-frequency transformation on the fourth audio data samples to obtain frame-level original speech features of the fourth audio data samples, the frame-level original speech features including amplitude spectrum vectors and phase spectrum vectors of each frame of the fourth audio data samples; based on the frame-level original speech features, fusing the context information of each frame of the audio signal in the amplitude spectrum vector and the phase spectrum vector to construct the training samples of the feature extraction network.
[0012] Optionally, based on the frame-level original speech features, the context information of each frame of the audio signal in the amplitude spectrum vector and the phase spectrum vector are integrated to construct a training sample for the feature extraction network, including: for any t-th frame of original speech features in the frame-level original speech features, the amplitude spectrum vectors and phase spectrum vectors of all frames from the tD-th frame to the t+D-th frame are spliced according to preset rules to obtain the training sample of the feature extraction network corresponding to the t-th frame, wherein D represents the preset context window length.
[0013] According to another aspect of an embodiment of the present invention, a speech recognition device is provided.
[0014] A speech recognition device comprises: a normalized weight vector determination module, used to extract pre-trained features corresponding to an unlabeled first audio data sample through a feature extraction network, and based on the pre-trained features corresponding to the first audio data sample, obtain a normalized weight vector of the phoneme of the first audio data sample through a feature mapping network, wherein the normalized weight vector represents the category of the phoneme of the first audio data sample; a speech recognition model training module, used to use the normalized weight vector as a training target corresponding to the first audio data sample, and use the label of the labeled second audio data sample as the training target corresponding to the second audio data sample, and use the first audio data sample and the second audio data sample to train a speech recognition model, so as to use the trained speech recognition model to perform speech recognition, wherein the label of the second audio data sample represents the category of the phoneme of the second audio data sample.
[0015] Optionally, it also includes a feature mapping network training module, which is used to: extract pre-trained features corresponding to the labeled third audio data samples through the feature extraction network; use the pre-trained features corresponding to the third audio data samples as input to the feature mapping network, and use the labels of the third audio data samples as training targets to train the feature mapping network, where the labels of the third audio data samples represent the categories of the phonemes of the third audio data samples.
[0016] Optionally, it also includes a feature extraction network training module, which is used to: construct training samples of the feature extraction network using unlabeled fourth audio data samples, wherein each combination of multiple training samples obtains a training sample subset; input the training sample subset into the feature extraction network to obtain a network output result corresponding to each training sample in the training sample subset; cluster the network output results corresponding to each training sample to obtain a pairing combination of training samples and cluster centers, and update the cluster centers according to the pairing combination; use a clustering criterion function as a loss function during training of the feature extraction network, and update the network parameters of the feature extraction network through back propagation, wherein the clustering criterion function is constructed based on the network output results and the cluster centers.
[0017] Optionally, the loss function is constructed in the following manner: the network output result of the i-th training sample is used as the i-th target sample y i , with c k Indicates the distance from the target sample y i The nearest cluster center, constructing the first relation: the i-th target sample y i and the distance target sample y i The nearest cluster center c k The square of the absolute value of the difference between the two is used to construct the second relationship: within the range of i from 1 to M, the first relationship corresponding to each value of i is calculated and summed, where M is the number of training samples in a single subset of the training samples; the second relationship is used as the loss function.
[0018] Optionally, the feature extraction network training module includes a training sample construction submodule, which is used to: perform time-frequency transformation on the fourth audio data sample to obtain frame-level original speech features of the fourth audio data sample, wherein the frame-level original speech features include amplitude spectrum vectors and phase spectrum vectors of each frame of the fourth audio data sample; based on the frame-level original speech features, fuse the context information of each frame of the audio signal in the amplitude spectrum vector and the phase spectrum vector to construct a training sample for the feature extraction network.
[0019] Optionally, the training sample construction submodule is also used to: for any t-th frame speech original features in the frame-level speech original features, splice the amplitude spectrum vectors and phase spectrum vectors of all frames from the tD-th frame to the t+D-th frame according to preset rules, to obtain the training sample corresponding to the t-th frame of the feature extraction network, where D represents the preset context window length.
[0020] According to yet another aspect of the embodiments of the present invention, an electronic device is provided.
[0021] An electronic device includes: one or more processors; a memory for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the speech recognition method provided by the embodiment of the present invention.
[0022] According to yet another aspect of an embodiment of the present invention, a computer readable medium is provided.
[0023] A computer-readable medium stores a computer program, which, when executed by a processor, implements the speech recognition method provided by an embodiment of the present invention.
[0024] One embodiment of the above invention has the following advantages or beneficial effects: extracting pre-trained features corresponding to the unlabeled first audio data sample through a feature extraction network, obtaining the normalized weight vector of the phoneme of the first audio data sample through a feature mapping network based on the pre-trained features corresponding to the first audio data sample, and the normalized weight vector represents the category of the phoneme of the first audio data sample; using the normalized weight vector as the training target corresponding to the first audio data sample, and using the label of the labeled second audio data sample as the training target corresponding to the second audio data sample, training a speech recognition model using the first audio data sample and the second audio data sample, and performing speech recognition using the trained speech recognition model, wherein the label of the second audio data sample represents the category of the phoneme of the second audio data sample. The invention can solve the data dependency and speech representation problems of speech recognition in various business fields and application scenarios, can effectively use a large amount of unlabeled audio data in existing speech recognition products to improve the performance of speech recognition, reduce the cost of manual labeling, reduce the time consumption of labeling, and improve the accuracy of labeling, and is suitable for the training of ultra-large-scale speech recognition, and solves the problems of ignoring speech phase information and the defect of modeling complex speech characteristics in the prior art.
[0025] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with the specific implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.
[0027] Figure 1 is a schematic diagram of the main steps of a speech recognition method according to an embodiment of the present invention;
[0028] Figure 2 is a schematic diagram of a speech feature pre-training framework using clustering criteria according to an embodiment of the present invention;
[0029] Figure 3 is a schematic diagram of main modules of a speech recognition device according to an embodiment of the present invention;
[0030] Figure 4 is an exemplary system architecture diagram to which embodiments of the present invention may be applied;
[0031] Figure 5 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0033] Figure 1 Schematic diagram of the main steps of a speech recognition method according to an embodiment of the present invention; Figure 1 As shown, the speech recognition method according to an embodiment of the present invention mainly includes the following steps S101 to S102:
[0034] Step S101: extracting pre-trained features corresponding to the unlabeled first audio data sample through a feature extraction network, and obtaining a normalized weight vector of the phoneme of the first audio data sample through a feature mapping network based on the pre-trained features corresponding to the first audio data sample, wherein the normalized weight vector represents the category of the phoneme of the first audio data sample;
[0035] Step S102: Using the normalized weight vector of the phoneme of the first audio data sample as the training target corresponding to the first audio data sample, and using the label of the annotated second audio data sample as the training target corresponding to the second audio data sample, train a speech recognition model using the first audio data sample and the second audio data sample, and perform speech recognition using the trained speech recognition model, wherein the label of the second audio data sample represents the category of the phoneme of the second audio data sample.
[0036] The feature extraction network can specifically adopt various neural networks that can extract features from audio data. The pre-trained features are the outputs of the feature extraction network.
[0037] The feature map network can be a neural network classifier.
[0038] Before extracting the pre-trained features corresponding to the unlabeled first audio data sample through the feature extraction network, the pre-trained features corresponding to the labeled third audio data sample can be extracted through the feature extraction network; the pre-trained features corresponding to the third audio data sample are used as the input of the feature mapping network, and the label of the third audio data sample is used as the training target to train the feature mapping network, and the label of the third audio data sample represents the category of the phoneme of the third audio data sample.
[0039] Before extracting the pre-trained features corresponding to the labeled third audio data samples through the feature extraction network, the unlabeled fourth audio data samples can be used to construct training samples for the feature extraction network, wherein each combination of multiple training samples obtains a training sample subset (batch), and the training sample subset is recorded as {x1, x2, x3, ..., x M}, where M is batchsize, which is a custom value; the training sample subset is input into the feature extraction network to obtain the network output result of each training sample in the corresponding training sample subset. The network output result of each training sample is recorded as {y1,y2,y3,...,y M}; Cluster the network output results corresponding to each training sample to obtain the paired combination of training samples and cluster centers, and update the cluster centers according to the paired combination; Use the clustering criterion function as the loss function during feature extraction network training, and update the network parameters of the feature extraction network through back propagation. The clustering criterion function is constructed based on the network output results and cluster centers.
[0040] Specifically, the loss function can be constructed as follows: the network output result of the i-th training sample is used as the i-th target sample y i , with c k Indicates the distance from the target sample y i The nearest cluster center, constructing the first relation: the i-th target sample y i and the distance target sample y i The nearest cluster center c k The square of the absolute value of the difference between |y i -c k | 2 .
[0041] The second relational expression is constructed as follows: within the range of i from 1 to M, the first relational expression corresponding to each value of i is calculated and summed, that is:
[0042] Where M is the number of training samples in a single training sample subset;
[0043] The second relation, i.e., the clustering criterion function, is used as the loss function.
[0044] The unlabeled fourth audio data samples are used to construct training samples for the feature extraction network. The specific steps include: performing time-frequency transformation on the fourth audio data samples to obtain frame-level original speech features of the fourth audio data samples, wherein the frame-level original speech features include amplitude spectrum vectors and phase spectrum vectors of each frame of the fourth audio data samples; based on the frame-level original speech features, the context information of each frame of the audio signal in the amplitude spectrum vector and the phase spectrum vector is integrated to construct training samples for the feature extraction network.
[0045] The fourth audio data sample is subjected to a time-frequency transform to obtain the original speech features at the frame level of the fourth audio data sample. Specifically, each speech time domain signal of the fourth audio data sample can be subjected to a Fourier transform to extract the two-dimensional amplitude spectrum and phase spectrum of the Fourier transform. The Fourier spectra of different frequency bands are combined, and the amplitude spectrum vector of the t-th frame is recorded as A(t), and the phase spectrum vector is recorded as P(t), to obtain the original speech features at the frame level.
[0046] Based on the original speech features at the frame level, the context information of each frame of the audio signal in the amplitude spectrum vector and the phase spectrum vector is integrated to construct a training sample for the feature extraction network. The specific steps include: for any t-th frame of original speech features at the frame level, the amplitude spectrum vectors and phase spectrum vectors of all frames from the tD-th frame to the t+D-th frame are spliced according to the preset rules to obtain the training sample corresponding to the t-th frame of the feature extraction network, where D represents the preset context window length. The training sample x(t) corresponding to the t-th frame of the feature extraction network is specifically obtained by the following splicing formula:
[0047] x(t)=[A T (tD),A T (t-D+1),...,A T (t+D),P T (tD),P T (t-D+1),...,P T (t+D)] T
[0048] Where D represents the context window length, and the above formula represents the amplitude spectrum and phase spectrum of all frames from time tD to time t+D.
[0049] In the above embodiments, audio specifically refers to speech.
[0050] The embodiment of the present invention proposes a new speech recognition framework based on semi-supervised learning. On the one hand, a clustering criterion is adopted to learn a pre-trained network for extracting audio features under the framework of unsupervised learning, so that the same phoneme has the same cluster center. On the other hand, based on the existing speech recognition model, the speech feature pre-trained network and the speech recognition model are further optimized by generating training labels for unlabeled data. The speech recognition model can be based on a traditional hybrid framework or an end-to-end speech recognition framework.
[0051] Feature extraction is one of the key factors that determine speech recognition performance. Traditionally, based on linear filter theory and integrating the characteristics of human auditory perception, a group of linear filters, called Mel filter banks, are designed to make the center frequency of the filter more densely distributed in the low-frequency range and have a larger bandwidth in the high-frequency area. Based on this filter, the filter bank features (filterbanks, FBANK) of speech can be extracted, and further discrete cosine transform can be performed to obtain Mel-frequency Cepstral Coefficients (MFCC). However, on the one hand, this type of feature ignores the speech phase information, and on the other hand, the ability to model complex speech characteristics is still somewhat defective. In recent years, neural network-based speech representation pre-training models such as WAV2VEC have been proposed, but this type of model is usually based on the modeling of speech temporal characteristics, and the output speech features are difficult to explicitly express the correspondence between speech signals and phonemes.
[0052] The embodiment of the present invention uses large-scale unlabeled speech audio data and proposes a speech feature pre-training framework using clustering criteria. A schematic diagram of a speech feature pre-training framework using clustering criteria according to an embodiment of the present invention is shown in FIG. Figure 2 shown.
[0053] First, initialize the clustering parameters in the clustering criterion function, including the number of cluster centers K, the vectors of each cluster center c1, c2, c3..., c K .
[0054] For large-scale speech data (i.e., audio signals, or audio data), time-frequency transformation is performed. First, each speech time domain signal is Fourier transformed to extract the two-dimensional amplitude spectrum and phase spectrum of the Fourier transform. The Fourier spectra of different frequency bands are combined, and the amplitude spectrum vector of the t-th frame is recorded as A(t), and the phase spectrum vector is recorded as P(t), and the original speech features at the frame level are obtained.
[0055] Based on the above frame-level original speech features, the context information of each frame of audio signal in the amplitude spectrum and phase spectrum is integrated to construct the training samples of the feature extraction network, namely:
[0056] x(t)=[A T (tD),AT (t-D+1),...,A T (t+D),P T (tD),P T (t-D+1),...,P T (t+D)] T
[0057] Where D represents the context window length, and the above formula represents the amplitude spectrum and phase spectrum of all frames from time tD to time t+D.
[0058] The multiple training samples are randomly combined into a subset (batch), denoted as {x1,x2,x3,...,x M}, where M is batchsize, which is a custom value. The input of each feature extraction network is a subset (batch). Each sample is sent to the current feature extraction network, and the network output results of each training sample are obtained, which are recorded as {y1,y2,y3,...,y M}, with each y as a target sample.
[0059] Calculate the clustering criterion function for the current batch output results:
[0060]
[0061] Among them, c k is the distance from the current target sample y i The nearest cluster center c k =argminc j |y i -c j | 2 Based on the shortest distance criterion, the pairing combination of each target sample y and the cluster center can be obtained.
[0062] According to the above clustering criterion function, the clustering criterion function is used as the loss function when training the feature extraction network, and the neural network (i.e., the feature extraction network) is back-propagated to update the network parameters of the feature extraction network. At the same time, based on the pairing combination of each target sample y and the cluster center obtained above, each cluster center is updated:
[0063]
[0064] N←N+P
[0065] Among them, N is the current cluster center c k The number of target samples included, P is the number of samples belonging to c in the current batch k The number of samples.
[0066] Since the context information is taken into consideration (i.e., the amplitude spectrum and phase spectrum of all frames from time tD to time t+D), the above-mentioned feature pre-training model method can effectively model the co-articulation phenomenon of phonemes, which can correspond to the three-factor model in speech recognition. For example, the syllables in the speech context have an impact on the pronunciation of the current syllable, such as continuous and weak pronunciation. The co-articulation phenomenon of phonemes can be modeled based on the context information.
[0067] The embodiment of the present invention uses unlabeled data to perform semi-supervised speech recognition training based on a feature extraction network. The semi-supervised speech recognition training process includes:
[0068] Step 1: For the currently labeled audio data, the phoneme label of each frame can be identified based on the existing speech recognition model. At the same time, based on the feature extraction network, the pre-trained features of each frame of audio data can be obtained at the same time. The pre-trained features are the network output results of the feature extraction network.
[0069] Step 2: Further train a neural network classifier as a feature mapping network to learn the mapping from pre-trained features to frame-level phoneme labels, and fine-tune the feature extraction network weights.
[0070] Step 3: After the above iterations are completed, for the current unlabeled audio data, the pre-trained features can be extracted according to the feature extraction network, and the softmax weight vector (normalized weight vector) of the phoneme can be obtained using the feature mapping network as the training target of the unlabeled data. The softmax weight vector of the phoneme is the output of the feature mapping network and is used to represent the category of the phoneme.
[0071] Step 4: Optimize the existing speech recognition model based on the training target of the existing labeled data (i.e., the label of the labeled data) and the training target of the unlabeled data generated above (i.e., the softmax weight vector of the phoneme).
[0072] Step 5: Repeat steps 1 to 4 once based on the updated speech recognition model, so that the unlabeled data can be trained with better labeled targets.
[0073] The feature mapping network in the above step 2 plays a role as a speech recognition acoustic model. It can play a role in model fusion by combining it with the speech recognition model through alternate training and updating.
[0074] The new speech feature pre-training network based on clustering criteria proposed in the embodiment of the present invention can generate a feature extraction network that is integrated with the speech recognition target without labeling; the proposed semi-supervised speech recognition method based on the speech feature pre-training network can realize the effective use of unlabeled data.
[0075] Figure 34 is a schematic diagram of main modules of a speech recognition device according to an embodiment of the present invention.
[0076] like Figure 3 As shown, a speech recognition device 300 according to an embodiment of the present invention mainly includes: a normalized weight vector determination module 301 and a speech recognition model training module 302 .
[0077] A normalized weight vector determination module 301 is used to extract pre-trained features corresponding to the unlabeled first audio data sample through a feature extraction network, and obtain a normalized weight vector of the phoneme of the first audio data sample through a feature mapping network based on the pre-trained features corresponding to the first audio data sample, wherein the normalized weight vector represents the category of the phoneme of the first audio data sample;
[0078] The speech recognition model training module 302 is used to use the normalized weight vector as the training target corresponding to the first audio data sample, and the label of the labeled second audio data sample as the training target corresponding to the second audio data sample, and use the first audio data sample and the second audio data sample to train the speech recognition model, so as to perform speech recognition using the trained speech recognition model, wherein the label of the second audio data sample represents the category of the phoneme of the second audio data sample.
[0079] The speech recognition device 300 may also include a feature mapping network training module, which is used to: extract pre-trained features corresponding to the labeled third audio data samples through a feature extraction network; use the pre-trained features corresponding to the third audio data samples as input to the feature mapping network, and use the labels of the third audio data samples as training targets to train the feature mapping network, where the labels of the third audio data samples represent the categories of the phonemes of the third audio data samples.
[0080] The speech recognition device 300 may also include a feature extraction network training module, which is used to: construct training samples of the feature extraction network using unlabeled fourth audio data samples, wherein each combination of multiple training samples obtains a training sample subset; input the training sample subset into the feature extraction network to obtain the network output result of each training sample in the corresponding training sample subset; cluster the network output results corresponding to each training sample to obtain a pairing combination of the training sample and the cluster center, and update the cluster center according to the pairing combination; use the clustering criterion function as the loss function during feature extraction network training, and update the network parameters of the feature extraction network through back propagation, and the clustering criterion function is constructed based on the network output results and the cluster center.
[0081] The loss function can be constructed as follows: the network output result of the i-th training sample is used as the i-th target sample y i , with c k Indicates the distance from the target sample yi The nearest cluster center, constructing the first relation: the i-th target sample y i and the distance target sample y i The nearest cluster center c k The square of the absolute value of the difference between the two is used to construct the second relational expression: within the range of i from 1 to M, the sum of the first relational expressions corresponding to the values of i is calculated, where M is the number of training samples in a single training sample subset; the second relational expression is used as the loss function.
[0082] The feature extraction network training module may include a training sample construction submodule, which is used to: perform time-frequency transformation on the fourth audio data sample to obtain the frame-level original speech features of the fourth audio data sample, the frame-level original speech features including the amplitude spectrum vector and phase spectrum vector of each frame of the fourth audio data sample; based on the frame-level original speech features, fuse the context information of each frame of the audio signal in the amplitude spectrum vector and the phase spectrum vector to construct a training sample for the feature extraction network.
[0083] The training sample construction submodule is also used for: for any t-th frame speech original features in the frame-level speech original features, concatenating the amplitude spectrum vectors and phase spectrum vectors of all frames from the tD-th frame to the t+D-th frame according to preset rules, to obtain the training sample corresponding to the t-th frame of the feature extraction network, where D represents the preset context window length.
[0084] In addition, the specific implementation content of the speech recognition device in the embodiment of the present invention has been described in detail in the speech recognition method described above, so the repeated content will not be described again here.
[0085] Figure 4 An exemplary system architecture 400 is shown to which the speech recognition method or speech recognition device according to the embodiment of the present invention can be applied.
[0086] like Figure 4 As shown, system architecture 400 may include terminal devices 401, 402, 403, network 404 and server 405. Network 404 is used to provide a medium for communication links between terminal devices 401, 402, 403 and server 405. Network 404 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0087] Users can use terminal devices 401, 402, 403 to interact with server 405 through network 404 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 401, 402, 403, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0088] The terminal devices 401 , 402 , and 403 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0089] The server 405 may be a server that provides various services, such as a background management server (only an example) that provides support for shopping websites browsed by users using the terminal devices 401, 402, and 403. The background management server may analyze and process the received audio information and other data, and feed back the processing results (such as speech recognition results - only an example) to the terminal device.
[0090] It should be noted that the speech recognition method provided in the embodiment of the present invention is generally executed by the server 405 , and accordingly, the speech recognition device is generally disposed in the server 405 .
[0091] It should be understood that Figure 4 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0092] Reference below Figure 5 , which shows a schematic diagram of the structure of a computer system 500 of a terminal device or server suitable for implementing an embodiment of the present application. Figure 5 The terminal device or server shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0093] like Figure 5 As shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage part 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0094] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed, so that a computer program read therefrom is installed into the storage section 508 as needed.
[0095] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the main step schematic diagram can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the main step schematic diagram. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above-mentioned functions defined in the system of the present application are executed.
[0096] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0097] The main step schematic diagram and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the main step schematic diagram or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or the main step schematic diagram, and the combination of the boxes in the block diagram or the main step schematic diagram can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0098] The modules involved in the embodiments of the present invention may be implemented in software or hardware. The modules described may also be arranged in a processor. For example, they may be described as: a processor including a normalized weight vector determination module and a speech recognition model training module. The names of these modules do not constitute limitations on the modules themselves in certain cases. For example, the normalized weight vector determination module may also be described as "a module for extracting pre-trained features corresponding to an unlabeled first audio data sample through a feature extraction network, and obtaining the normalized weight vector of the phonemes of the first audio data sample through a feature mapping network based on the pre-trained features corresponding to the first audio data sample".
[0099] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes: extracting pre-trained features corresponding to the unlabeled first audio data sample through a feature extraction network, and obtaining a normalized weight vector of the phoneme of the first audio data sample through a feature mapping network based on the pre-trained features corresponding to the first audio data sample, wherein the normalized weight vector represents the category of the phoneme of the first audio data sample; using the normalized weight vector as the training target corresponding to the first audio data sample, and using the label of the labeled second audio data sample as the training target corresponding to the second audio data sample, using the first audio data sample and the second audio data sample to train a speech recognition model, so as to perform speech recognition using the trained speech recognition model, wherein the label of the second audio data sample represents the category of the phoneme of the second audio data sample.
[0100] The technical solution according to the embodiments of the present invention can solve the data dependence and speech representation problems of speech recognition in various business fields and application scenarios, and can effectively utilize the large amount of unlabeled audio data generated every day in existing speech recognition products to improve the performance of existing speech recognition, reduce the cost of manual labeling, reduce the time consumption of labeling, and improve the accuracy of labeling. It is suitable for the training of ultra-large-scale speech recognition and solves the problems of ignoring speech phase information and the defects in the ability to model complex speech characteristics in the prior art.
[0101] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A speech recognition method, characterized in that: include: Extracting pre-trained features corresponding to the unlabeled first audio data sample through a feature extraction network, and obtaining a normalized weight vector of the phoneme of the first audio data sample through a feature mapping network based on the pre-trained features corresponding to the first audio data sample, wherein the normalized weight vector represents a category of the phoneme of the first audio data sample; Using the normalized weight vector as a training target corresponding to the first audio data sample, and using the label of the annotated second audio data sample as a training target corresponding to the second audio data sample, training a speech recognition model using the first audio data sample and the second audio data sample, so as to perform speech recognition using the trained speech recognition model, wherein the label of the second audio data sample represents a category of a phoneme of the second audio data sample; The method also includes: before training the feature mapping network, training the feature extraction network through the following steps: constructing training samples of the feature extraction network using unlabeled fourth audio data samples, wherein each combination of multiple training samples obtains a training sample subset; inputting the training sample subset into the feature extraction network to obtain a network output result corresponding to each training sample in the training sample subset; clustering the network output results corresponding to each training sample to obtain a pairing combination of training samples and cluster centers, and updating the cluster centers according to the pairing combination; using a clustering criterion function as a loss function during training of the feature extraction network, and updating the network parameters of the feature extraction network through back propagation, wherein the clustering criterion function is constructed based on the network output results and the cluster centers.
2. The method according to claim 1, characterized in that: Before extracting the pre-trained features corresponding to the unlabeled first audio data sample through the feature extraction network and after training the feature extraction network, the method includes: Extracting pre-trained features corresponding to the labeled third audio data samples through the feature extraction network; The feature mapping network is trained using the pre-trained features corresponding to the third audio data sample as input to the feature mapping network and using the label of the third audio data sample as a training target, wherein the label of the third audio data sample represents the category of the phoneme of the third audio data sample.
3. The method according to claim 1, characterized in that The loss function is constructed as follows: Take the network output result of the i-th training sample as the i-th target sample y i , with c k Indicates the distance from the target sample y i The nearest cluster center, constructing the first relation: the i-th target sample y i and the distance target sample y i The nearest cluster center c k The square of the absolute value of the difference between the two is used to construct the second relationship: within the range of i from 1 to M, the first relationship corresponding to each value of i is calculated and summed, where M is the number of training samples in a single subset of the training samples; the second relationship is used as the loss function.
4. The method according to claim 1, characterized in that: The step of constructing a training sample of the feature extraction network using the unlabeled fourth audio data sample includes: Performing a time-frequency transformation on the fourth audio data sample to obtain a frame-level original speech feature of the fourth audio data sample, wherein the frame-level original speech feature includes an amplitude spectrum vector and a phase spectrum vector of each frame of the fourth audio data sample; Based on the frame-level original speech features, context information of each frame of the audio signal in the amplitude spectrum vector and the phase spectrum vector is fused to construct a training sample for the feature extraction network.
5. The method according to claim 4, characterized in that The step of fusing the context information of each frame of the audio signal in the amplitude spectrum vector and the phase spectrum vector based on the original speech features at the frame level to construct a training sample for the feature extraction network includes: For any t-th frame of speech original features in the frame-level speech original features, the amplitude spectrum vectors and phase spectrum vectors of all frames from the tD-th frame to the t+D-th frame are spliced according to preset rules to obtain the training sample of the feature extraction network corresponding to the t-th frame, where D represents the preset context window length.
6. A speech recognition device, characterized in that: include: a normalized weight vector determination module, configured to extract pre-trained features corresponding to the unlabeled first audio data sample through a feature extraction network, and obtain a normalized weight vector of the phoneme of the first audio data sample through a feature mapping network based on the pre-trained features corresponding to the first audio data sample, wherein the normalized weight vector represents the category of the phoneme of the first audio data sample; a speech recognition model training module, configured to use the normalized weight vector as a training target corresponding to the first audio data sample, and use the label of the annotated second audio data sample as a training target corresponding to the second audio data sample, and use the first audio data sample and the second audio data sample to train a speech recognition model, so as to perform speech recognition using the trained speech recognition model, wherein the label of the second audio data sample represents a category of a phoneme of the second audio data sample; A feature extraction network training module is used to: train the feature extraction network through the following steps before training the feature mapping network: construct training samples of the feature extraction network using unlabeled fourth audio data samples, wherein each combination of multiple training samples obtains a training sample subset; input the training sample subset into the feature extraction network to obtain a network output result corresponding to each training sample in the training sample subset; cluster the network output results corresponding to each training sample to obtain a pairing combination of training samples and cluster centers, and update the cluster centers according to the pairing combination; use a clustering criterion function as a loss function during the training of the feature extraction network, and update the network parameters of the feature extraction network through back propagation, wherein the clustering criterion function is constructed based on the network output results and the cluster centers.
7. The device according to claim 6, characterized in that It also includes a feature mapping network training module, which is used to: before extracting the pre-trained features corresponding to the unlabeled first audio data sample through the feature extraction network and after training the feature extraction network, Extracting pre-trained features corresponding to the labeled third audio data samples through the feature extraction network; The feature mapping network is trained using the pre-trained features corresponding to the third audio data sample as input to the feature mapping network and using the label of the third audio data sample as a training target, wherein the label of the third audio data sample represents the category of the phoneme of the third audio data sample.
8. The device according to claim 6, characterized in that The loss function is constructed as follows: Take the network output result of the i-th training sample as the i-th target sample y i , with c k Indicates the distance from the target sample y i The nearest cluster center, constructing the first relation: the i-th target sample y i and the distance target sample y i The nearest cluster center c k The square of the absolute value of the difference between the two is used to construct the second relationship: within the range of i from 1 to M, the first relationship corresponding to each value of i is calculated and summed, where M is the number of training samples in a single subset of the training samples; the second relationship is used as the loss function.
9. The device according to claim 6, characterized in that The feature extraction network training module includes a training sample construction submodule, which is used to: Performing a time-frequency transformation on the fourth audio data sample to obtain a frame-level original speech feature of the fourth audio data sample, wherein the frame-level original speech feature includes an amplitude spectrum vector and a phase spectrum vector of each frame of the fourth audio data sample; Based on the frame-level original speech features, context information of each frame of the audio signal in the amplitude spectrum vector and the phase spectrum vector is fused to construct a training sample for the feature extraction network.
10. The device according to claim 9, characterized in that The training sample construction submodule is also used for: For any t-th frame of speech original features in the frame-level speech original features, the amplitude spectrum vectors and phase spectrum vectors of all frames from the tD-th frame to the t+D-th frame are spliced according to preset rules to obtain the training sample of the feature extraction network corresponding to the t-th frame, where D represents the preset context window length.
11. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 5.
12. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.