A method, apparatus, computer device, and medium for training a speech recognition model
By performing frame-based windowing and feature extraction on the speech signal, a data set is formed, and the initial model is iteratively trained, the problem of low confidence in the speech recognition task is solved, and high-quality speech recognition model training is achieved.
Patent Information
- Application Number
- CN202210139450.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-02-16
AI Technical Summary
In the prior art, the confidence of speech recognition tasks is low, mainly due to the small number of high-quality speech samples, resulting in a lack of sufficient data support during model training.
The voice signal segments are processed in frames and windows through multiple window lengths, and the signal segment set is obtained, and the audio features are extracted to form a data set. Then, the initial model is iteratively trained to obtain a high-quality speech recognition model.
The data enhancement of speech signals is realized, more signal segments and audio features used for model training are obtained, the recognition performance of speech recognition models is improved, and confidence is improved.
Smart Images

Figure CN114420108B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, computer device and medium for training a speech recognition model. Background Art
[0002] With the rapid development and wide application of artificial intelligence technology, speech recognition technology has gradually been applied to scenarios such as consumer business and financial service business. The tasks of speech recognition include tasks such as voiceprint recognition, semantic recognition or emotion recognition. An algorithm model can be trained through a data set to obtain an ideal recognition model. The training process of the recognition model requires a large number of labeled speech samples. However, there are few high-quality speech samples. Training with a small number of speech samples is not conducive to obtaining a recognition model with high confidence. Summary of the Invention
[0003] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method, device, computer device and medium for training a speech recognition model, which is used to solve the problem of low confidence in speech recognition tasks.
[0004] To achieve the above object and other related objects, the present invention provides a method for training a speech recognition model, including:
[0005] Collect the time-domain signal of the speech through multiple sampling points, perform channel conversion on the time-domain signal to obtain a single-channel time-domain signal, and perform frequency-domain conversion on the single-channel time-domain signal to obtain a frequency-domain signal;
[0006] Traverse the frequency-domain signal to determine the correspondence between the data frames and the spectrum of the frequency-domain signal, obtain the energy value of the data frames, determine the paragraph points of the frequency-domain signal through the energy value, and obtain signal paragraphs through the paragraph points and the data frames;
[0007] Perform frame windowing processing on the signal paragraphs with multiple window lengths to obtain a set of signal segments, respectively extract the audio features of each signal segment in the set of signal segments to obtain a data set, where the frame shift in the frame windowing processing is 1 / n of the window length, n>1;
[0008] Obtain an initial model, and perform iterative training on the data set and the initial model to obtain a speech recognition model.
[0009] Optionally, the step of traversing the frequency-domain signal to determine the correspondence between the data frames and the spectrum of the frequency-domain signal, obtaining the energy value of the data frames, determining the paragraph points of the frequency-domain signal through the energy value, and obtaining signal paragraphs through the paragraph points and the data frames includes:
[0010] Traverse the frequency-domain signal, determine the correspondence between the current data frame and the current spectrum, obtain the short-time energy entropy ratio of the current data frame, and determine whether the short-time energy entropy ratio is greater than a preset value. If the short-time energy entropy ratio is greater than the preset value, use the current data frame as the paragraph point of the frequency-domain signal, and output the data frames between two adjacent paragraph points as the signal paragraph.
[0011] Optionally, perform channel conversion on the time-domain signal to obtain a single-channel time-domain signal, including:
[0012] Determine the number of channels of the time-domain signal. When the number of channels is multi-channel, perform channel conversion through array enhancement or generalized sidelobe canceller enhancement to obtain the single-channel time-domain signal.
[0013] Optionally, the initial model includes a convolutional neural network, the convolutional neural network includes an input layer, an intermediate layer, and an output layer, the intermediate layer includes a convolutional layer and multiple weight layers, and the input end of one weight layer is connected to the output end of another weight layer.
[0014] Optionally, perform iterative training on the data set and the initial model through a loss function. The loss function includes: the loss of the samples and weight values of the i-th category, the loss of the weights of the labels of the i-th category and the labels of the i-th sample, and the loss of the samples of the i-th category and the k-th weight value.
[0015] Optionally, the steps of obtaining an initial model, performing iterative training on the data set and the initial model, and obtaining a speech recognition model further include:
[0016] Divide the data set into K sub-data sets, use one of the sub-data sets as the validation set, and use the remaining K - 1 sub-data sets as the training set for iterative training to obtain K - 1 training models;
[0017] Select the speech recognition model from the K - 1 training models according to the accuracy rate or recall rate.
[0018] Optionally, before the steps of obtaining an initial model, performing iterative training on the data set and the initial model, and obtaining a speech recognition model, it includes:
[0019] Perform re-frame windowing processing on the signal paragraph with multiple window lengths to obtain a signal segment set, extract the audio features of each signal segment in the signal segment set respectively to obtain a test set, where the frame shift in the re-frame windowing processing is 1 / p of the window length, p > 1 and p ≠ n.
[0020] The present invention provides a speech recognition model training device, including:
[0021] The acquisition module is used to acquire the time-domain signal of speech through multiple sampling points, perform channel conversion on the time-domain signal to obtain a single-channel time-domain signal, and perform frequency-domain conversion on the single-channel time-domain signal to obtain a frequency-domain signal;
[0022] The preprocessing module is used to traverse the frequency-domain signal, determine the correspondence between the data frame and the spectrum of the frequency-domain signal, obtain the energy value of the data frame, determine the paragraph points of the frequency-domain signal through the energy value, and obtain signal paragraphs through the paragraph points and the data frames;
[0023] The data module is used to perform frame addition and windowing processing on the signal paragraphs with multiple window lengths to obtain a set of signal segments, respectively extract the audio features of each signal segment in the set of signal segments to obtain a data set, where the frame shift in the frame addition and windowing processing is 1 / n of the window length, n>1;
[0024] The processing module is used to obtain an initial model, perform iterative training on the data set and the initial model to obtain a speech recognition model.
[0025] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the speech recognition model training method when executing the computer program.
[0026] The present invention provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the speech recognition model training method when executed by a processor.
[0027] As described above, the speech recognition model training method, device, computer device, and readable storage medium of the present invention have the following beneficial effects:
[0028] Perform frame addition and windowing processing on the signal paragraphs with multiple window lengths to obtain signal segments, realize data augmentation of high-quality speech signals, obtain more signal segments and audio features that can be used for model training, and obtain a data set based on the set of signal segments after data augmentation to meet the requirements of model training. Through iterative training of the model, using a high recognition performance as an index, obtain a preferred training model as the speech recognition model, and process the current or real-time speech signal through the trained speech recognition model to complete the speech recognition task. Description of the Drawings
[0029] Figure 1 It shows a schematic diagram of the application environment of the speech recognition model training method according to an embodiment of the present invention;
[0030] Figure 2Schematic flowchart of the speech recognition model training method according to an embodiment of the present invention;
[0031] Figure 3 Schematic structural diagram of a convolutional neural network according to another embodiment of the present invention;
[0032] Figure 4 Schematic structural diagram of an intermediate layer according to an embodiment of the present invention;
[0033] Figure 5 Schematic structural diagram of a speech recognition model training apparatus according to another embodiment of the present invention;
[0034] Figure 6 Schematic structural diagram of a computer device according to an embodiment of the present invention. Detailed implementation manners
[0035] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.
[0036] It should be noted that the diagrams provided in this embodiment only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex. The structures, ratios, sizes, etc. shown in the diagrams of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have technical substance. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed in the present invention. At the same time, the terms such as "upper", "lower", "left", "right", "middle", and "one" cited in this specification are only for the convenience of clear narration, rather than used to limit the scope under which the present invention can be implemented. The change or adjustment of their relative relationships, without substantial change in the technical content, should also be regarded as the scope under which the present invention can be implemented.
[0037] The speech recognition model training method provided by this solution can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. For example, voice recognition processing can be performed through the terminal 102, and its processing results are transmitted to the server 104 for data transfer. For example, voice recognition processing can be performed through the server 104, and its processing results are fed back to the terminal 102. For another example, the initial model can be iteratively trained in the server 104, the trained voice recognition model is stored locally, and the trained voice recognition model can also be configured in the terminal 102. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0038] In some business scenarios, for example, financial services, it is necessary to complete and support business processing through voice interaction. With the rapid development of voice recognition technology, processing tasks such as voiceprint recognition, speech recognition, and emotion recognition have also penetrated into various sub-fields. In the process of speech recognition, it is necessary to train the algorithm model for speech recognition through a data set. However, there are few high-quality speech samples, and it is necessary to perform enhancement processing on the speech samples to obtain more speech samples to meet the training requirements of the algorithm model.
[0039] As Figure 2 shown, the present invention provides a method for training a voice recognition model, including:
[0040] S1: Collect the time-domain signal of the voice through multiple sampling points, perform channel conversion on the time-domain signal to obtain a single-channel time-domain signal, and perform frequency-domain conversion on the single-channel time-domain signal to obtain a frequency-domain signal;
[0041] S2: Traverse the frequency-domain signal, determine the correspondence between the data frame of the frequency-domain signal and the spectrum, obtain the energy value of the data frame, determine the paragraph points of the frequency-domain signal through the energy value, and obtain signal paragraphs through the paragraph points and the data frames;
[0042] S3: Perform frame windowing processing on the signal paragraphs with multiple window lengths to obtain a set of signal segments, respectively extract the audio features of each signal segment in the set of signal segments to obtain a data set, where the frame shift in the frame windowing processing is 1 / n of the window length, n>1;
[0043] S4: Obtain an initial model, and perform iterative training on the data set and the initial model to obtain a voice recognition model.
[0044] In step S1, by way of example, voice acquisition can be completed through multiple sampling points. For example, it can be acquired through a recording device such as a microphone to obtain the time-domain signal of the voice. The time-domain signal can also be subjected to enhanced noise reduction processing to remove the interference of the noise signal on the voice and improve the quality of the time-domain signal. Since the multi-channel time-domain signal is not convenient for the extraction and recognition of voice features, the multi-channel time-domain signal can be subjected to channel conversion to convert the multi-channel time-domain signal into a single-channel time-domain signal. For the convenience of voice feature extraction and recognition, the single-channel time-domain signal is also subjected to frequency-domain conversion to obtain the frequency-domain signal.
[0045] In step S2, by way of example, the data frames and spectra of the frequency-domain signal are traversed to determine the correspondence between the data frames and spectra of the frequency-domain signal, and the energy values of each of the data frames are obtained. The paragraph points of the frequency-domain signal are determined through the energy values. For example, the data frames with high or low energy values are used as paragraph points, and the signal paragraphs are obtained through the paragraph points and the data frames. The audio features carried within the signal paragraphs are relatively continuous and relatively rich in audio features, so as to facilitate the extraction of more voice feature information.
[0046] In step S3, by way of example, the signal segment is subjected to frame division and windowing with multiple window lengths. For example, the provided window lengths include 200 ms (milliseconds), 100 ms, 50 ms, 30 ms, etc. A set of signal segments is obtained, and the audio features of each signal segment in the set of signal segments are extracted respectively to obtain a data set. Among them, the frame shift in the frame division and windowing process is 1 / n of the window length, where n > 1. For example, n includes 2, 3, 4, etc. The window length and the frame shift can be adjusted according to the number requirement of the voice samples to be enhanced. When more voice samples are needed, a shorter window length can be selected and the value of n can be set larger. When fewer voice samples are needed, a longer window length can be selected and the value of n can be set smaller. Voice is collected through a limited number of sampling points to achieve data augmentation of voice samples to meet the training requirements of the algorithm model. For example, four window lengths can be selected, namely: 200 ms, 100 ms, 50 ms, 30 ms, and the frame overlap is 1 / 2, that is, the corresponding frame overlap lengths (frame shifts) are: 100 ms, 50 ms, 25 ms, 15 ms respectively. The frame overlap can also be 1 / 4, that is, the corresponding frame overlap lengths (frame shifts) are: 50 ms, 25 ms, 12.5 ms, 7.5 ms respectively. The Mel Frequency Cepstrum Coefficient (MFCC) or chromaticity features can also be extracted as audio features. The extraction process of the Mel Frequency Cepstrum Coefficient includes: pre-emphasis, frame division, windowing, fast Fourier transform, obtaining the MEL spectrum through a band-pass filter, and obtaining cepstrum features. The extraction process of the chromaticity features includes: obtaining the frequency-domain signal through Fourier transform, and extracting the difference signals of the formant features and harmonic frequency features from the frequency-domain signal. The difference signals include first-order difference signals and / or second-order difference signals.
[0047] In step S4, by way of example, the obtained initial model can be a neural network algorithm or a support vector machine algorithm, or an algorithm model that has not been trained completely. The samples in the data set are input into the initial model for iteration, and a certain learning rate, model parameters, and number of training times are set to obtain the recognition result. The performance of the trained model is measured according to indicators such as the accuracy (Precision) and recall rate (Recall) of the recognition result, and the preferred trained model is obtained as the speech recognition model.
[0048] The signal segment is subjected to frame division and windowing with multiple window lengths to obtain signal segments, realizing data augmentation of high-quality voice signals, obtaining more signal segments and audio features that can be used for model training, and obtaining a data set based on the set of signal segments after data augmentation to meet the requirements of model training.
[0049] In some embodiments, the steps of traversing the frequency-domain signal to determine the correspondence between the data frames and the spectrum of the frequency-domain signal, obtaining the energy value of the data frames, determining the paragraph points of the frequency-domain signal based on the energy value, and obtaining the signal paragraphs based on the paragraph points and the data frames include:
[0050] Traverse the frequency-domain signal to determine the correspondence between the current data frame and the current spectrum, obtain the short-time energy-entropy ratio of the current data frame, use the short-time energy-entropy ratio as the determination index for the energy magnitude of the data frame, determine whether the short-time energy-entropy ratio is greater than a preset value. If the short-time energy-entropy ratio is greater than the preset value, move the current data frame to a preset speech frame buffer, use the current data frame as the paragraph point of the frequency-domain signal, and output the data frames between two adjacent paragraph points as the signal paragraphs. Speech segmentation can also be achieved through other algorithms. For example, through the response and pause of speech, the signal paragraphs can be obtained.
[0051] To reduce the processing difficulty of the time-domain signal, the multi-channel time-domain signal can be converted into a single-channel time-domain signal. In some embodiments, the steps of performing channel conversion on the time-domain signal include:
[0052] Determine the number of channels of the time-domain signal. When the number of channels is multi-channel, perform channel conversion through an Array Intensity Estimator (AIE) or a Generalized Sidelobe Canceller (GSC) to obtain the single-channel time-domain signal. The single-channel frequency-domain signal can also be input into a fixed beamformer for fixed beamforming to obtain a speech signal containing residual noise. The single-channel frequency-domain signal is input into a blocking matrix and, after being processed by an adaptive filter connected to the blocking matrix, a reference noise signal is obtained. The speech signal containing residual noise and the reference noise signal are input into an adaptive noise canceller for adaptive filtering to obtain a frequency-domain signal.
[0053] With the development of Artificial Intelligence (AI) technology, neural network algorithms are widely used in processing tasks such as classification and recognition. However, with the complication of processing tasks, general neural networks are not convenient for identifying and analyzing deep-level information. Especially in the process of speech recognition, when processing speech recognition tasks through neural network algorithms, problems such as model overfitting and difficulty in convergence are likely to occur, such as Figure 3 and Figure 4As shown, in some embodiments, the initial model includes a convolutional neural network, which includes an input layer 10, an intermediate layer 20, and an output layer 30. The intermediate layer includes a convolutional layer and multiple weight layers. Among them, the input end of a weight layer 21 is connected to the output end 22 of another weight layer. Since only a small number of hidden units in the intermediate layer of the neural network change their activation values for different inputs, while most hidden units have the same response to different inputs, the rank of the entire weight matrix is not high at this time. And as the number of network layers increases, the rank becomes even lower after successive multiplications. By connecting the input end and the output end in the weight layer of the intermediate layer, the weight distribution in backpropagation is improved, and the accuracy of the speech recognition model is increased.
[0054] In some embodiments, the dataset and the initial model are iteratively trained through a loss function, and the loss function includes: the loss of the samples and weight values of the i-th category, the loss of the weights of the labels of the i-th category and the labels of the i-th sample, and the loss of the samples of the i-th category and the k-th weight value. For example, the mathematical expression of the loss function is:
[0055]
[0056] where Loss is the loss function, N is the number of samples in the dataset, θ i is the angle between the samples of the i-th category and the weight values, y i is the label of the i-th sample, m is the cosine margin, s is the scale variable, e is the natural logarithm, k is the label sequence number of the sample, is the angle between the weights of the labels of the i-th category and the labels of the i-th sample, θ k,i is the angle between the samples of the i-th category and the k-th weight value, and i, k, and N are positive integers. Through this loss function, the initial model can be trained with higher accuracy in backpropagation, the weights of the neurons in the intermediate layer can be reasonably allocated, and the overfitting phenomenon during the training process can also be avoided.
[0057] To improve the accuracy and reliability of model training, in some embodiments, the steps of obtaining an initial model, iteratively training the dataset and the initial model, and obtaining a speech recognition model further include:
[0058] Dividing the dataset into K sub-datasets, using one subset of data as the validation set, and the remaining K - 1 sub-datasets as the training set, and performing iterative training to obtain K - 1 training models;
[0059] Select the speech recognition model from the K-1 training models according to the accuracy rate or recall rate. For example, when K is 10, divide the data set into ten equal parts, use one of the sub-data sets as the validation set, and use the other nine sub-data sets as the training set. Through training and comparison, obtain nine speech recognition models. Through the cross-validation strategy, the accuracy of the training model is better recognized and determined, avoiding the situation where the accuracy of the obtained speech recognition model is relatively low.
[0060] In order to verify the accuracy of the obtained speech recognition model, in some embodiments, before the step of obtaining the speech recognition model by iteratively training the data set and the initial model, it includes:
[0061] Perform re-frame windowing processing on the signal segment with multiple window lengths to obtain a set of signal segments, extract the audio features of each signal segment in the set of signal segments respectively to obtain a test set, where the frame shift in the re-frame windowing processing is 1 / p of the window length, p>1 and p≠n. During the process of obtaining the data set and the test set, different frame shifts are adopted to avoid a relatively high overlap between the data set and the test set. Different frame shifts can be used to meet the data acquisition requirements of the test set. For example, input the data set into the initial model for training to obtain a speech recognition model, and use the speech recognition model to recognize the speech samples in the test set to obtain the accuracy rate and recall rate of the speech recognition model, so as to test the performance of the recognition model.
[0062] It should be understood that although Figures 2-4 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figures 2-4 at least a part of the steps in
[0063] such as Figure 5 shown, the present invention provides a speech recognition model training device, including:
[0064] An acquisition module, configured to acquire the time-domain signal of the speech through multiple sampling points, perform channel conversion on the time-domain signal to obtain a single-channel time-domain signal, and perform frequency-domain conversion on the single-channel time-domain signal to obtain a frequency-domain signal;
[0065] A preprocessing module for traversing the frequency-domain signal, determining the correspondence between the data frames and the spectrum of the frequency-domain signal, obtaining the energy values of the data frames, determining the paragraph points of the frequency-domain signal based on the energy values, and obtaining signal paragraphs based on the paragraph points and the data frames;
[0066] A data module for performing frame windowing processing on the signal paragraphs with multiple window lengths to obtain a set of signal segments, respectively extracting the audio features of each signal segment in the set of signal segments, and obtaining a data set, where the frame shift in the frame windowing processing is 1 / n of the window length, and n>1;
[0067] A processing module for obtaining an initial model, iteratively training the data set and the initial model to obtain a speech recognition model. Performing frame windowing processing on the signal paragraphs with multiple window lengths to obtain signal segments, realizing data augmentation of high-quality speech signals, obtaining more signal segments and audio features that can be used for model training, and obtaining a data set based on the set of signal segments after data augmentation to meet the requirements of model training.
[0068] The window length and the frame shift can be adjusted according to the number requirement of the speech samples to be enhanced. When more speech samples are needed, a shorter window length can be selected and the value of n can be set larger. When fewer speech samples are needed, a longer window length can be selected and the value of n can be set smaller. Speech is collected through limited sampling points to realize data augmentation of speech samples to meet the training requirements of the algorithm model. For example, four window lengths can be selected, which are 200ms, 100ms, 50ms, and 30ms respectively, and the frame overlap is 1 / 2, that is, the corresponding frame overlap lengths (frame shifts) are 100ms, 50ms, 25ms, and 15ms respectively. The frame overlap can also be 1 / 4, that is, the corresponding frame overlap lengths (frame shifts) are 50ms, 25ms, 12.5ms, and 7.5ms respectively. The Mel Frequency Cepstrum Coefficient (MFCC) or chromaticity features can also be extracted as audio features. The extraction process of the Mel Frequency Cepstrum Coefficient includes pre-emphasis, framing, windowing, fast Fourier transform, obtaining the MEL spectrum through a band-pass filter, and obtaining cepstral features. The extraction process of the chromaticity features includes obtaining the frequency-domain signal through Fourier transform and extracting the difference signals of the formant features and harmonic frequency features from the frequency-domain signal. The difference signals include first-order difference signals and / or second-order difference signals.
[0069] In some embodiments, the steps of traversing the frequency-domain signal, determining the correspondence between the data frames and the spectrum of the frequency-domain signal, obtaining the energy values of the data frames, determining the paragraph points of the frequency-domain signal based on the energy values, and obtaining the signal paragraphs based on the paragraph points and the data frames include:
[0070] Traverse the frequency-domain signal, determine the correspondence between the current data frame and the current spectrum, obtain the short-time energy entropy ratio of the current data frame, determine whether the short-time energy entropy ratio is greater than a preset value. If the short-time energy entropy ratio is greater than the preset value, use the current data frame as the paragraph point of the frequency-domain signal, and output the data frames between two adjacent paragraph points as the signal paragraphs.
[0071] In some embodiments, the steps of performing channel conversion on the time-domain signal include:
[0072] Determine the number of channels of the time-domain signal. When the number of channels is multi-channel, perform channel conversion through array enhancement or generalized sidelobe canceller enhancement to obtain the single-channel time-domain signal.
[0073] In some embodiments, the initial model includes a convolutional neural network. The convolutional neural network includes an input layer, an intermediate layer, and an output layer. The intermediate layer includes a convolutional layer and multiple weight layers. The input end of one weight layer is connected to the output end of another weight layer. The speech recognition model can be configured in the processing module. When a processing task needs to be executed, the current speech information is processed by the processing module to obtain the recognition result.
[0074] In some embodiments, iterative training is performed on the data set and the initial model through a loss function. The loss function includes: the loss between the samples of the i-th category and the weight value, the loss between the label of the i-th category and the weight value of the label of the i-th sample, and the loss between the samples of the i-th category and the k-th weight value. For example, the mathematical expression of the loss function is:
[0075]
[0076] where Loss is the loss function, N is the number of samples in the data set, θ i is the angle between the samples of the i-th category and the weight value, y i is the label of the i-th sample, m is the cosine margin, s is the scale variable, e is the natural logarithm, k is the label sequence number, is the angle between the label of the i-th category and the weight value of the label of the i-th sample, θ k,i is the angle between the samples of the i-th category and the k-th weight value, and i, k, N are positive integers.
[0077] In some embodiments, the steps of obtaining an initial model, iteratively training the dataset and the initial model, and obtaining a speech recognition model further include:
[0078] Dividing the dataset into K sub-datasets, using one subset of data as a validation set, and the remaining K - 1 sub-datasets as training sets, and performing iterative training to obtain K - 1 training models;
[0079] Selecting the speech recognition model from the K - 1 training models according to the accuracy rate or recall rate.
[0080] In some embodiments, before the steps of obtaining an initial model, iteratively training the dataset and the initial model, and obtaining a speech recognition model, it includes:
[0081] Performing re-frame windowing processing on the signal segment with multiple window lengths to obtain a set of signal segments, respectively extracting the audio features of each signal segment in the set of signal segments to obtain a test set, where the frame shift in the re-frame windowing processing is 1 / p of the window length, p > 1 and p ≠ n. During the process of obtaining the dataset and the test set, different frame shifts are adopted to avoid a high overlap degree between the dataset and the test set. Different frame shifts can be used to meet the data acquisition requirements of the test set. For example, by inputting the dataset into the initial model for training to obtain a speech recognition model, and using the speech recognition model to recognize the speech samples in the test set to obtain the accuracy rate and recall rate of the speech recognition model, so as to test the performance of the recognition model.
[0082] According to the embodiments of the present disclosure, there is also provided an electronic device or a computer device capable of implementing the above method. Those skilled in the art can understand that the execution subject of the present invention includes but is not limited to a system, a method, or a program product. Therefore, the execution subject of the present invention can be specifically implemented in the following execution or implementation forms, that is: a hardware implementation manner, a software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module", "system", "device", or "apparatus" here.
[0083] Next, refer to Figure 6 to describe the electronic device 600 according to this embodiment of the present invention. Figure 6 The shown electronic device 600 is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present invention. As Figure 6As shown, the electronic device 600 is presented in the form of a general-purpose computer device. The components of the electronic device 600 may include, but are not limited to: at least one of the above-mentioned processing units 610, at least one of the above-mentioned storage units 620, and a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610). Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the "Speech Recognition Model Training Method" section of the present specification. The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 621 and / or a cache storage unit 622, and may further include a read-only storage unit (ROM) 623. The storage unit 620 may also include a program / utilities 624 having a set (at least one) of program modules 625. Such program modules 625 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The bus 630 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures. The electronic device 600 may also communicate with one or more external devices 800 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or may communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 650, such as communicating with a display unit 640. And, the electronic device 600 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 660. As shown in the figure, the network adapter 660 communicates with other modules of the electronic device 600 through the bus 630. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID (Redundant Arrays of Independent Disks / disk arrays) systems, tape drives, and data backup storage systems, etc.
[0084] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computer device (such as a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0085] According to an embodiment of the present disclosure, there is also provided a computer-readable storage medium storing computer-readable instructions, which when executed by a computer, cause the computer to execute the speech recognition model training method described above in this specification.
[0086] In some possible embodiments, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the "Exemplary Method" section above in this specification.
[0087] In the present invention, the readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above. The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0088] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for the purpose of limitation. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed, for example, synchronously or asynchronously in multiple modules. It should be understood that the present invention is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A method for training a speech recognition model, characterized in that, Including: Collect the time-domain signal of speech through multiple sampling points, perform channel conversion on the time-domain signal to obtain a single-channel time-domain signal, and perform frequency-domain conversion on the single-channel time-domain signal to obtain a frequency-domain signal; Traverse the frequency-domain signal, determine the correspondence between the data frames and the spectrum of the frequency-domain signal, obtain the energy value of the data frames, determine the paragraph points of the frequency-domain signal through the energy value, and obtain signal paragraphs through the paragraph points and the data frames; Perform frame windowing processing on the signal paragraphs with multiple window lengths to obtain a set of signal segments, and extract the audio features of each signal segment in the set of signal segments respectively to obtain a data set, where the frame shift in the frame windowing processing is 1 / n of the window length, and n>1; Obtain an initial model, perform iterative training on the data set and the initial model to obtain a speech recognition model; before the step of obtaining the speech recognition model, it includes: performing frame windowing processing on the signal paragraphs again with multiple window lengths to obtain a set of signal segments, and extracting the audio features of each signal segment in the set of signal segments respectively to obtain a test set, where the frame shift in the frame windowing processing again is 1 / p of the window length, and p>1 and p≠n.
2. The method for training a speech recognition model according to claim 1, characterized in that, The step of traversing the frequency-domain signal, determining the correspondence between the data frames and the spectrum of the frequency-domain signal, obtaining the energy value of the data frames, determining the paragraph points of the frequency-domain signal through the energy value, and obtaining signal paragraphs through the paragraph points and the data frames includes: Traverse the frequency-domain signal, determine the correspondence between the current data frame and the current spectrum, obtain the short-time energy entropy ratio of the current data frame, and judge whether the short-time energy entropy ratio is greater than a preset value. If the short-time energy entropy ratio is greater than the preset value, then use the current data frame as the paragraph point of the frequency-domain signal, and output the data frames between two adjacent paragraph points as the signal paragraph.
3. The method for training a speech recognition model according to claim 1, characterized in that, Performing channel conversion on the time-domain signal to obtain a single-channel time-domain signal includes: Judge the number of channels of the time-domain signal. When the number of channels is multi-channel, perform channel conversion through array enhancement or generalized sidelobe canceller enhancement to obtain the single-channel time-domain signal.
4. The method for training a speech recognition model according to claim 1, characterized in that, The initial model includes a convolutional neural network, and the convolutional neural network includes an input layer, an intermediate layer, and an output layer. The intermediate layer includes a convolutional layer and multiple weight layers, where the input end of one weight layer is connected to the output end of another weight layer.
5. The method for training a speech recognition model according to claim 4, characterized in that, Perform iterative training on the data set and the initial model through a loss function, and the loss function includes: the loss of the samples and weight values of the i-th category, the loss of the weights of the labels of the i-th category and the labels of the i-th sample, and the loss of the samples of the i-th category and the k-th weight value.
6. The method for training a speech recognition model according to claim 4, characterized in that, The step of obtaining an initial model, performing iterative training on the data set and the initial model to obtain a speech recognition model further includes: Divide the data set into K sub-data sets, use one sub-data set as a validation set, and use the remaining K-1 sub-data sets as training sets for iterative training to obtain K-1 training models; Select the speech recognition model from the K-1 training models according to the accuracy rate or the recall rate.
7. A speech recognition model training device, characterized in that, It includes: An acquisition module, configured to acquire the time-domain signal of the speech through a plurality of sampling points, perform channel conversion on the time-domain signal to obtain a single-channel time-domain signal, and perform frequency-domain conversion on the single-channel time-domain signal to obtain a frequency-domain signal; A preprocessing module, configured to traverse the frequency-domain signal, determine the correspondence between the data frame and the spectrum of the frequency-domain signal, obtain the energy value of the data frame, determine the paragraph points of the frequency-domain signal through the energy value, and obtain the signal paragraphs through the paragraph points and the data frames; A data module, configured to perform frame windowing processing on the signal paragraphs with multiple window lengths to obtain a set of signal segments, respectively extract the audio features of each signal segment in the set of signal segments to obtain a data set, wherein the frame shift in the frame windowing processing is 1 / n of the window length, and n>1; A processing module, configured to obtain an initial model, perform iterative training on the data set and the initial model to obtain a speech recognition model; before the step of obtaining the speech recognition model, it includes: performing frame windowing processing on the signal paragraphs again with multiple window lengths to obtain a set of signal segments, respectively extract the audio features of each signal segment in the set of signal segments to obtain a test set, wherein the frame shift in the re-frame windowing processing is 1 / p of the window length, and p>1 and p≠n.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the speech recognition model training method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the speech recognition model training method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Audio beat detection method and device and storage medium
CN109256147A
Air traffic control speech recognition method and device for small number of labeled samples
CN111785257A
Voice endpoint detection method, device and equipment and computer readable storage medium
CN112102851A