Audio recognition method, model training method, device, equipment and storage medium

By determining the low-frequency triphone and the corresponding low-frequency text in speech recognition technology, and training the audio recognition model, the problem of low-frequency word recognition is solved, and the overall accuracy of speech recognition is improved.

CN115497460BActive Publication Date: 2025-05-09IFLYTEK CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211096150.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-08
Publication Date
2025-05-09
Estimated Expiration
2042-09-08

AI Technical Summary

Technical Problem

In speech recognition technology, low-frequency words have less weight in training data, resulting in lower recognition accuracy.

Method used

By determining the low-frequency triphones in the audio dataset, the low-frequency text containing the low-frequency triphones is determined from the preset corpus, and the audio recognition model is trained based on this text.

Benefits of technology

It improves the diversity and accuracy of training data, enables the audio recognition model to effectively recognize low-frequency words, and improves the accuracy of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497460B_ABST
    Figure CN115497460B_ABST
Patent Text Reader

Abstract

The present application provides an audio recognition method, a model training method, an apparatus, a device and a storage medium, and the specific implementation scheme is: determining low-frequency triphones in a first audio data set; based on the low-frequency triphones, determining low-frequency texts containing low-frequency triphones from a preset corpus; and training an audio recognition model based on the low-frequency texts. According to the technical scheme of the present application, the diversity and accuracy of the low-frequency data content in the training data can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of deep learning technology, and in particular to the field of speech recognition technology. Background Art

[0002] In recent years, with the rapid development of speech recognition technology, speech enhancement, speech recognition, voice question answering, information extraction and other related tasks have received more and more attention, but speech recognition technology relies heavily on training data, and low-frequency words account for a small weight in the collected training data. Therefore, the amount of training data is a key factor affecting low-frequency word recognition. Summary of the invention

[0003] According to a first aspect of an embodiment of the present application, a method for training an audio recognition model is provided, comprising:

[0004] identifying low frequency triphones in a first audio data set;

[0005] Based on the low-frequency triphones, determining low-frequency texts containing the low-frequency triphones from a preset corpus;

[0006] Train an audio recognition model based on low-frequency text.

[0007] According to a second aspect of an embodiment of the present application, there is provided an audio recognition method, comprising:

[0008] The audio data to be processed is recognized using an audio recognition model to obtain a recognition result in the audio data to be processed; wherein the audio recognition model is trained based on audio data synthesized from text containing low-frequency triphonemes.

[0009] According to a third aspect of an embodiment of the present application, a training device for an audio recognition model is provided, comprising:

[0010] A determination module, configured to determine low-frequency triphones in a first audio data set;

[0011] A search module, used for determining low-frequency texts containing low-frequency triphones from a preset corpus based on the low-frequency triphones;

[0012] The training module is used to train the audio recognition model based on low-frequency text.

[0013] According to a fourth aspect of an embodiment of the present application, there is provided an audio recognition device, including:

[0014] The audio processing module is used to use an audio recognition model to recognize the audio data to be processed and obtain a recognition result in the audio data to be processed; wherein the audio recognition model is trained based on audio data synthesized from text containing low-frequency triphonemes.

[0015] According to a fifth aspect of the embodiments of the present application, there is provided an electronic device, including:

[0016] at least one processor; and

[0017] a memory communicatively connected to at least one processor; wherein,

[0018] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor so that the at least one processor can execute any audio recognition model training method or audio recognition method in the embodiments of the present application.

[0019] According to the sixth aspect of the embodiments of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute any one of the audio recognition model training methods or audio recognition methods in the embodiments of the present application.

[0020] One embodiment of the above application has the following advantages or beneficial effects: by determining low-frequency triphones in the first audio data set, low-frequency texts containing low-frequency triphones are determined from a preset corpus, thereby increasing the diversity of training data. At the same time, since triphones take co-pronunciation into consideration, the accuracy of low-frequency audio data in the training data can be improved by using low-frequency triphones to determine low-frequency texts. The audio recognition model is trained based on low-frequency texts, thereby improving the training effect of the audio recognition model, so that the audio recognition model can effectively recognize low-frequency words. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0022] Figure 1 is a flowchart of a method for training an audio recognition model according to an embodiment of the present application;

[0023] Figure 2 is a flowchart of a method for training an audio recognition model according to another embodiment of the present application;

[0024] Figure 3 is a specific flow chart of step S130 in the method for training an audio recognition model according to another embodiment of the present application;

[0025] Figure 4 is a schematic diagram of a specific process of a method for training an audio recognition model according to another embodiment of the present application;

[0026] Figure 5 is a schematic diagram of an audio recognition method according to another embodiment of the present application;

[0027] Figure 6 is a block diagram of a training device for an audio recognition model according to an embodiment of the present application;

[0028] Figure 7 is a block diagram of an audio recognition device according to an embodiment of the present application;

[0029] Figure 8 It is a block diagram of an electronic device used to implement the training method of the audio recognition model and the audio recognition method of the embodiment of the present application. DETAILED DESCRIPTION

[0030] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0031] Exemplary Methods

[0032] Figure 1 FIG. 1 is a flow chart of a method for training an audio recognition model according to an embodiment of the present application. Figure 1 As shown, in an exemplary embodiment, the method may include:

[0033] S110, determining low-frequency triphones in a first audio data set;

[0034] S120, based on the low-frequency triphones, determining a low-frequency text containing the low-frequency triphones from a preset corpus;

[0035] S130. Train an audio recognition model based on low-frequency text.

[0036] In step S110, illustratively, the first audio data set is a set of audio data used to train the audio recognition model. The audio data in the first audio data set can be collected in advance in any scenario, and can include voice data in real scenarios, and can also include voice data synthesized by various AI voice platforms, which is not limited here. Among them, the voice data can be from a general voice training set open in the industry, or from voice recorded by a specific user, etc.

[0037] Optionally, the speech data can be represented by a single phoneme or a triphone. However, a single phoneme does not take into account co-pronunciation, that is, the context phoneme will affect the pronunciation of the current central phoneme and produce a co-variation, so the use of triphones makes the representation of audio data more accurate.

[0038] Optionally, the low-frequency triphone is used to represent the triphone whose occurrence frequency is lower than a preset threshold, and the preset threshold is set to 10, which can also be set according to actual conditions. Specifically, the audio data in the first audio data set can be represented by the triphone, and the occurrence frequency of the triphone is screened to determine the low-frequency triphone.

[0039] In step S120, illustratively, the preset corpus is an electronic library built by collecting naturally occurring continuous language texts or voice fragments according to certain linguistic principles. It is understandable that the preset corpus can be a corpus obtained in any field and in any way. Optionally, the preset corpus may include voice data and may also include text data. Optionally, when the corpus includes voice data, the voice data in the corpus may be converted into text data, the text data may be segmented to obtain whole words, and then the whole words may be converted into triphones to obtain a triphone sequence. Optionally, when the corpus includes text data, the text data may be directly segmented to obtain whole words, and then the whole words may be converted into triphones to obtain a triphone sequence.

[0040] Optionally, a phoneme is the smallest unit of speech divided according to the natural properties of speech, and is analyzed according to the pronunciation action in a syllable, and one action constitutes a phoneme. Therefore, the method of converting a whole word into a triphone may include: first using a monophone to represent the whole word to obtain a monophone sequence, and then copying the monophone into a triphone to convert the monophone sequence into a triphone sequence. For example, a monophone is represented as t, o, ng, and a triphone is represented as t-o+ng.

[0041] Another example is: Study hard and make progress every day in Chinese;

[0042] The pinyin is: hao3 hao3 xue2 xi2 tian1 tian1 xiang4 shang4;

[0043] The single phonemes are represented as: sil, h, ao3, h, ao3, x, ue2, x, i2, t, ian1, t, ian1, x, iang4, sh, ang4, sil;

[0044] The three phonemes are expressed as: sil sil-h+ao3 h-ao3+h ao3-h+ao3 h-ao3+x…sh-ang4+silsil; among them, the phoneme sil is used to express silence, and the position covered by the phoneme sil will be silent. In "sil-h+ao3", "sil" is on the left side of "-", and "h+ao3" is on the right side of "-", which means that the pronunciation is biased towards "h+ao3". "h" is on the left side of "+", and "ao3" is on the right side of "+", which means that the pronunciation is biased towards "h".

[0045] Then, the text data corresponding to the triphone sequence containing the low-frequency triphone extracted from the preset corpus is determined as the low-frequency text. Optionally, the relationship between the triphone sequence and the text data of the corpus can be used to extract the text data corresponding to the triphone sequence containing the low-frequency triphone from the corpus; or the triphone sequence containing the low-frequency triphone is converted into text data.

[0046] Exemplarily, since the same low-frequency triphone can correspond to multiple whole words, the low-frequency triphone can be first converted into multiple whole words, and then the low-frequency text containing the above whole words can be extracted from the preset corpus based on the multiple whole words. It should be noted that the low-frequency text finally determined by the above processing process can be a text of any length composed of any number of characters, for example, it can be a single word, a whole word, or a sentence, chapter, etc. Among them, the whole word can be a two-character word, or a four-character idiom, etc., which is not limited here.

[0047] In step S130, illustratively, the audio recognition model is a network model with audio recognition function, for example, it can be a neural network model (NN), an encoder-decoder (Encoder-Decoder), or a hidden Markov model (HMM), etc.

[0048] When the low-frequency text determined in step S120 is used to train the audio recognition model, any feasible training method may be used. For example, the audio corresponding to the low-frequency text may be retrieved from the audio library, and then the low-frequency text and its corresponding audio may be used to form a training sample, and then the audio recognition model may be trained for audio recognition using the training sample.

[0049] It can be understood that, compared with conventional audio data samples, for example, compared with the first audio data set mentioned above, the above low-frequency text contains more and more concentrated low-frequency triphones. Using low-frequency text to train the audio recognition model can increase the proportion of low-frequency audio in the training samples, so that the model can learn more low-frequency audio features.

[0050] Therefore, in the technical solution of the present application, low-frequency triphones are determined in the first audio data set, and low-frequency texts containing low-frequency triphones are determined from the preset corpus, which increases the diversity of the training data. At the same time, because the triphones take co-pronunciation into consideration, the accuracy of the low-frequency audio data in the training data can be improved by using low-frequency triphones to determine the low-frequency texts. The audio recognition model is trained based on low-frequency texts, which improves the training effect of the audio recognition model and enables the audio recognition model to effectively recognize low-frequency words.

[0051] In one embodiment, if Figure 2 As shown, before training the audio recognition model based on low-frequency text, it also includes:

[0052] S210, determining low-frequency words corresponding to low-frequency triphones in the low-frequency text;

[0053] S220: Based on the types of low-frequency words, adjust the number of low-frequency texts corresponding to the low-frequency words.

[0054] Exemplarily, based on the low-frequency triphones, at least one low-frequency text containing the low-frequency triphones is determined from a preset corpus, and the low-frequency words corresponding to the low-frequency triphones in each low-frequency text are determined.

[0055] Exemplarily, low-frequency words with the same text can be determined as low-frequency words of the same type. In this way, after obtaining the low-frequency words in at least one low-frequency text, the obtained low-frequency words can be aggregated to obtain at least one type of low-frequency words. Then, the number of low-frequency texts corresponding to each type of low-frequency words is adjusted to a preset range, wherein the preset range can be set according to actual conditions. This ensures the balance of the number of different low-frequency words in the training data (i.e., low-frequency texts), thereby improving the training effect of the audio recognition model.

[0056] In this embodiment, after multiple low-frequency texts are determined for the same low-frequency triphone in the preset corpus, since the same low-frequency triphone can correspond to multiple low-frequency words, and the number of low-frequency words in multiple low-frequency texts varies, the number of low-frequency texts corresponding to each low-frequency word is determined. It is determined whether the number of low-frequency texts corresponding to each low-frequency word exceeds the set value. If it does not exceed the set value, all low-frequency texts are retained; if it exceeds, a certain number of texts are randomly selected from the low-frequency texts, and generally the number of low-frequency texts corresponding to the set value is selected. Among them, the set value can be set according to actual needs.

[0057] For example, suppose a low-frequency triphone is A, and there are several low-frequency texts containing A retrieved, and the Chinese characters corresponding to A are A1, A2, A3, and A4, respectively. The low-frequency texts corresponding to A1, A2, A3, and A4 are screened. Assuming that the number of low-frequency texts to be screened is set to 30, if there are less than 30 low-frequency texts containing A1, all of them are retained. If there are 50 low-frequency texts containing A1, 30 are randomly selected and retained. The same screening operation is performed on A2, A3, and A4, so that the four Chinese characters can be evenly covered in the training data, thereby ensuring the balance of the training data.

[0058] For another example, suppose a low-frequency triphone is A, and there are several low-frequency texts containing A retrieved, and the Chinese characters corresponding to A are A1, A2, A3, and A4 respectively. The low-frequency texts corresponding to A1, A2, A3, and A4 are screened respectively. If there are 30 low-frequency texts containing A1, 50 low-frequency texts containing A2, 45 low-frequency texts containing A3, and 40 low-frequency texts containing A4, the number of texts corresponding to A1 can be used as the set value. In this way, 30 low-frequency texts are randomly selected from A2, A3, and A4 and retained, so that the four Chinese characters can be evenly covered in the training data, thereby ensuring the balance of the training data.

[0059] In one embodiment, if Figure 3 As shown, the audio recognition model is trained based on low-frequency text, including:

[0060] S310, synthesizing audio training data based on low-frequency text;

[0061] S320, determining a second audio data set based on the audio training data and the first audio data set;

[0062] S330: Train an audio recognition model based on the second audio data set.

[0063] Exemplarily, speech synthesis is performed on low-frequency text, and the synthesized audio can be used as audio training data.

[0064] In order to improve the synthetic audio training data to cover more speaking styles and increase the diversity and comprehensiveness of audio training data, low-frequency text can be synthesized through different types of speaker resources. Specifically, audio training data with different voice styles can be synthesized through the AI ​​voice platform, or speech synthesis can be performed through speakers with different voice characteristics.

[0065] Since the audio training data are all audio data synthesized based on low-frequency text, the proportion of low-frequency speech in the synthesized audio data is high. It should be noted that the proportion of low-frequency speech is low in real scenes. Then, if the audio recognition model is trained with synthesized audio data, the problem of inaccurate recognition of audio data in real scenes by the audio recognition model will occur. Therefore, the first audio data set is mixed with the audio training data to obtain a second audio data set. That is, the audio data collected in the real scene and the synthesized audio data are combined to obtain the second audio data set. Optionally, the first audio data set can be mixed with the audio training data in equal proportions, or they can be mixed in different proportions, which is not limited here. In this way, the audio data after the audio data collected in the real scene and the synthesized audio data are combined as training data, so that the training data is more authentic, and at the same time, the diversity of the training data is enriched, thereby improving the training effect of the audio recognition model, and then the audio recognition model can improve the accuracy of speech recognition.

[0066] In one embodiment, training an audio recognition model based on a second audio data set includes:

[0067] Generate a first loss function based on the correct sequence labeling and the observation sequence obtained by recognizing the audio data in the second audio data set by the audio recognition model;

[0068] Generate a target loss function based on a cross entropy loss function and a first loss function for identifying audio data in a second audio data set using an audio recognition model;

[0069] Train the audio recognition model using the objective loss function.

[0070] Exemplarily, since the essence of audio recognition is a sequence classification problem, the first loss function is generated based on the recognition result obtained by recognizing the audio data in the second audio data set based on the correct sequence labeling and the audio recognition model.

[0071] Optionally, sequence-discriminative training (SDT) may be introduced, that is, maximum mutual information, enhanced maximum mutual information, minimum phoneme error, and minimum Bayesian risk training criteria are introduced as the loss function (ie, the first loss function) of model training.

[0072] In this embodiment, the maximum mutual information is adopted, and the formula is as follows:

[0073]

[0074] Among them, m and w mare the observed sequence and correct sequence annotation of the mth audio sample, θ is the acoustic model parameter, and s m Yes m The corresponding state sequence, K is the acoustic scaling factor. In this formula, the numerator represents the possibility of the correct sequence, and the denominator is the sum of the possibilities of all possible word sequences.

[0075] For example, since machine learning (such as decoder-encoder) is generally used to generate an audio recognition model, a cross entropy loss function can be used when training the model to minimize the frame error rate. The cross entropy loss function and the first loss function are combined to generate a target loss function, which improves the accuracy of the sequence, thereby improving the training effect of the audio recognition model.

[0076] In this embodiment, the target loss function is obtained by adding the cross entropy loss function to the first loss function, and the formula is as follows:

[0077]

[0078] Among them, L ASR That is the cross entropy loss function, L SDT is the first loss function.

[0079] In one embodiment, determining low-frequency triphones in a first audio data set includes:

[0080] Determine a first phoneme set corresponding to the first audio data set using the triphones;

[0081] Based on the frequencies of the triphones in the first phoneme set, low-frequency triphones in the first phoneme set are determined.

[0082] In this embodiment, the audio data in the first audio data set is annotated, and then word segmentation is performed on it to obtain a text sequence, and then the text sequence is represented by a monophone to obtain a monophone sequence, and the monophone is copied into a triphone to convert the monophone sequence into a triphone sequence, thereby achieving the text sequence represented by the triphone, and then obtaining the first phoneme set based on the triphone sequence. According to the frequency of occurrence of each triphone in the first phoneme set, the triphone whose frequency of occurrence is lower than a preset threshold is defined as a low-frequency triphone, so as to more accurately determine the low-frequency words in the first audio data set.

[0083] In order to more thoroughly understand the features and technical contents of the embodiments of the present application, a specific application example is provided below for illustration. It should be understood that the following application example is only for reference and does not limit the specific implementation process.

[0084] In an application example, Figure 4 The training method of the audio recognition model may include:

[0085] The first step: segment the labeled data corresponding to the training audio in the original training data (ie, the first audio data set) to obtain a text string with whole word information.

[0086] Step 2: Convert the segmented text (i.e., the text string with whole word information) into single phonemes, i.e., phone strings, according to the phoneme dictionary.

[0087] Step 3: Since monophones do not take co-articulation into account, that is, the context phonemes will affect the pronunciation of the current central phoneme and produce co-articulation changes, the monophones must be converted into triphones, that is, the phone string is converted into a triphone string.

[0088] Step 4: Count the occurrence frequency of the triphones, set a threshold T, and triphones with a frequency lower than the threshold are defined as low-frequency triphones. T can be set according to actual conditions.

[0089] Step 5: Process the text in the corpus into triphone strings according to the first to third steps above, and then use low-frequency triphones to search the triphone strings in the corpus to see if they contain low-frequency triphones. If they do, convert the triphone strings into Chinese text and convert the low-frequency triphones into corresponding Chinese characters.

[0090] Step 6: After obtaining the text containing low-frequency words, it is found that the same low-frequency triphone corresponds to multiple Chinese characters, and the number of texts containing each Chinese character is different. In order to reduce the amount of text to be screened and ensure that the triphone covers every Chinese character as much as possible, the texts corresponding to these low-frequency words are screened. Determine whether the number of texts corresponding to each low-frequency word exceeds the set value. If it does not exceed the set value, all low-frequency texts are retained; if it exceeds, a certain number of texts are randomly selected from the low-frequency texts, generally according to the number corresponding to the set value. Among them, the set value can be set according to actual needs.

[0091] Step 7: Select seventy different speaker resources to synthesize audio from the filtered low-frequency word text, so that the speaking style of the synthesized audio is diverse and simulates the real audio data as much as possible.

[0092] Step 8: The original training data are all real audio data. In order not to affect the recognition effect of the recognition model on the real data, a part of the real audio data should be extracted and mixed with the low-frequency word audio data synthesized in the seventh step according to a certain ratio. After extracting the features, the model training is carried out. The training is based on the encoder-decoder framework, and the details will not be repeated. Traditional encoder-decoder usually adopts the cross entropy loss function (CE), which can independently process each frame of speech vector and minimize the frame error rate. However, speech recognition is essentially a sequence classification problem, so after two rounds of training based on the encoder-decoder framework, a sequence identification training method is introduced. The commonly used criteria are maximum mutual information, enhanced maximum mutual information, minimum phoneme error and minimum Bayesian risk training criteria. In this embodiment, the criterion used is the maximum mutual information criterion, which maximizes the mutual information between the observed sequence distribution and the word sequence distribution and reduces the sentence error rate. In frame-based speech recognition, the word error rate (WER) is generally used directly to evaluate the speech recognition accuracy, where the maximum mutual information criterion is:

[0093]

[0094] Among them, m and w m are the observed sequence and correct sequence annotation of the mth audio sample, θ is the acoustic model parameter, and s m Yes m The corresponding state sequence, K is the acoustic scaling factor. In this formula, the numerator represents the possibility of the correct sequence, and the denominator is the sum of the possibilities of all possible word sequences.

[0095] The overall loss function of the entire recognition framework is:

[0096]

[0097] Among them, L ASR is the cross entropy loss function, L SDT Losses calculated for the MMI criterion.

[0098] When the model loss satisfies the overall loss function, the trained model is output.

[0099] Figure 5 FIG. 1 is a flow chart of an audio recognition method according to an embodiment of the present application. Figure 5 As shown, in an exemplary embodiment, the method may include:

[0100] S510. Use an audio recognition model to recognize the audio data to be processed to obtain a recognition result in the audio data to be processed; wherein the audio recognition model is trained based on audio data synthesized from text containing low-frequency triphonemes.

[0101] In the technical solution of the present application, since the audio recognition model is trained based on audio data synthesized from text containing low-frequency triphonemes, the audio recognition model can accurately identify low-frequency words in the audio data to be processed.

[0102] In one embodiment, the audio recognition model is trained according to any audio recognition model training method in the application examples.

[0103] For example, low-frequency triphones can be determined in the first audio data set, and low-frequency texts containing low-frequency triphones can be determined from a preset corpus, thereby increasing the diversity of training data. At the same time, since triphones take co-pronunciation into consideration, the accuracy of low-frequency data content in the training data can be improved by using low-frequency triphones to determine low-frequency texts. The audio recognition model is trained based on low-frequency texts, which improves the training effect of the audio recognition model and enables the audio recognition model to effectively recognize low-frequency words.

[0104] In the technical solution of this application, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0105] Exemplary Devices

[0106] Correspondingly, such as Figure 6 As shown, the embodiment of the present application also provides a training device for an audio recognition model, including:

[0107] A determination module 610, configured to determine low frequency triphones in a first audio data set;

[0108] A search module 620 is used to determine a low-frequency text containing the low-frequency triphone from a preset corpus based on the low-frequency triphone;

[0109] The training module 630 is used to train an audio recognition model based on the low-frequency text.

[0110] In one embodiment, the apparatus, before the training module 630, further includes:

[0111] A first processing module, configured to determine in the low-frequency text the low-frequency words corresponding to the low-frequency triphonemes;

[0112] The second processing module is used to adjust the number of low-frequency texts corresponding to the low-frequency words based on the types of the low-frequency words.

[0113] In one embodiment, the training module 630 is further configured to:

[0114] synthesizing audio training data based on the low-frequency text;

[0115] Determine a second audio data set based on the audio training data and the first audio data set;

[0116] The audio recognition model is trained based on the second audio data set.

[0117] In one implementation, training the audio recognition model based on the second audio data set includes:

[0118] Generate a first loss function based on the correct sequence labeling and the observation sequence obtained by recognizing the audio data in the second audio data set by the audio recognition model;

[0119] Generate a target loss function based on a cross entropy loss function and the first loss function for identifying the audio data in the second audio data set by the audio recognition model;

[0120] The audio recognition model is trained using the target loss function.

[0121] In one implementation, the determination module 610 is further configured to:

[0122] Determine a first phoneme set corresponding to the first audio data set using the triphones;

[0123] Based on the frequency of the triphones in the first phoneme set, low-frequency triphones in the first phoneme set are determined.

[0124] Correspondingly, such as Figure 7 The embodiment of the present application also provides an audio recognition device, including:

[0125] The audio recognition module 710 is used to recognize the audio data to be processed using an audio recognition model to obtain a recognition result in the audio data to be processed; wherein the audio recognition model is trained based on audio data synthesized from text containing low-frequency triphonemes.

[0126] The device provided in this embodiment belongs to the same application concept as the method provided in the above embodiment of this application, can execute the method provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For the technical details not fully described in this embodiment, please refer to the specific processing content of the method provided in the above embodiment of this application, which will not be repeated here.

[0127] Exemplary Electronic Devices

[0128] Another embodiment of the present application also provides an electronic device, such as Figure 8 As shown, the device includes:

[0129] Memory 800 and processor 810;

[0130] The memory 800 is connected to the processor 810 and is used to store programs;

[0131] The processor 810 is used to implement the audio recognition model training method or audio recognition method disclosed in any of the above embodiments by running the program stored in the memory 800.

[0132] Specifically, the electronic device may further include: a bus, a communication interface 820 , an input device 830 and an output device 840 .

[0133] The processor 810, the memory 800, the communication interface 820, the input device 830 and the output device 840 are connected to each other via a bus.

[0134] A bus may include a pathway that transfers information between components of a computer system.

[0135] The processor 810 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the scheme of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0136] The processor 810 may include a main processor, and may also include a baseband chip, a modem, and the like.

[0137] The memory 800 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes a computer operation instruction. More specifically, the memory 800 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.

[0138] The input device 830 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.

[0139] Output device 840 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0140] The communication interface 820 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0141] The processor 810 executes the program stored in the memory 800 and calls other devices, which can be used to implement any audio recognition model training method or each step of the audio recognition method provided in the above embodiments of the present application.

[0142] Exemplary computer program products and storage media

[0143] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the training method of the audio recognition model or the steps in the audio recognition method according to various embodiments of the present application described in the above-mentioned "Exemplary Method" section of this specification.

[0144] The computer program product may be written in any combination of one or more programming languages ​​to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0145] In addition, an embodiment of the present application may also be a storage medium on which a computer program is stored. The computer program is executed by a processor to execute the training method of the audio recognition model or the steps in the audio recognition method according to various embodiments of the present application described in the above "Exemplary Method" section of this specification. Specifically, the following steps can be implemented: S110, determining low-frequency triphones in a first audio data set; S120, based on the low-frequency triphones, determining low-frequency text containing low-frequency triphones from a preset corpus; S130, training the audio recognition model based on the low-frequency text.

[0146] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0147] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0148] The steps in the methods of each embodiment of the present application can be adjusted in order, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.

[0149] The modules and sub-modules in the devices and terminals of the various embodiments of the present application can be combined, divided and deleted according to actual needs.

[0150] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0151] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0152] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.

[0153] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0154] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0155] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0156] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for training an audio recognition model, characterized in that: include: identifying low frequency triphones in a first audio data set; Based on the low-frequency triphones, determining low-frequency texts containing the low-frequency triphones from a preset corpus; Determining the low-frequency words corresponding to the low-frequency triphones in the low-frequency text; Based on the type of the low-frequency words, the number of low-frequency texts corresponding to the low-frequency words is adjusted; wherein low-frequency words of the same type refer to low-frequency words with the same text; An audio recognition model is trained based on the low-frequency text.

2. The method according to claim 1, characterized in that The step of training the audio recognition model based on the low-frequency text comprises: synthesizing audio training data based on the low-frequency text; Determine a second audio data set based on the audio training data and the first audio data set; The audio recognition model is trained based on the second audio data set.

3. The method according to claim 2, characterized in that The step of training the audio recognition model based on the second audio data set comprises: Generate a first loss function based on the correct sequence labeling and the observation sequence obtained by recognizing the audio data in the second audio data set by the audio recognition model; Generate a target loss function based on a cross entropy loss function and the first loss function for identifying the audio data in the second audio data set by the audio recognition model; The audio recognition model is trained using the target loss function.

4. The method according to any one of claims 1 to 3, characterized in that The step of determining low-frequency triphones in the first audio data set comprises: Determine a first phoneme set corresponding to the first audio data set using the triphones; Based on the frequency of the triphones in the first phoneme set, low-frequency triphones in the first phoneme set are determined.

5. An audio recognition method, characterized in that: include: The audio data to be processed is recognized using an audio recognition model to obtain a recognition result in the audio data to be processed; wherein the audio recognition model is trained based on audio data synthesized from text containing low-frequency triphones; and the audio recognition model is trained according to the training method for the audio recognition model as described in any one of claims 1 to 4.

6. A training device for an audio recognition model, characterized in that: include: A determination module, configured to determine low-frequency triphones in a first audio data set; A search module, configured to determine, based on the low-frequency triphones, a low-frequency text containing the low-frequency triphones from a preset corpus; A first processing module, configured to determine in the low-frequency text the low-frequency words corresponding to the low-frequency triphonemes; A second processing module is used to adjust the number of low-frequency texts corresponding to the low-frequency words based on the types of the low-frequency words; wherein low-frequency words of the same type refer to low-frequency words with the same text; A training module is used to train an audio recognition model based on the low-frequency text.

7. An audio recognition device, characterized in that: include: An audio processing module is used to use an audio recognition model to identify the audio data to be processed and obtain a recognition result in the audio data to be processed; wherein the audio recognition model is trained based on audio data synthesized from text containing low-frequency triphones; the audio recognition model is trained according to the training method of the audio recognition model as described in any one of claims 1-4.

8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the audio recognition model training method or the audio recognition method described in any one of claims 1 to 5.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the audio recognition model training method or the audio recognition method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Training method and system for end-to-end speech recognition model

    CN109346064A

  • Voice recognition method and device, electronic equipment and storage medium

    CN114067786A

  • Expansion method and speech recognition method for Cantonese audio

    CN114694655A

  • Generation method and device of speech recognition model, recognition method and device, medium and equipment

    CN114765025A