An expansion method for Cantonese audio and a speech recognition method
In the speech recognition model training, the first average word frequency and expansion weight are calculated based on the phoneme word frequency, and the audio with low phoneme coverage is selected for expansion, the problem of unbalanced training data of the small language speech recognition model is solved and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202210314205.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-03-28
AI Technical Summary
When training a speech recognition model for identifying a small language, the audio used for training is collected for a long time and uneven distribution, which affects the recognition accuracy.
By obtaining the phoneme text of the sample audio, counting the phoneme word frequency, calculating the first average word frequency, determining the expansion weight, selecting audio with low phoneme coverage for expansion, enriching the training data.
It improves the phoneme coverage balance of the training data and improves the recognition accuracy of the speech recognition model.
Smart Images

Figure CN114694655B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing, and more specifically, to a method for expanding Cantonese audio and a speech recognition method. Background Art
[0002] With the rapid development of computer technology, artificial intelligence has become more and more common in people's daily lives. Speech recognition is to convert voice signals into corresponding texts, which is one of the very important ways to achieve human-computer interaction. In recent years, with the great improvement of speech recognition accuracy and the continuous popularization of intelligent devices, voice input has become one of the main ways of text input, and voice interaction has also been applied in more and more scenarios.
[0003] Under different languages, there are differences in pronunciation rules. If you want to train an audio recognition model for a certain language, you need to pre-collect a number of audio of that language to train the speech recognition model. In practice, if you need to train a speech recognition model for a small language, the time to collect the audio for training this speech recognition model is relatively long, and moreover, the collected training audio may be unevenly distributed in terms of pronunciation, thereby affecting the speech recognition accuracy of the speech recognition model. Summary of the Invention
[0004] In view of the above problems, embodiments of the present application propose a method, device, electronic device and storage medium for expanding Cantonese audio to improve the above problems.
[0005] In a first aspect, an embodiment of the present application provides a method for expanding Cantonese audio, the method including: obtaining the phoneme text corresponding to each sample audio in the sample audio set; the phoneme text includes at least one phoneme; the sample audio is Cantonese audio; according to the phoneme text corresponding to each sample audio in the sample audio set, counting the phoneme word frequency of each phoneme; for each sample audio, calculating the average value of the phoneme word frequencies corresponding to the phonemes in the phoneme text corresponding to the sample audio to obtain the first average word frequency corresponding to the sample audio; according to the first average word frequency corresponding to the sample audio, determining the expansion weight corresponding to the sample audio, where the expansion weight is negatively correlated with the first average word frequency; according to the expansion weights corresponding to each sample audio, determining the target sample audio to be expanded in the sample audio set; performing audio expansion on the target sample audio to obtain an expanded audio; the expanded audio and the sample audio in the sample audio set are used to train a speech recognition model.
[0006] Second aspect, embodiments of the present application provide an expansion device for Cantonese audio, including: an acquisition module, configured to acquire phoneme texts corresponding to each sample audio in a sample audio set; the phoneme text includes at least one phoneme; a phoneme word frequency statistics module, configured to statistically calculate the phoneme word frequency of each phoneme according to the phoneme texts corresponding to each sample audio in the sample audio set; a first average word frequency determination module, configured to, for each sample audio, calculate the average value of the phoneme word frequencies corresponding to the phonemes in the phoneme text corresponding to the sample audio, to obtain the first average word frequency corresponding to the sample audio; an expansion weight determination module, configured to determine the expansion weight corresponding to the sample audio according to the first average word frequency corresponding to the sample audio, wherein the expansion weight has a negative correlation with the first average word frequency; a target sample audio determination module, configured to determine a target sample audio to be expanded in the sample audio set according to the expansion weights corresponding to each sample audio; an expansion audio determination module, configured to perform audio expansion on the target sample audio to obtain an expansion audio; the expansion audio and the sample audios in the sample audio set are used to train a speech recognition model.
[0007] In some embodiments, the target sample audio determination module includes: a sample audio subset determination unit, configured to classify the sample audios in the sample audio set according to a preset plurality of expansion weight intervals and the expansion weights corresponding to each sample audio, to obtain sample audio subsets corresponding to each expansion weight interval; an expansion quantity determination unit, configured to determine the expansion quantity corresponding to each expansion weight interval, the expansion quantity having a positive correlation with the expansion weight; a sample audio selection unit, configured to, for each expansion weight interval, select sample audios from the sample audio subset corresponding to the expansion weight interval according to the expansion quantity corresponding to the expansion weight interval; a first target sample audio determination unit, configured to use the selected sample audios as the target sample audio to be expanded.
[0008] In some embodiments, the expansion quantity determination unit includes: an expansion ratio acquisition subunit, configured to acquire the expansion ratio set for each expansion weight interval; an expansion quantity determination subunit, configured to determine the expansion quantity corresponding to each expansion weight interval according to the expansion ratio set for each expansion weight interval and the total number of sample audios in the sample audio set.
[0009] In some embodiments, the expansion audio determination module includes a processing unit, configured to perform specified processing on the target sample audio to obtain the corresponding expansion audio, the specified processing including at least one of noise addition processing, speech rate acceleration processing, and reverberation addition processing.
[0010] In some embodiments, the extended weight determination module includes: a difference calculation unit configured to calculate the difference between the first average word frequency corresponding to the sample audio and the word frequency threshold; a numerical interval determination unit configured to determine the numerical interval to which the difference belongs; and a first extended weight determination unit configured to use, as the extended weight corresponding to the sample audio, the extended weight corresponding to the numerical interval to which the difference belongs according to the correspondence between the numerical interval and the extended weight.
[0011] In some other embodiments, the extended weight determination module includes: a target word frequency interval determination unit configured to determine the target word frequency interval to which the first average word frequency corresponding to the sample audio belongs; a second extended weight determination unit configured to determine the extended weight corresponding to the target word frequency interval according to the correspondence between the word frequency interval and the extended weight; and a third extended weight determination unit configured to use, as the extended weight corresponding to the sample audio, the extended weight corresponding to the target word frequency interval.
[0012] In some other embodiments, the target sample audio determination module includes: a sample audio determination unit configured to determine, according to the extended weights corresponding to the sample audios, the sample audios in the sample audio set whose extended weights are greater than the weight threshold; and a second target sample audio determination unit configured to use the sample audios whose extended weights are greater than the weight threshold as the target sample audios to be extended.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor; and a memory storing computer-readable instructions thereon, where when the computer-readable instructions are executed by the processor, the extended method for Cantonese audio as described above is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-readable instructions thereon, where when the computer-readable instructions are executed by a processor, the extended method for Cantonese audio as described above is implemented.
[0015] In the solution of the present application, first, according to the phoneme texts corresponding to the sample audios in the sample audio set, then, according to the phoneme texts corresponding to the sample audios, the phoneme word frequencies of each phoneme are statistically calculated, based on the statistically calculated phoneme word frequencies of each phoneme, the average value of the phoneme word frequencies of the corresponding phonemes in each sample audio is calculated to obtain the first average word frequency corresponding to each sample audio, and then, according to the first average word frequency of each sample audio, the extended weight of each sample audio is determined. Thus, according to the extended weights of the sample audios, the target sample audios to be extended are determined in the sample audio set, and finally, the target sample audios are extended to obtain extended audios, and the extended audios and the sample audios in the sample audio set are used to train the speech recognition model.
[0016] In the solution of the present application, by determining the phoneme word frequencies of each phoneme in each sample audio in the sample audio set, then calculating the corresponding first average word frequency of each sample audio, and determining the expansion weight of the sample audio based on the first average word frequency. Since the first average word frequency can reflect the overall situation of phoneme coverage in the phoneme text corresponding to each sample audio, and there is a negative correlation between the first average word frequency and the expansion weight, it is possible to determine a higher expansion weight for the sample audio with a lower phoneme coverage rate. Furthermore, the probability that the sample audio with a lower phoneme coverage rate is selected for audio expansion is higher. In this way, since the probability that the sample audio with a lower phoneme coverage rate is expanded is higher, compared with only using the sample audio set as the training data of the speech recognition model, both the expanded audio obtained by expansion and the sample audio in the sample audio set are used as the training data of the speech recognition model, which improves the balance of phoneme coverage in the training data. Therefore, training the speech recognition model based on the expanded audio and the sample audio in the sample audio set can ensure the recognition accuracy of the trained speech recognition model, and solve the problem of low recognition accuracy caused by the uneven distribution of the sample audio used for training the speech recognition model in terms of pronunciation.
[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0019] Figure 1 It is a schematic flowchart of a method for expanding Cantonese audio according to an embodiment of the present application.
[0020] Figure 2 It is a schematic flowchart of the specific steps of step 140 according to an embodiment of the present application.
[0021] Figure 3 It is a schematic flowchart of the specific steps of step 140 according to another embodiment of the present application.
[0022] Figure 4 It is a schematic flowchart of the specific steps of step 150 according to an embodiment of the present application.
[0023] Figure 5 It is a schematic flowchart of the specific steps of step 420 according to an embodiment of the present application.
[0024] Figure 6 It is a schematic flowchart of specific steps before speech recognition using a speech recognition model according to an embodiment of the present application.
[0025] Figure 7 It is a block diagram of an expansion device for Cantonese audio according to an embodiment of the present application.
[0026] Figure 8 It is a hardware structure diagram of an electronic device according to an embodiment of the present application.
[0027] Through the above-mentioned drawings, specific embodiments of the present invention have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the inventive concept in any way, but to illustrate the concept of the present invention to those skilled in the art through specific embodiments. Detailed implementation manners
[0028] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0029] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0030] Figure 1 It is a schematic flowchart of an expansion method for Cantonese audio according to an embodiment of the present application. This method can be executed by an electronic device with computing and processing capabilities, such as terminal devices like desktop computers and laptop computers. This method can also be executed interactively by a processing system including a server and a terminal. As Figure 1 shown, this method includes the following steps:
[0031] Step 110, obtaining the phoneme text corresponding to each sample audio in the sample audio set; the phoneme text includes at least one phoneme; the sample audio is Cantonese audio.
[0032] In other application scenarios, the sample audios in the sample audio set may be audios of the same language, for example, all Cantonese audios. In addition, the sample audios in the sample audio set may also be Tibetan audios, Northeastern dialect audios, Sichuan dialect audios, etc., which are not specifically limited here.
[0033] Phone is the smallest unit of speech divided according to the natural properties of speech. Phonemes are analyzed based on the pronunciation action in a syllable. One action constitutes one phoneme. For example, in Mandarin, 啊(ā) has one phoneme ā; 爱(ài) has two phonemes, à and i; 代(dài) has three phonemes, d, à and i.
[0034] In different languages, the pronunciation corresponding to the same text is different. In the present application, the language to which the sample audio belongs can be determined in advance, and then the sample audio is converted into a phoneme text in the language to which the sample audio belongs. In other words, the phoneme text corresponding to the sample audio reflects the pronunciation of the text content corresponding to the sample audio in the language to which the sample audio belongs.
[0035] In some embodiments, the phonemes corresponding to the sample audio can be determined based on the pronunciation of the sample audio, and then the text form of each phoneme can be determined. For example, dai (dài) has three phonemes, and the corresponding phoneme texts are d, à and i.
[0036] In some embodiments, each sample audio corresponds to an identifier, which is used to distinguish each sample audio. It can be understood that the text content corresponding to each sample audio also has the same identifier as the sample audio. When the sample audio is used to train the speech recognition model in the subsequent process, the sample audio and the text content corresponding to each sample audio can be found quickly and accurately, and the identifiers of the sample audios are different, and the identifier of each sample audio is unique.
[0037] Step 120 , counting the phoneme frequency of each phoneme according to the phoneme text corresponding to each sample audio in the sample audio set.
[0038] The phoneme word frequency of each phoneme is the number of times each phoneme appears in the phoneme text corresponding to all the sample audios in the sample audio set. In other embodiments, the phoneme word frequency of each phoneme can also be the frequency of occurrence of the phoneme in the phoneme text corresponding to all the sample audios in the sample audio set, that is, the phoneme word frequency of the phoneme = K / N, where K is the number of times a phoneme appears in the phoneme text corresponding to all the sample audios in the sample audio set, and N is the sum of the number of times all phonemes appear in the phoneme text corresponding to all the sample audios in the sample audio set.
[0039] Step 130: For each sample audio, average the phoneme word frequencies corresponding to the phonemes in the phoneme text corresponding to the sample audio to obtain a first average word frequency corresponding to the sample audio.
[0040] In some embodiments, according to the phoneme word frequencies of each phoneme counted in step 120, the phoneme word frequency corresponding to each phoneme in each sample audio can be obtained. Then, the average value of the phoneme word frequencies of all phonemes in the phoneme text corresponding to the sample audio is calculated. For example, if the text content of an audio is "personnel" in Cantonese, and its corresponding phoneme text is j, an3 (3 represents the tone, an3 represents the third tone of an), s, and i4 (this phoneme represents the fourth tone of i), according to the phoneme word frequencies of each phoneme counted in step 120, the specific phoneme word frequencies of each phoneme in this sample audio are: j(10), an3(23), s(12), and i4(75). Then, the average value of the phoneme word frequencies of this sample audio is calculated as: (10 + 23 + 12 + 75) / 4 = 30, that is, the first average word frequency corresponding to this sample audio is 30.
[0041] Step 140, determine the expansion weight corresponding to the sample audio according to the first average word frequency corresponding to the sample audio, where the expansion weight is negatively correlated with the first average word frequency.
[0042] The expansion weight is used to reflect the magnitude of the probability that the sample audio needs to be audio-expanded. In this application, the expansion weight is negatively correlated with the first average word frequency, that is, if the first average word frequency corresponding to a sample audio is larger, the expansion weight corresponding to the sample audio is smaller.
[0043] For a phoneme, if the phoneme word frequency of a phoneme is higher, it indicates that the frequency of this phoneme appearing in the phoneme text set corresponding to the sample audio set is higher. Since the first average word frequency corresponding to a sample audio is obtained by calculating the average value of the phoneme word frequencies of all phonemes in the phoneme text corresponding to the sample audio, the first average word frequency can reflect the overall situation of phoneme coverage in the phoneme text corresponding to the sample audio, that is, the higher the first average word frequency, the higher the phoneme coverage rate in the phoneme text corresponding to the sample audio.
[0044] For a speech recognition model, in order to ensure the recognition accuracy of the speech recognition model, it is necessary to ensure the balance of the training data, for example, the audio covering different phonemes is basically balanced. If the coverage rate of a certain phoneme or phoneme text is low, it indicates that it is necessary to increase the coverage rate of the audio corresponding to this phoneme or phoneme text. Based on this, in this application, since the first average word frequency corresponding to the sample audio reflects the phoneme coverage of the phoneme text corresponding to the sample audio, the smaller the first average word frequency, the more it indicates that it is necessary to increase the coverage rate of the phoneme text corresponding to the sample audio. Therefore, for the sample audio with a smaller first average word frequency, a higher expansion weight is given, and then in the subsequent process, the probability of expanding this sample audio is higher.
[0045] In some embodiments, such as Figure 2As shown, step 140 includes:
[0046] Step 210, calculating the difference between the first average word frequency corresponding to the sample audio and the word frequency threshold.
[0047] In some embodiments, the average phoneme word frequency corresponding to the phonemes in the phoneme text corresponding to all the sample audios in the sample audio set can be calculated to obtain a second average word frequency, and the obtained second average word frequency can be used as the word frequency threshold. In this embodiment, the second average word frequency is used as the word frequency threshold, and then the difference between the first average word frequency and the word frequency threshold is calculated. In other embodiments, the word frequency threshold can also be set by the user, and no specific limitation is provided here.
[0048] Step 220, determining the numerical interval to which the difference belongs.
[0049] In some embodiments, multiple numerical intervals can be pre-divided, and the expansion weights corresponding to each numerical interval can be set. For example, the numerical intervals can be divided in units of 10. Specifically, 0 - 9 can be the first numerical interval, 10 - 19 can be the second numerical interval, and so on. If the difference is 15, it is determined that the difference belongs to the second numerical interval. The specific number of numerical intervals divided and the values included in each numerical interval can be set according to actual needs, and no specific limitation is provided here.
[0050] Step 230, according to the corresponding relationship between the numerical interval and the expansion weight, using the expansion weight corresponding to the numerical interval to which the difference belongs as the expansion weight corresponding to the sample audio.
[0051] As described above, since the expansion weights corresponding to each numerical interval are preset, after determining the numerical interval to which the difference belongs, the expansion weight corresponding to the numerical interval to which the difference belongs can be correspondingly determined to obtain the expansion weight corresponding to the sample audio.
[0052] In some embodiments, the corresponding relationship between the numerical interval and the expansion weight can be that one numerical interval corresponds to one expansion weight. In other embodiments, the corresponding relationship between the numerical interval and the expansion weight can also be that multiple numerical intervals correspond to one expansion weight. The corresponding relationship between the numerical interval and the expansion weight can be set according to actual needs, and no specific limitation is provided here.
[0053] In other embodiments, as Figure 3 shown, step 140 includes:
[0054] Step 310, determining the target word frequency interval to which the first average word frequency corresponding to the sample audio belongs.
[0055] The target word frequency interval refers to the word frequency interval to which the first average word frequency belongs.
[0056] In some embodiments, the maximum and minimum values of the phoneme word frequencies can be determined according to the phoneme word frequencies of each phoneme counted in step 120. Based on the maximum and minimum values, multiple word frequency intervals are divided. Specifically, taking the minimum and maximum values as the starting and ending points, the range defined by the maximum and minimum values is divided into multiple word frequency intervals. Optionally, 10 word frequency intervals can be divided according to the maximum and minimum values of the phoneme word frequencies. Based on the first average word frequency corresponding to the sample audio calculated in step 130, the target word frequency interval to which the first average word frequency belongs is determined. For example, if the maximum value of the phoneme word frequency is 80 and the minimum value is 21, 21 - 25 can be the first word frequency interval, 26 - 30 can be the second word frequency interval, and so on, to obtain 10 word frequency intervals. If the first average word frequency calculated in step 130 is 33, the target word frequency interval to which the first average word frequency belongs can be determined as the third word frequency interval (the word frequency interval with a phoneme word frequency of 31 - 35). In some other embodiments, the user can customize the number of word frequency intervals and the corresponding ranges, which are not specifically limited herein.
[0057] Step 320: Determine the expansion weight corresponding to the target word frequency interval according to the correspondence between the word frequency interval and the expansion weight.
[0058] In this embodiment, multiple word frequency intervals can be preset, and the expansion weights corresponding to each word frequency interval can be set. In this way, when determining the target word frequency interval to which the first average word frequency corresponding to the sample audio belongs, the expansion weight corresponding to the target word frequency interval can be determined accordingly.
[0059] In some embodiments, the correspondence between the word frequency interval and the expansion weight can be that one word frequency interval corresponds to one expansion weight. In some other embodiments, the correspondence between the word frequency interval and the expansion weight can also be that multiple word frequency intervals correspond to one expansion weight. The correspondence between the word frequency interval and the expansion weight can be set according to actual needs, which is not specifically limited herein.
[0060] Step 330: Use the expansion weight corresponding to the target word frequency interval as the expansion weight corresponding to the sample audio.
[0061] Please continue to refer to Figure 1 , step 150: Determine the target sample audio to be expanded in the sample audio set according to the expansion weights corresponding to each sample audio.
[0062] As described above, the higher the expansion weight corresponding to a sample audio, the higher the probability that the sample audio is used for expansion. Therefore, after determining the expansion weights corresponding to each sample audio through the above steps 110 - 140, the sample audio to be expanded can be selected from the sample audio set based on the expansion weights.
[0063] In some embodiments, sample audio with an extended weight exceeding a weight threshold can be selected from a sample audio set, and the selected sample audio with an extended weight exceeding the weight threshold is used as the target sample audio to be extended.
[0064] In some other embodiments, as Figure 4 shown, step 150 includes:
[0065] Step 410, classifying the sample audio in the sample audio set according to a plurality of preset extended weight intervals and the extended weights corresponding to each sample audio, to obtain a sample audio subset corresponding to each extended weight interval.
[0066] In some embodiments, a plurality of extended weight intervals can be preset in advance. Each extended weight interval corresponds to a numerical range of weight values. The extended weight interval to which the extended weight of each sample audio determined in step 140 belongs can be determined according to the weight value of the extended weight of each sample audio.
[0067] Specifically, a plurality of extended weight intervals can be preset in advance. The sample audio in the sample audio set can be classified according to the extended weight intervals to determine the sample audio belonging to the same extended weight interval. The sample audio belonging to the same extended weight interval constitutes the sample audio subset corresponding to this extended weight interval.
[0068] Step 420, determining the extension quantity corresponding to each extended weight interval, where the extension quantity is positively correlated with the extended weight.
[0069] That the extension quantity is positively correlated with the extended weight means that the higher the extended weight, the more the extension quantity corresponding to the extended weight interval where the extended weight is located. The extension quantity corresponding to each extended weight interval refers to the number of sample audio to be extended selected from the sample audio subset corresponding to each extended weight interval.
[0070] In some embodiments, the extension quantity corresponding to each extended weight interval can be preset in advance.
[0071] In some other embodiments, as Figure 5 shown, step 420 includes:
[0072] Step 510, obtaining the extension ratio set for each extended weight interval.
[0073] The extension ratio refers to the proportion of the number of sample audio selected in each extension interval in the total number of all sample audio in the sample audio set. The specific value of this extension ratio is set by the user and is not specifically limited here.
[0074] Step 520: Determine the expansion quantity corresponding to each expansion weight interval according to the expansion ratio set for each expansion weight interval and the total number of sample audios in the sample audio set.
[0075] In some embodiments, the expansion quantity is obtained by multiplying the expansion ratio by the total number of sample audios in the sample audio set. For example, if the expansion ratio set for an expansion interval is 30% and the total number of sample audios in the sample audio set is 200, and the expansion quantity corresponding to this expansion weight interval is a, then a = 30% * 200 = 60.
[0076] Step 430: For each expansion weight interval, select sample audios from the sample audio subset corresponding to the expansion weight interval according to the expansion quantity corresponding to the expansion weight interval.
[0077] In some embodiments, sample audios corresponding to the expansion quantity can be selected from the sample audio subsets corresponding to each expansion weight interval. During the selection process, it can be random selection or selection in descending order of the expansion weight, which is not specifically limited herein.
[0078] Step 440: Use the selected sample audios as the target sample audios to be expanded.
[0079] Please continue to refer to Figure 1 , Step 160: Expand the target sample audio to obtain an expanded audio; the expanded audio and the sample audios in the sample audio set are used to train the speech recognition model.
[0080] The essence of speech recognition is a pattern recognition based on speech feature parameters, that is, through learning, the speech recognition model can classify the input speech according to a certain pattern, and then find the best matching result according to the decision criterion. In the speech recognition task in the solution of this application, the speech recognition model is used to recognize the text content corresponding to the audio.
[0081] Before applying the speech recognition model, the speech recognition model needs to be trained. Since a large amount of training data is required to train the speech recognition model, the sample audios used to train the speech recognition model can be expanded, and both the expanded audio and the sample audios are used as training data for training the speech recognition model to enrich the training data.
[0082] In some embodiments, Step 160 includes:
[0083] Perform specified processing on the target sample audio to obtain the corresponding expanded audio, and the specified processing includes at least one of noise addition processing, speech speed acceleration processing, and reverberation addition processing.
[0084] By performing at least one of adding noise, increasing the speech rate, and adding reverberation to the sample audio, the content scenario of the sample audio can be made more diverse, and it can also ensure that the speech recognition model can perform speech recognition based on the collected noisy audio data, reverberant audio data, and fast-paced audio data during application.
[0085] The speech recognition model is used to recognize the text content corresponding to an audio segment. In some embodiments, the specified language for the speech recognition model to output can be the same as or different from the language of the audio used to train the speech recognition model, which can be specifically set according to actual needs. For example, if the input audio of the speech recognition model is Cantonese audio, the output text of the speech recognition model can be the Cantonese text corresponding to this audio, or Mandarin text, or English text. The language type of the text content output by the speech recognition model can be set according to actual needs, and no specific limitation is provided here.
[0086] Before training the speech recognition model using the sample audio, it is necessary to select multiple sample audios for data expansion to ensure that the amount of training data used to train the speech recognition model is rich enough. For a large-scale sample audio set (such as a sample audio set in Mandarin), randomly selecting multiple sample audios for data expansion can obtain richer training data for training the speech recognition model. However, for a small-scale sample audio set, randomly selecting multiple sample audios for data expansion cannot ensure that sample audios with low phoneme word frequencies (such as sample audios containing multiple rare characters) can be selected, thereby increasing the sparsity of the training data used to train the speech recognition model. To reduce the sparsity of the training data, multiple sample audios can be purposefully selected as the target sample audios to be expanded according to the expansion weights corresponding to each sample audio.
[0087] In the solution of this application, first, according to the phoneme text corresponding to each sample audio in the sample audio set, then, according to the phoneme text corresponding to each sample audio, the phoneme word frequency of each phoneme is statistically calculated. Based on the statistically calculated phoneme word frequency of each phoneme, the average value of the phoneme word frequency of the corresponding phoneme in each sample audio is calculated to obtain the first average word frequency corresponding to each sample audio. Then, according to the first average word frequency of each sample audio, the expansion weight of each sample audio is determined. Thus, according to the expansion weights of each sample audio, the target sample audios to be expanded are determined in the sample audio set. Finally, the target sample audios are expanded to obtain expanded audios, and the expanded audios and the sample audios in the sample audio set are used to train the speech recognition model.
[0088] In the solution of this application, by determining the phoneme word frequency of each phoneme in each sample audio in the sample audio set, then calculating the corresponding first average word frequency of each sample audio, and determining the expansion weight of the sample audio based on the first average word frequency. Since the first average word frequency can reflect the overall situation of phoneme coverage in the phoneme text corresponding to each sample audio, and there is a negative correlation between the first average word frequency and the expansion weight, it is possible to determine a higher expansion weight for the sample audio with a lower phoneme coverage rate. Furthermore, the probability that the sample audio with a lower phoneme coverage rate is selected for audio expansion is higher. In this way, since the probability that the sample audio with a lower phoneme coverage rate is expanded is higher, compared with only using the sample audio set as the training data of the speech recognition model, using both the expanded audio obtained by expansion and the sample audio in the sample audio set as the training data of the speech recognition model improves the balance of phoneme coverage in the training data. Therefore, training the speech recognition model based on the expanded audio and the sample audio in the sample audio set can ensure the recognition accuracy of the trained speech recognition model, and solve the problem of low recognition accuracy caused by the uneven distribution of the sample audio used to train the speech recognition model in terms of pronunciation.
[0089] According to an aspect of this application, a speech recognition method is shown. This method can be executed by an electronic device with computing and processing capabilities, such as terminal devices like desktop computers and laptop computers. This method can also be executed interactively by a processing system including a server and a terminal. The method includes: obtaining the speech to be recognized, where the speech to be recognized is Cantonese speech; performing speech recognition on the speech to be recognized by the speech recognition model to obtain the text content of the speech to be recognized, and the speech recognition model is trained by using the expanded audio obtained by the above-mentioned expansion method for Cantonese audio and the sample audio in the sample audio set.
[0090] In some other embodiments, the speech to be recognized can also be Tibetan speech, Northeast dialect speech, Sichuan dialect speech, etc., which are not specifically limited here.
[0091] In some embodiments, as Figure 6 shown, before performing speech recognition using the speech recognition model, the method further includes:
[0092] Step 610, obtaining the text content corresponding to each sample audio.
[0093] In some embodiments, the text content corresponding to the sample audio is the Mandarin text corresponding to the sample audio. In other embodiments, the text content corresponding to the sample audio can also be English text, Cantonese text, or text in other languages, and the language type of the text content corresponding to the sample audio can be selected according to actual needs. Step 620: Perform speech recognition on the training audio by the speech recognition model to obtain the predicted text content corresponding to the training audio, where the training audio is the sample audio or the extended audio corresponding to the sample audio.
[0094] In some embodiments, a neural network model for speech recognition can be constructed, and the neural network model for speech recognition is referred to as a speech recognition model. Optionally, the speech recognition model can be constructed by a convolutional neural network, a recurrent neural network, a long short-term memory neural network, a fully connected neural network, a feedforward neural network, etc.
[0095] Step 630: Calculate the model loss according to the predicted text content corresponding to the training audio and the text content corresponding to the training audio; if the training audio is the extended audio corresponding to the sample audio, the text content corresponding to the training audio is the text content corresponding to the sample audio from which the extended audio is derived.
[0096] In some embodiments, the loss of the speech recognition model can be the difference between the predicted text content and the actual text content. Optionally, the loss of the speech recognition model can be CTC (Connectionist Temporal Classification) loss, cross-entropy loss, absolute value loss, etc., which are not specifically limited herein. CTC loss means that during the training process of the speech recognition model, the predicted text content and the actual text content are corresponded one by one, and then the difference between the predicted text content and the actual text content is calculated, and this difference is determined as the loss of the speech recognition model.
[0097] Step 640: Adjust the parameters of the speech recognition model in the reverse direction according to the model loss until the training end condition is reached.
[0098] In some embodiments, the training end condition can be that the number of iterations of the speech recognition model reaches the number threshold, or the loss value of the speech recognition model loss is not greater than the loss threshold, and the loss threshold can be set according to actual needs.
[0099] When the training end condition is reached, it means that the training of the speech recognition model is completed, and the speech recognition model can be applied in real life.
[0100] Figure 7 It is a block diagram of an extended device for Cantonese audio according to an embodiment of the present application, as Figure 7As shown, the expansion device 700 for Cantonese audio includes: an acquisition module 710, a phoneme word frequency statistics module 720, a first average word frequency determination module 730, an expansion weight determination module 740, a target sample audio determination module 750, and an expanded audio determination module 760.
[0101] The acquisition module 710 is configured to acquire the phoneme text corresponding to each sample audio in the sample audio set; the phoneme text includes at least one phoneme; the phoneme word frequency statistics module 720 is configured to statistically calculate the phoneme word frequency of each phoneme according to the phoneme text corresponding to each sample audio in the sample audio set; the first average word frequency determination module 730 is configured to, for each sample audio, calculate the average value of the phoneme word frequencies corresponding to the phonemes in the phoneme text corresponding to the sample audio, to obtain the first average word frequency corresponding to the sample audio; the expansion weight determination module 740 is configured to determine the expansion weight corresponding to the sample audio according to the first average word frequency corresponding to the sample audio, wherein the expansion weight has a negative correlation with the first average word frequency; the target sample audio determination module 750 is configured to determine the target sample audio to be expanded in the sample audio set according to the expansion weight corresponding to each sample audio; the expanded audio determination module 760 is configured to perform audio expansion on the target sample audio to obtain the expanded audio; the expanded audio and the sample audio in the sample audio set are used to train the speech recognition model.
[0102] In some embodiments, the target sample audio determination module 750 includes: a sample audio subset determination unit configured to classify the sample audio in the sample audio set according to a preset plurality of expansion weight intervals and the expansion weight corresponding to each sample audio, to obtain a sample audio subset corresponding to each expansion weight interval; an expansion quantity determination unit configured to determine the expansion quantity corresponding to each expansion weight interval, the expansion quantity having a positive correlation with the expansion weight; a sample audio selection unit configured to, for each expansion weight interval, select sample audio from the sample audio subset corresponding to the expansion weight interval according to the expansion quantity corresponding to the expansion weight interval; a first target sample audio determination unit configured to use the selected sample audio as the target sample audio to be expanded.
[0103] In some embodiments, the expansion quantity determination unit includes: an expansion ratio acquisition sub-unit configured to acquire the expansion ratio set for each expansion weight interval; an expansion quantity determination sub-unit configured to determine the expansion quantity corresponding to each expansion weight interval according to the expansion ratio set for each expansion weight interval and the total number of sample audio in the sample audio set.
[0104] In some embodiments, the expanded audio determination module 760 includes a processing unit configured to perform specified processing on the target sample audio to obtain the corresponding expanded audio, and the specified processing includes at least one of noise addition processing, speech rate acceleration processing, and reverberation addition processing.
[0105] In some embodiments, the extended weight determination module 740 includes: a difference calculation unit configured to calculate the difference between the first average word frequency corresponding to the sample audio and the word frequency threshold; a numerical range determination unit configured to determine the numerical range to which the difference belongs; and a first extended weight determination unit configured to use, as the extended weight corresponding to the sample audio, the extended weight corresponding to the numerical range to which the difference belongs according to the correspondence between the numerical range and the extended weight.
[0106] In some other embodiments, the extended weight determination module 740 includes: a target word frequency range determination unit configured to determine the target word frequency range to which the first average word frequency corresponding to the sample audio belongs; a second extended weight determination unit configured to determine the extended weight corresponding to the target word frequency range according to the correspondence between the word frequency range and the extended weight; and a third extended weight determination unit configured to use, as the extended weight corresponding to the sample audio, the extended weight corresponding to the target word frequency range.
[0107] In some other embodiments, the target sample audio determination module 750 includes: a sample audio determination unit configured to determine, from the sample audio set, a sample audio whose extended weight is greater than the weight threshold according to the extended weight corresponding to each sample audio; and a second target sample audio determination unit configured to use the sample audio whose extended weight is greater than the weight threshold as the target sample audio to be extended.
[0108] According to one aspect of the present application, there is provided a speech recognition device, which includes: a to-be-recognized speech acquisition module configured to acquire to-be-recognized speech, where the to-be-recognized speech is Cantonese speech; and a speech recognition module configured to perform speech recognition on the to-be-recognized speech by a speech recognition model to obtain the text content of the to-be-recognized speech, where the speech recognition model is trained by using the extended audio obtained by the above-mentioned extension method for Cantonese audio and the sample audio in the sample audio set.
[0109] In some embodiments, the speech recognition device further includes: a text content acquisition module configured to acquire the text content corresponding to each sample audio; a speech recognition module configured to perform speech recognition on the training audio by the speech recognition model to obtain the predicted text content corresponding to the training audio, where the training audio is a sample audio or the extended audio corresponding to the sample audio; a model loss calculation module configured to calculate the model loss according to the predicted text content corresponding to the training audio and the text content corresponding to the training audio, where if the training audio is the extended audio corresponding to the sample audio, the text content corresponding to the training audio is the text content of the sample audio from which the extended audio is derived; and an adjustment module configured to reversely adjust the parameters of the speech recognition model according to the model loss until the training end condition is reached. In this embodiment, the text content corresponding to the sample audio is the Mandarin text corresponding to the sample audio.
[0110] According to an aspect of the present application, an electronic device is further provided, which includes: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods in any of the above embodiments are implemented.
[0111] According to an aspect of an embodiment of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods in any of the above embodiments.
[0112] According to an aspect of an embodiment of the present application, an electronic device is further provided, as Figure 8 shown. The electronic device 800 includes a processor 810 and one or more memories 820. The one or more memories 820 are used to store program instructions to be executed by the processor 810. When the processor 810 executes the program instructions, the above object recognition method is implemented.
[0113] Further, the processor 810 may include one or more processing cores. The processor 810 runs or executes instructions, programs, code sets or instruction sets stored in the memory 820, and calls data stored in the memory 820. Optionally, the processor 810 may be implemented in at least one of the hardware forms of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 810 may integrate one or a combination of several of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor and may be implemented separately by a communication chip.
[0114] According to an aspect of the present application, a computer-readable storage medium is further provided. The computer-readable medium may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by the processor, the methods in any of the above embodiments are implemented.
[0115] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0116] The units described in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.
[0117] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.
[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0119] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0120] It should be understood that the present application is not limited to the exact structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. An expansion method for Cantonese audio, characterized in that The method includes: Obtaining the phoneme text corresponding to each sample audio in the sample audio set; the phoneme text includes at least one phoneme; the sample audio is a Cantonese audio; According to the phoneme text corresponding to each sample audio in the sample audio set, counting the phoneme word frequencies of each phoneme; For each sample audio, calculating the average value of the phoneme word frequencies corresponding to the phonemes in the phoneme text corresponding to the sample audio to obtain the first average word frequency corresponding to the sample audio; According to the first average word frequency corresponding to the sample audio, determining the expansion weight corresponding to the sample audio, wherein the expansion weight has a negative correlation with the first average word frequency; According to a plurality of preset expansion weight intervals and the expansion weights corresponding to each sample audio, classifying the sample audios in the sample audio set to obtain a sample audio subset corresponding to each expansion weight interval; Determining the expansion quantity corresponding to each expansion weight interval, where the expansion quantity is positively correlated with the expansion weight; For each expansion weight interval, selecting sample audios from the sample audio subset corresponding to the expansion weight interval according to the expansion quantity corresponding to the expansion weight interval; Taking the selected sample audio as the target sample audio to be expanded; Performing audio expansion on the target sample audio to obtain an expanded audio; the expanded audio and the sample audios in the sample audio set are used to train a speech recognition model.
2. The method according to claim 1, characterized in that, The determining the expansion quantity corresponding to each expansion weight interval includes: Obtaining the expansion ratio set for each expansion weight interval; According to the expansion ratio set for each expansion weight interval and the total number of sample audios in the sample audio set, determining the expansion quantity corresponding to each expansion weight interval.
3. The method according to claim 1, wherein The performing audio expansion on the target sample audio to obtain an expanded audio includes: Performing specified processing on the target sample audio to obtain the corresponding expanded audio, where the specified processing includes at least one of noise addition processing, speech rate acceleration processing, and reverberation addition processing.
4. The method according to claim 1, wherein The according to the first average word frequency corresponding to the sample audio, determining the expansion weight corresponding to the sample audio includes: Calculating the difference between the first average word frequency corresponding to the sample audio and a word frequency threshold; Determining the numerical interval to which the difference belongs; According to the correspondence between the numerical interval and the expansion weight, taking the expansion weight corresponding to the numerical interval to which the difference belongs as the expansion weight corresponding to the sample audio.
5. The method according to claim 1, characterized in that The according to the first average word frequency corresponding to the sample audio, determining the expansion weight corresponding to the sample audio includes: Determining the target word frequency interval to which the first average word frequency corresponding to the sample audio belongs; According to the correspondence between the word frequency interval and the expansion weight, determining the expansion weight corresponding to the target word frequency interval; Taking the expansion weight corresponding to the target word frequency interval as the expansion weight corresponding to the sample audio.
6. A voice recognition method, characterized in that, The method includes: Obtaining the speech to be recognized, where the speech to be recognized is Cantonese speech; The speech recognition model performs speech recognition on the speech to be recognized to obtain the text content of the speech to be recognized. The speech recognition model is trained using the sample audio in the extended audio and sample audio set obtained by the method described in any one of claims 1-5.
7. The method according to claim 6, wherein Before the speech recognition model performs speech recognition on the speech to be recognized to obtain the text content of the speech to be recognized, the method further includes: Obtaining the text content corresponding to each sample audio; The speech recognition model performs speech recognition on the training audio to obtain the predicted text content corresponding to the training audio. The training audio is the sample audio or the extended audio corresponding to the sample audio; Calculating the model loss according to the predicted text content corresponding to the training audio and the text content corresponding to the training audio; if the training audio is the extended audio corresponding to the sample audio, the text content corresponding to the training audio is the text content of the sample audio from which the extended audio is derived; Adjusting the parameters of the speech recognition model in the reverse direction according to the model loss until the training end condition is reached.
8. The method according to claim 7, wherein The text content corresponding to the sample audio is the Mandarin text corresponding to the sample audio.
Citation Information
Patent Citations
Voice recognition model training method and device
CN111768761A
End-to-end mandarin and low-resource Cantonese unified identification method
CN119274539A
Method and device for speech recognition, terminal and storage medium
WO2021135611A1