Data augmentation method and speech recognition method
By acquiring and labeling the first speech data and using the speech conversion network to generate expanded speech, the problem of insufficient training data in the multilingual speech recognition model is solved, and the recognition accuracy of the model in a multilingual environment is improved.
Patent Information
- Application Number
- CN202410261957.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-07
- Publication Date
- 2025-09-09
AI Technical Summary
Existing multilingual speech recognition models have difficulty effectively recognizing the speech of non-native speakers due to insufficient training data, resulting in reduced recognition accuracy. In particular, in a mixed Chinese-English environment, the model is prone to recognition errors.
By obtaining multiple first speech data, annotated with text labels, and using the speech conversion model obtained by training the speech conversion network, the second speech is speech converted to generate an expanded speech with the native accent of the first language, thereby enriching the training data and improving the accuracy of the model.
The generated augmented speech is close to real data, effectively achieving data augmentation and improving the recognition accuracy of the speech recognition model, especially in a multilingual environment, and can better recognize the speech of non-native speakers.
Smart Images

Figure CN120612924A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of speech recognition technology, and in particular to a data augmentation method and a speech recognition method. Background Art
[0002] Automatic Speech Recognition (ASR) technology is a core technology that transcribes input speech data into corresponding text content. Currently, its application in e-commerce, finance, logistics, and other fields is becoming increasingly widespread. Typically, ASR models only support monolingual speech recognition tasks, meaning that an ASR model can only recognize a specific language. For example, a Chinese ASR model can only be used for Chinese speech recognition, and an English ASR model can only be used for English speech recognition.
[0003] In multilingual speech recognition applications, traditional automatic speech recognition technology requires maintaining multiple automatic speech recognition models to identify every possible language, which is limited. With the increasing popularity of automatic speech recognition technology, the need to use a single speech recognition model to recognize multiple languages is becoming increasingly urgent. However, the main challenge facing multilingual speech recognition is the lack of training data. Summary of the Invention
[0004] In view of this, some embodiments of the present application provide a data augmentation method, so that the expanded speech has the accent of the original speech and is close to the real data, which can effectively realize data augmentation for the speech recognition model and is conducive to improving the accuracy of the model.
[0005] In a first aspect, some embodiments of the present application provide a data augmentation method, comprising: obtaining multiple first voices, where the first voices are voices in a first language with a first accent, the first accent being the native accent of the first language, and the first voices are annotated with text labels; using the multiple first voices to train a preset voice conversion network, and performing voice conversion on multiple second voices using the trained voice conversion model to obtain multiple expanded voices; wherein the second voices are voices in a second language with a second accent, the second accent being the native accent of the second language, and the expanded voices are voices in the second language with the first accent.
[0006] For example, the first speech is speech data of a native Chinese speaker speaking Chinese, and the second speech is speech data of a native English speaker speaking English. In this embodiment, a preset speech conversion network is trained using multiple first speech sounds. The trained speech conversion model learns the first accent, so that the second speech sound can be converted to an expanded speech sound with the first accent, for example, mimicking speech data of a native Chinese speaker speaking English. This means that the expanded speech sound has the same accent as the original speech sound (the collected real speech sound), which is closer to the real data. This effectively enables data augmentation for the speech recognition model, which helps improve the accuracy of the model.
[0007] In some embodiments, the speech conversion model includes a speech feature extraction module and a vocoder, wherein the speech feature extraction module is used to extract speech features from input speech, and the vocoder is used to convert the speech features into audio data.
[0008] In this embodiment, the speech conversion model is configured to include a speech feature extraction module and a vocoder, so that after the input second speech is subjected to feature extraction by the speech feature extraction module, the extracted speech features are input into the vocoder and converted into audio data (i.e., the expanded speech). Because the speech conversion model is trained using multiple first speech voices and learns the first accent features of multiple first speech voices, during the process of extracting speech features from the second speech voice and converting it into audio data, the first accent is added to replace the original second accent of the second speech voice, so that the output expanded speech has the first accent. This effectively simulates the speech of a non-native speaker (a native speaker of the first language) speaking the second language, making the expanded speech more realistic.
[0009] In some embodiments, the speech feature extraction module includes a content extraction module, a speaker information extraction module, and an acoustic module; wherein the content extraction module is used to extract content features from the input speech; the speaker information extraction module is used to extract speaker information from the input speech; and the acoustic module is used to fuse content features with speaker information to obtain speech features.
[0010] In this embodiment, the speech feature extraction module includes a content extraction module, a speaker information extraction module, and an acoustic module. After the first input speech undergoes content extraction and speaker information extraction, the extracted content features and speaker information are fused to generate speech features. This ensures that the speech features accurately represent the textual meaning of the input speech while also possessing timbre characteristics consistent with the speaker information, facilitating the training of an accurate speech conversion model.
[0011] In some embodiments, the content extraction module includes multiple convolutional layers, each configured with a stride of 1. In this embodiment, the content extraction module does not downsample the input second speech. During the feature extraction process, the dimensionality of the speech data does not change, which helps maintain the intelligibility and naturalness of the speech conversion, thereby improving the speech conversion effect.
[0012] In some embodiments, the trained speech conversion model performs speech conversion on multiple second speech to obtain multiple expanded speech, including: for any second speech, inputting the second speech and speaker information into the speech conversion model, and outputting an expanded speech.
[0013] In this embodiment, the second speech and speaker information are input into the speech conversion model, which converts the second accent of the second speech into the first accent and adds timbre features according to the speaker information, so that the expanded speech has the first accent and timbre features that match the speaker information.
[0014] In some embodiments, the method further comprises: constructing a speaker Gaussian mixture model using speaker information extracted from the plurality of first voices by a speaker information extraction module in the voice conversion model; sampling and outputting virtual speaker information using the speaker Gaussian mixture model;
[0015] The step of inputting the second voice and speaker information into the voice conversion model to output an expanded voice includes: inputting the second voice and virtual speaker information into the voice conversion model to obtain the expanded voice.
[0016] In this embodiment, a Gaussian mixture model is constructed to represent the speaker information corresponding to multiple first speech signals. This allows the virtual speaker information input into the speech conversion model to be determined based on the Gaussian mixture model, thereby making the timbre characteristics of the expanded speech closer to those of the real data (the first speech). This effectively addresses the scarcity of speech in a second language for non-native speakers. Furthermore, the expanded speech has the first accent, which is more realistic.
[0017] In some embodiments, the speaker information extracted from the multiple first voices by the speaker information extraction module in the aforementioned speech conversion model is used to construct a speaker Gaussian mixture model, including: grouping the speaker information corresponding to the multiple first voices according to multiple feature dimensions to obtain multiple groups of speaker information; and constructing a corresponding speaker Gaussian mixture model for each group of speaker information.
[0018] In this embodiment, speaker information corresponding to multiple first voices is grouped and modeled, and each group of speaker information is clustered, so that the Gaussian mixture model of each group of speakers is more accurate and the convergence of the speaker Gaussian mixture model is accelerated.
[0019] In a second aspect, some embodiments of the present application provide a method for training a speech recognition model, comprising: augmenting a training set using the method in the first aspect; iteratively training a speech recognition network using the augmented training set until convergence to obtain a speech recognition model.
[0020] In a third aspect, some embodiments of the present application provide a speech recognition method, comprising: acquiring speech data; inputting the speech data into the speech recognition model in the second aspect to obtain recognized text.
[0021] In a fourth aspect, some embodiments of the present application provide an electronic device, including:
[0022] at least one processor, and
[0023] a memory communicatively coupled to at least one processor, wherein:
[0024] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method in the first aspect, the second aspect, or the third aspect.
[0025] In a fifth aspect, some embodiments of the present application provide a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable a computer device to execute the method in the first aspect, the second aspect, or the third aspect.
[0026] The beneficial effects of the embodiments of the present application are as follows: Unlike the prior art, the data augmentation method provided in the embodiments of the present application obtains multiple first speech sounds, wherein the first speech sounds are speech sounds in a first language with a first accent, the first accent being the native accent of the first language, and the first speech sounds are annotated with text labels. A preset speech conversion network is trained using the multiple first speech sounds. The trained speech conversion model performs speech conversion on multiple second speech sounds to obtain multiple augmented speech sounds. The second speech sounds are speech sounds in a second language with a second accent, the second accent being the native accent of the second language, and the augmented speech sounds are speech sounds in the second language with the first accent. For example, the first speech sounds are speech data of a native Chinese speaker speaking Chinese, and the second speech sounds are speech data of a native English speaker speaking English. The preset speech conversion network is trained using the multiple first speech sounds, and the trained speech conversion model learns the first accent. Thus, the speech conversion of the second speech sounds is performed so that the resulting augmented speech sounds have the first accent, for example, imitating speech data of a native Chinese speaker speaking English. That is, the expanded speech has the accent of the original speech (the collected real speech), is close to the real data, can effectively realize the data augmentation for the speech recognition model, and is conducive to improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0028] Figure 1 This is a schematic diagram of the structure of the speech recognition system in some embodiments of the present application;
[0029] Figure 2 This is a schematic diagram of the structure of an electronic device in some embodiments of the present application;
[0030] Figure 3 Schematic diagram of the flow of data augmentation methods in some embodiments of the present application;
[0031] Figure 4 This is a schematic diagram of the process of model training and data augmentation in some embodiments of the present application;
[0032] Figure 5 This is a flowchart of a method for training a speech recognition model in some embodiments of the present application;
[0033] Figure 6 This is a flowchart of the speech recognition method in some embodiments of the present application. DETAILED DESCRIPTION
[0034] The present application is described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but are not intended to limit the present application in any form. It should be noted that those skilled in the art may make several variations and improvements without departing from the scope of the present application. These all fall within the scope of protection of the present application.
[0035] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0036] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other and are all within the scope of protection of the present application. In addition, although the functional modules are divided in the device schematic and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flow chart. In addition, the words "first", "second", "third", etc. used herein do not limit the data and execution order, but only distinguish between the same items or similar items with basically the same functions and effects.
[0037] Unless otherwise defined, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification and in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the relevant listed items.
[0038] In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0039] In multilingual speech recognition applications, users may speak non-native languages in addition to their native language, and the speech recognition model must recognize speech in multiple languages. For example, a speech recognition model for Chinese speakers may need to recognize English speech in addition to Chinese speech; and a speech recognition model for English speakers may need to recognize Chinese speech in addition to English speech.
[0040] Some existing multilingual speech recognition models are trained using multilingual sample speech data and the corresponding multilingual sample speech recognition results. The inventors of this application are aware of several improvements that primarily address the network structure and training methods of speech recognition models. For example, these improvements involve using a hybrid expert network as the speech recognition model, or first training an acoustic phoneme classifier as the target network on labeled data in a high-resource language, and then training a master network to approximate the acoustic phoneme classifier's representation across multiple languages.
[0041] However, the multilingual sample speech data used in these training programs do not take accent issues into consideration. For example, the Chinese speech in the multilingual sample speech data is produced by native Chinese speakers, and the English speech is produced by native English speakers.
[0042] Understandably, every language has its own unique phonetic characteristics and linguistic structure. For example, Chinese and English have different phonetic characteristics and linguistic structures. Due to language acquisition methods and phonetic habits, speakers of different native languages may develop accents when speaking a non-native language. Understandably, when someone learns a non-native language, they may be influenced by aspects of their native language, such as phonetics, intonation, grammar, and rhythm. This can lead non-native speakers to retain some native-language characteristics in pronunciation and intonation, forming an accent.
[0043] In a multilingual environment, if accents are not considered, speech recognition models are prone to errors. For example, if a model learns the characteristics of English speech from native English speakers, it will easily make recognition errors when it encounters English speech from native Chinese speakers.
[0044] For any machine learning model, data is the foundation of training and the guarantee of model performance. For example, in a mixed Chinese-English environment, each language has its own unique speech characteristics and language structure, which requires a large amount of labeled data to train and optimize the model. However, in reality, speech from non-native speakers is difficult to collect. For example, obtaining a large amount of speech and labeled data of Chinese people speaking English, or foreigners speaking Chinese, requires a lot of manpower, material resources, and time. Without corresponding data, the model cannot effectively learn and understand the speech characteristics and differences between Chinese and English. This leads to a decline in model performance in mixed Chinese-English speech recognition, and it is unable to accurately recognize speech in different languages.
[0045] In view of this, some embodiments of the present application provide a data augmentation method, which obtains multiple first speech sounds, wherein the first speech sounds are speech sounds in a first language with a first accent, the first accent being the native accent of the first language, and the first speech sounds are annotated with text labels. A preset speech conversion network is trained using the multiple first speech sounds, and the trained speech conversion model is used to perform speech conversion on multiple second speech sounds, thereby obtaining multiple augmented speech sounds. The second speech sounds are speech sounds in a second language with a second accent, the second accent being the native accent of the second language, and the augmented speech sounds are speech sounds in the second language with the first accent. For example, the first speech sounds are speech data of a native Chinese speaker speaking Chinese, and the second speech sounds are speech data of a native English speaker speaking English. The preset speech conversion network is trained using the multiple first speech sounds, and the trained speech conversion model learns the first accent. Thus, the speech conversion of the second speech sounds is performed so that the resulting augmented speech sounds have the first accent, thereby mimicking speech data of a native Chinese speaker speaking English. That is, the expanded speech has the accent of the original speech (the collected real speech), is close to the real data, can effectively realize the data augmentation for the speech recognition model, and is conducive to improving the accuracy of the model.
[0046] Some embodiments of the present application also provide a method for training a speech recognition model, wherein a training set is augmented using the data augmentation method described in the above embodiments. The augmented training set is then used to iteratively train a speech recognition network until convergence, thereby obtaining a speech recognition model. Because the augmented training set is rich and authentic, and the augmented speech has the accent of the original speech (the collected real speech), the trained speech recognition model learns the accent characteristics of the speaker in multiple languages, thereby achieving better speech recognition accuracy.
[0047] Some embodiments of the present application also provide a speech recognition method, which obtains speech data, inputs the speech data into the speech recognition model trained in the above embodiments, and obtains recognized text.
[0048] The following describes exemplary applications of electronic devices for data augmentation, for training speech recognition models, or for speech recognition provided in embodiments of the present application. It can be understood that the electronic device can perform data augmentation, train a speech recognition model based on a training set after data augmentation, and use the speech recognition model for speech recognition.
[0049] The electronic device provided in the embodiment of the present application can be a server, such as a server deployed in the cloud. When the server is used for data augmentation, based on the multiple voice data provided by other devices or those skilled in the art, the multiple voice data are augmented to obtain multiple expanded voices. When the server is used to train a speech recognition model, based on the training set and neural network provided by other devices or those skilled in the art, the training set is first augmented according to the data augmentation method in the above embodiment, and the neural network is iteratively trained using the augmented training set to determine the final model parameters, so that the neural network configures the final model parameters to obtain a speech recognition model. When the server is used for speech recognition, the built-in speech recognition model is called to perform corresponding calculations on the voice data provided by other devices or users to obtain corresponding recognition text.
[0050] The electronic device provided in some embodiments of the present application may be a terminal of various types, such as a laptop computer, a desktop computer, or a mobile device. The above-mentioned data augmentation method, the method for training a speech recognition model, or the speech recognition method may also be implemented on the terminal.
[0051] For example, see Figure 1 , Figure 1 This is a schematic diagram of an application scenario of the speech recognition system provided in an embodiment of the present application. The terminal 10 is connected to the server 20 via a network, wherein the network can be a wide area network or a local area network, or a combination of the two.
[0052] The terminal 10 can be used to obtain a training set and construct a neural network. For example, a person skilled in the art downloads a prepared training set on the terminal and builds a network structure of the neural network.
[0053] In some embodiments, the terminal 10 locally executes the data augmentation method provided in the embodiments of the present application to complete the augmentation processing of the training set, uses the augmented training set to train the designed neural network, determines the final model parameters, and thus the neural network configures the final model parameters to obtain a speech recognition model. In some embodiments, the terminal 10 can also send the training set and the constructed neural network stored on the terminal by a person skilled in the art to the server 20 through the network. The server 20 receives the training set and the neural network, first performs augmentation processing on the training set, uses the augmented training set to train the designed neural network, determines the final model parameters, and then sends the final model parameters to the terminal 10. The terminal 10 saves the final model parameters so that the neural network can configure the final model parameters to obtain a speech recognition model.
[0054] In some embodiments, the terminal 10 locally executes the speech recognition method provided in the embodiments of the present application to provide speech recognition services to the user, calling the built-in speech recognition model to perform corresponding computations on the speech data to obtain recognized text. In some embodiments, the terminal 10 may also send the speech data to the server 20 via the network. After receiving the speech data, the server 20 calls the built-in speech recognition model to perform corresponding computations on the speech data to obtain recognized text.
[0055] The structure of the electronic device in the embodiment of the present application is described below. Figure 2 5 is a schematic diagram of the structure of an electronic device 500 in an embodiment of the present application. The electronic device 500 includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the electronic device 500 are coupled together through a bus system 540. It can be understood that the bus system 540 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 3 Various buses are labeled as bus system 540 .
[0056] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0057] The user interface 530 includes one or more output devices 531 that enable presentation of media content, including one or more speakers and / or one or more visual displays. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, and other input buttons and controls.
[0058] Memory 550, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the data augmentation method, the method for training a speech recognition model, or the program instructions / modules corresponding to the speech recognition method in the embodiments of the present application. By executing the non-transitory software programs, instructions, and modules stored in memory 550, processor 510 can implement the data augmentation method, the method for training a speech recognition model, or the speech recognition method in any of the following method embodiments.
[0059] The memory 550 may include a volatile memory (VM), such as a random access memory (RAM); the memory 550 may also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); the memory 550 may also include a combination of the above types of memory.
[0060] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0061] A network communication module 552 for reaching other computing devices via one or more (wired or wireless) network interfaces 520 , exemplary network interfaces 520 including Bluetooth, WiFi, and USB;
[0062] a display module 553 for enabling presentation of information via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with the user interface 530 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0063] The input processing module 554 is used to detect one or more user inputs or interactions from one of the one or more input devices 532 and translate the detected inputs or interactions.
[0064] It can be understood from the above that the data augmentation method, the method for training a speech recognition model, or the speech recognition method provided in the embodiments of the present application can be implemented by various types of electronic devices with computing and processing capabilities, such as smart terminals and servers.
[0065] The following describes the data augmentation method provided by the embodiment of the present application in conjunction with the exemplary application and implementation of the server provided by the embodiment of the present application. Figure 3 , Figure 3 Schematic diagram of the data augmentation method provided in the embodiment of the present application.
[0066] S10: Acquire multiple first voices, where the first voices are voices in a first language with a first accent, where the first accent is a native accent of the first language, and the first voices are annotated with text labels.
[0067] The first speech is audio data representing at least one sentence or word. For example, any first speech is produced by a native Chinese speaker. In this embodiment, the first language is Chinese, and the first accent is a native Chinese accent. Each first speech is annotated with a text label, which is a textual representation of the first speech content.
[0068] For multiple first voices, these first voices are not exactly the same. They may be voices uttered by different people, and the content of the voices may not be exactly the same. These different people may be distributed differently in terms of gender, age, region, etc., so that the multiple first voices cover rich timbre features. It is understandable that these different people have the same mother tongue. Exemplarily, the multiple first voices are all Chinese spoken by Chinese people, and the text label of each voice annotation is the corresponding Chinese text. In this embodiment, the target group of the speech recognition model is Chinese, and the mother tongue is Chinese, so the first language is Chinese. The speech recognition model is mainly used to recognize Chinese or other languages (such as English) spoken by Chinese people.
[0069] In other embodiments, if the target group of the speech recognition model is a population whose native language is English, the first language is English, and the corresponding text label is English text.
[0070] It is understood that the first accent of the first language in the first speech is the native accent of the user, making collection easier. These multiple first speech sounds can be sourced from open source datasets. This application does not impose any restrictions on the method for collecting the first speech sounds.
[0071] S20: A preset speech conversion network is trained using the plurality of first speech sounds, and the trained speech conversion model is used to perform speech conversion on the plurality of second speech sounds to obtain a plurality of expanded speech sounds. The second speech sounds are speech sounds in a second language with a second accent, the second accent being the native accent of the second language, and the expanded speech sounds are speech sounds in the second language with a first accent.
[0072] It is understandable that the second language is different from the first language, and the second accent is different from the first accent. The second accent in the second speech is the native accent of the speaker, making it easier to collect. For example, the first speech is Chinese spoken by a native Chinese speaker, and the second speech is English spoken by a native English speaker. Using the speech conversion model, the second speech is converted into an expanded speech, which is English spoken by a native Chinese speaker. This reduces the difficulty of collecting audio from non-native speakers of the second language.
[0073] Compared with directly combining the second speech with the first speech to form training data without considering the accent problem, in this embodiment, by performing voice conversion on the second speech, only the accent is changed without changing the speech content, so that the expanded speech obtained by conversion conforms to the accent characteristics of the target group, thereby facilitating the training of a more accurate speech recognition model.
[0074] Among them, see Figure 4 The speech conversion model is obtained by training a preset speech conversion network on multiple first speech sounds. It is understood that the speech conversion network is a neural network that includes neural network components (such as convolutional layers, pooling layers, or activation functions). After learning the accent features (features of the first accent) of the multiple first speech sounds, the speech conversion network determines model parameters to obtain the speech conversion model. The speech conversion model performs accent conversion on the input second speech sound, so that the output expanded speech sound has the first accent.
[0075] In some embodiments, as Figure 4 As shown, the speech conversion model includes a speech feature extraction module and a vocoder, wherein the speech feature extraction module is used to extract speech features in the input speech, and the vocoder is used to convert the speech features into audio data.
[0076] It is understood that the speech feature extraction module includes several convolutional layers, and the input second speech is spatially mapped and transformed by these convolutional layers to obtain speech features. Speech features include speech content features or speaker timbre features.
[0077] A vocoder is a signal processing system that converts speech features into audio data, or in other words, encodes speech signals into digital or analog signals. In some embodiments, the vocoder can be a BigVGAN model. This BigVGAN model includes a generator and multiple discriminators. The generator introduces periodic nonlinearity and anti-aliasing representations, which bring the desired inductive bias to waveform synthesis, thereby significantly improving the quality of audio data.
[0078] In this embodiment, the speech conversion model is configured to include a speech feature extraction module and a vocoder, so that after the input second speech is subjected to feature extraction by the speech feature extraction module, the extracted speech features are input into the vocoder and converted into audio data (i.e., the expanded speech). Because the speech conversion model is trained using multiple first speech voices and learns the first accent features of multiple first speech voices, during the process of extracting speech features from the second speech voice and converting it into audio data, the first accent is added to replace the original second accent of the second speech voice, so that the output expanded speech has the first accent. This effectively simulates the speech of a non-native speaker (a native speaker of the first language) speaking the second language, making the expanded speech more realistic.
[0079] In some embodiments, as Figure 4 As shown, the speech feature extraction module includes a content extraction module, a speaker information extraction module, and an acoustic module. It is understandable that different modules are used to implement different functions. These modules all include components of a neural network, differing in their hierarchical structure, intra-layer structure, or parameters.
[0080] The content extraction module is used to extract content features from the input speech. Here, content features refer to the actual text or information expressed by the speaker in the speech. In some embodiments, the content extraction module includes multiple convolutional layers, each configured with a step size of 1. This means that the content extraction module does not downsample the input second speech. During the feature extraction process, the dimensionality of the speech data remains unchanged, which helps maintain the intelligibility and naturalness of the speech conversion, thereby improving the speech conversion effect.
[0081] The speaker information extraction module is used to extract speaker information from the input speech. The speaker information may include the speaker's gender, age, or region (for example, south or north). In some embodiments, the speaker information extraction module is built using the ECAPA TDNN network structure. The full name of the ECAPA TDNN network is Emphasized Channel-Attention and Position-Aware Temporal Convolutional Network, which is a deep neural network structure used for speech recognition. It introduces channel attention and position perception mechanisms based on TDNN (Time-Delay Neural Network), which can better capture the timing information in the speech signal. In other words, the ECAPA TDNN network is a deep neural network that combines timing modeling, channel attention, and position perception. The network structure of the speaker information extraction module is the ECAPA TDNN network, which can better extract speaker information accurately from the time series data of speech.
[0082] The acoustic module is used to fuse the content features output by the content extraction module with the speaker information output by the speaker information extraction module to obtain speech features. It is understandable that the speech features can reflect the content features and speaker information. In some embodiments, the acoustic module is built using the same network architecture as the Grad-TTS network model, which includes a speech content encoder constructed by a 6-layer transformer feedforward network and a decoder based on a diffusion model. The Grad-TTS network model is an existing network model in the field and will not be introduced in detail here. In some embodiments, the transformer feedforward network can be replaced by a conformer network structure.
[0083] In this embodiment, the speech feature extraction module includes a content extraction module, a speaker information extraction module, and an acoustic module. After the first input speech undergoes content extraction and speaker information extraction, the extracted content features and speaker information are fused to generate speech features. This ensures that the speech features accurately represent the textual meaning of the input speech while also possessing timbre characteristics consistent with the speaker information, facilitating the training of an accurate speech conversion model.
[0084] In some embodiments, the aforementioned “performing speech conversion on the plurality of second speech signals by the trained speech conversion model to obtain a plurality of expanded speech signals” specifically includes:
[0085] S21: For any second speech, the second speech and speaker information are input into a speech conversion model, and an expanded speech is output.
[0086] The speaker information may include the speaker's gender, age, or region (e.g., south or north). It is understood that the speaker information is not the information of the real speaker corresponding to the second voice. In some embodiments, the speaker information may be the speaker information corresponding to any one of the multiple first voices. In some embodiments, the speaker information is random information of a native speaker of the first language.
[0087] like Figure 4 As shown, the second voice and speaker information are input into the voice conversion model. The voice conversion model converts the original second accent of the second voice into the first accent, and adds timbre features according to the speaker information, so that the expanded voice has the first accent and timbre features that are consistent with the speaker information.
[0088] It is understandable that, for multiple second voices, the speaker information added to each second voice is not exactly the same, so that the multiple expanded voices are rich and diverse, and close to real data.
[0089] In some embodiments, the method S100 further includes:
[0090] S30: Constructing a speaker Gaussian mixture model using the speaker information extracted from the plurality of first voices by the speaker information extraction module in the voice conversion model.
[0091] Here, as Figure 4 As shown, the speaker information corresponding to multiple first speech sounds is used to construct a speaker Gaussian mixture model. A speaker Gaussian mixture model is a probabilistic model used to describe the distribution of a random variable. It is composed of a linear combination of multiple Gaussian distribution functions, with each Gaussian distribution function (also called a component) representing a category or cluster in the data. In clustering, each Gaussian distribution corresponds to a cluster, and the goal of the model is to describe the data distribution using appropriate parameters.
[0092] In some embodiments, the parameters of the speaker Gaussian mixture model are typically estimated using an expectation-maximization algorithm (EM algorithm). The EM algorithm iteratively updates the parameters to maximize the likelihood of the model for the speaker information (observation data) corresponding to the multiple first speech sounds.
[0093] In some embodiments, the aforementioned step S30 specifically includes:
[0094] S31: Grouping speaker information corresponding to the multiple first voices according to multiple feature dimensions to obtain multiple groups of speaker information;
[0095] S32: Constructing a corresponding speaker Gaussian mixture model for each group of speaker information.
[0096] Among them, the feature dimensions may include age, gender, and region. In some embodiments, according to age, those under 18 years old can be divided into teenagers, those aged 18 to 44 years old can be divided into young people, those aged 45 to 59 years old can be divided into middle-aged people, and those aged 60 and above can be divided into elderly people. According to gender, they can be divided into males and females. According to the region, they can be divided into the south and the north. Thus, there can be 16 combinations based on age, gender, and region. In other embodiments, the speaker information corresponding to multiple first voices can also be subjected to K-means cluster analysis according to age, gender, and region to be divided into groups.
[0097] The speaker information corresponding to the multiple first voices is divided into 16 groups according to the above combination, and a corresponding speaker Gaussian mixture model is constructed for each group of speaker information. For example, a Gaussian mixture model is constructed for the speaker information whose region is the South, whose age is young, and whose gender is male. In some embodiments, the specific modeling steps include: selecting an appropriate number of Gaussian components using the Bayesian Information Criterion (BIC); then initializing the mean and variance of each component using random initialization; and finally fitting the Gaussian mixture model using the expectation maximization algorithm (EM algorithm) until convergence. In this way, 16 speaker Gaussian mixture models can be obtained, and each speaker Gaussian mixture model is targeted at a different group.
[0098] In this embodiment, speaker information corresponding to multiple first voices is grouped and modeled, and each group of speaker information is clustered, so that the Gaussian mixture model of each group of speakers is more accurate and the convergence of the speaker Gaussian mixture model is accelerated.
[0099] S40: Use the speaker Gaussian mixture model to sample and output virtual speaker information.
[0100] For example, Figure 4 As shown, a Gaussian mixture model can be randomly selected from 16 speaker Gaussian mixture models, and the selection follows a uniform distribution. The speaker Gaussian mixture model outputs virtual speaker information. It is understandable that the virtual speaker information is speaker information that does not exist in reality and does not appear in the speaker information corresponding to the first speech.
[0101] In this embodiment, the aforementioned step S21 specifically includes: inputting the second speech and virtual speaker information into the speech conversion model to obtain the expanded speech.
[0102] The voice conversion model converts the second voice's original second accent into the first accent and adds timbre characteristics based on the virtual speaker information, giving the expanded voice the first accent and timbre characteristics consistent with the virtual speaker information. Because the virtual speaker information is rich and realistic, the expanded voices contain richer speaker information and greater diversity.
[0103] In this embodiment, a Gaussian mixture model is constructed to characterize the speaker information corresponding to multiple first speech sounds. This allows the virtual speaker information for the input speech conversion model to be determined based on the Gaussian mixture model. This allows the timbre characteristics of the expanded speech to approximate those of the real data (the first speech). This effectively addresses the scarcity of second language speech for non-native speakers. Furthermore, the expanded speech, with its first accent, is more realistic.
[0104] In summary, some embodiments of the present application provide a data augmentation method that obtains multiple first speech sounds, where the first speech sounds are speech sounds in a first language with a first accent, the first accent being the native accent of the first language, and the first speech sounds are annotated with text labels. A preset speech conversion network is trained using the multiple first speech sounds. The trained speech conversion model then performs speech conversion on multiple second speech sounds to obtain multiple augmented speech sounds. The second speech sounds are speech sounds in a second language with a second accent, the second accent being the native accent of the second language, and the augmented speech sounds are speech sounds in the second language with the first accent. For example, the first speech sounds are speech data of a native Chinese speaker speaking Chinese, and the second speech sounds are speech data of a native English speaker speaking English. The preset speech conversion network is trained using the multiple first speech sounds, and the trained speech conversion model learns the first accent. Thus, the speech conversion can be performed on the second speech sounds, resulting in augmented speech sounds with the first accent, for example, speech data that mimics speech data of a native Chinese speaker speaking English. That is, the expanded speech has the accent of the original speech (the collected real speech), is close to the real data, can effectively realize the data augmentation for the speech recognition model, and is conducive to improving the accuracy of the model.
[0105] Some embodiments of the present application also provide a method for training a speech recognition model. Figure 5 , the method S200 includes:
[0106] S201: Augment the training set.
[0107] Here, the training set is data used to train the speech recognition network. In some embodiments, the training set includes multiple first speech sounds described in the above embodiments. The first speech sounds are speech sounds in a first language with a first accent, where the first accent is the native accent of the first language. Each first speech sound is annotated with a text label. In other embodiments, the training set may also include several other first speech sounds.
[0108] The training set is augmented using the data augmentation method described in the above embodiment. Specifically, the speech conversion model described in the above embodiment is used to perform speech conversion on the collected multiple second speech sounds to obtain multiple augmented speech sounds. The second speech sounds are speech sounds in a second language with a second accent, where the second accent is the native accent of the second language, and the second speech sounds are annotated with text labels. The augmented speech sounds are speech sounds in the second language with a first accent.
[0109] The text label of the expanded speech is the text label of the second speech. Multiple expanded speech are added to the training set to obtain an augmented training set.
[0110] S202: Iteratively train the speech recognition network using the augmented training set until convergence, thereby obtaining a speech recognition model.
[0111] It can be understood that the speech recognition network predicts the text label of the input speech, and adjusts the network parameters based on the loss between the predicted text label and the real text label until the loss is minimized or fluctuates within a certain range, and the network converges. The parameters at this time are used as model parameters and configured to the speech recognition network to obtain the speech recognition model.
[0112] Because the augmented training set includes multilingual speech (e.g., a first speech in a first language, and an expanded speech in a second language), and these speech texts share the same accent (e.g., the first accent), it closely resembles real-world data. Consequently, the model can learn the relationship between multilingual speech texts with the same accent, resulting in more accurate recognition performance for target groups with the first accent. For example, for a target group whose native language is Chinese, the model can accurately recognize not only Chinese speech but also English speech spoken by that target group.
[0113] Some embodiments of the present application also provide a speech recognition method, see Figure 6 , the method S300 includes:
[0114] S301: Acquire voice data.
[0115] The terminal (such as a smart phone) has a built-in voice recognition assistant (application), and the voice recognition model is encapsulated in the voice recognition assistant. After the processor of the terminal obtains the voice data, it is provided to the voice recognition assistant. The voice data can be collected by the microphone on the terminal, or it can be collected by other devices with microphones and sent to the terminal. No limitation is imposed on the method of obtaining voice data here. It is understandable that the voice data may be a voice with a first accent in a first language, or it may be a voice with a first accent in a second language. For example, it can be Chinese speech spoken by a native Chinese speaker, or it can be English speech spoken by a native Chinese speaker.
[0116] S302: Input the speech data into the speech recognition model to obtain the recognized text.
[0117] The speech recognition model is trained using any of the aforementioned methods for training a speech recognition model. After acquiring speech data, the terminal inputs the speech data into the speech recognition model, performs speech recognition, and outputs recognized text. The recognized text is displayed on the terminal's display screen for the user to view or use.
[0118] It can be understood that the speech recognition model is trained by the method of training the speech recognition model in the above embodiment, and has the same structure and function as the speech recognition model in the above embodiment, which will not be described in detail here.
[0119] The present application also provides a computer-readable storage medium, which stores computer-executable instructions. The computer-executable instructions are used to enable an electronic device to execute the method provided in the present application, for example, Figure 3-4 The data augmentation methods shown, such as Figure 5 The method of training the speech recognition model shown, such as Figure 6 The speech recognition method is shown.
[0120] In some embodiments, the storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EE PROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.
[0121] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0122] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0123] As an example, executable instructions may be deployed to be executed on one computing device (including devices such as smart terminals and servers), or on multiple computing devices located in one location, or on multiple computing devices distributed in multiple locations and interconnected by a communication network.
[0124] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the data augmentation method, the method for training a speech recognition model, or the speech recognition method in the aforementioned embodiments.
[0125] An embodiment of the present application also provides a computer program product, including a computer program stored on a non-volatile computer-readable storage medium, wherein the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the data augmentation method, the method for training a speech recognition model, or the speech recognition method in the above embodiments.
[0126] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course by hardware. Those skilled in the art can understand that all or part of the processes in the above embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Based on the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present application as described above. For the sake of simplicity, they are not provided in detail. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A data augmentation method, characterized in that: include: Acquire multiple first speech sounds, where the first speech sounds are speech sounds in a first language with a first accent, the first accent being a native accent of the first language, and the first speech sounds are annotated with text labels; The plurality of first voices are used to train a preset voice conversion network, and the trained voice conversion model is used to perform voice conversion on the plurality of second voices to obtain a plurality of expanded voices; wherein the second voice is a voice in a second language with a second accent, the second accent is the native accent of the second language, and the expanded voice is a voice in the second language with the first accent.
2. The method according to claim 1, characterized in that The speech conversion model includes a speech feature extraction module and a vocoder, wherein the speech feature extraction module is used to extract speech features from input speech, and the vocoder is used to convert the speech features into audio data.
3. The method according to claim 2, characterized in that The speech feature extraction module includes a content extraction module, a speaker information extraction module, and an acoustic module; Among them, the content extraction module is used to extract content features in the input speech; the speaker information extraction module is used to extract speaker information in the input speech; and the acoustic module is used to fuse the content features with the speaker information to obtain the speech features.
4. The method according to claim 3, characterized in that The content extraction module includes multiple convolutional layers, and the step size of each convolutional layer is 1.
5. The method according to claim 1, wherein The trained speech conversion model performs speech conversion on the plurality of second speech sounds to obtain a plurality of expanded speech sounds, including: For any second speech, the second speech and speaker information are input into the speech conversion model, and an expanded speech is output.
6. The method according to claim 5, characterized in that The method further comprises: constructing a speaker Gaussian mixture model using the speaker information extracted from the plurality of first voices by the speaker information extraction module in the voice conversion model; Using the speaker Gaussian mixture model to sample and output virtual speaker information; The step of inputting the second speech and speaker information into the speech conversion model and outputting an expanded speech comprises: The second speech and the virtual speaker information are input into the speech conversion model to obtain the expanded speech.
7. The method according to claim 6, characterized in that The method of constructing a speaker Gaussian mixture model by using the speaker information extracted from the plurality of first voices by the speaker information extraction module in the voice conversion model comprises: Grouping the speaker information corresponding to the multiple first voices according to multiple feature dimensions to obtain multiple groups of speaker information; A corresponding speaker Gaussian mixture model is constructed for each group of speaker information.
8. A method for training a speech recognition model, characterized in that: include: Augmenting the training set using the method described in any one of claims 1 to 7; The speech recognition network is iteratively trained using the augmented training set until convergence, thereby obtaining the speech recognition model.
9. A speech recognition method, characterized in that: include: Get voice data; The speech data is input into the speech recognition model as claimed in claim 8 to obtain a recognition text.
10. An electronic device, characterized in that: include: at least one processor, and a memory communicatively coupled to the at least one processor, wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer device to execute the method according to any one of claims 1 to 9.