Speech recognition model training method, device, electronic device and storage medium

By training a speech recognition model containing acoustic sub-models and transformation sub-models, and using generalization processing and error update technology, the problem of poor speech recognition robustness in the existing technology is solved, and higher speech recognition accuracy is achieved.

CN115116443BActive Publication Date: 2025-05-13WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110287757.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-17
Publication Date
2025-05-13
Estimated Expiration
2041-03-17

AI Technical Summary

Technical Problem

The existing speech recognition technology is poor in the face of environmental noise, accent differences between different users and emotional fluctuations of speakers, resulting in low accuracy of speech recognition.

Method used

By training a speech recognition model including acoustic sub-model and conversion sub-model, generalization processing technology is used to process the original audio samples, multiple audio samples carrying the same text tags are generated, phoneme prediction and text conversion are performed separately, and model parameters are updated based on errors to improve the robustness of the model.

Benefits of technology

Through generalization processing and model update technology, the robustness of the speech recognition model is improved, and the ability to adapt to environmental noise, accent differences and mood fluctuations is enhanced, thereby improving the accuracy of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116443B_ABST
    Figure CN115116443B_ABST
Patent Text Reader

Abstract

The present application provides a training method, device, electronic device, computer-readable storage medium and computer program product for a speech recognition model; the method includes: obtaining an original audio sample, the original audio sample carries a first text label; generalizing the original audio sample to obtain a plurality of audio samples corresponding to the original audio sample and carrying the first text label; performing phoneme prediction on each of the audio samples through the acoustic sub-model to obtain a phoneme sequence corresponding to each of the audio samples; performing text conversion on each of the phoneme sequences through the conversion sub-model to obtain a converted text corresponding to each of the audio samples; obtaining the error between each of the converted texts and the first text label, and updating the model parameters of the speech recognition model based on the obtained error. Through the present application, a speech recognition model with strong robustness can be trained to improve the accuracy of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to speech recognition technology, and in particular to a training method, device, electronic device and storage medium for a speech recognition model. Background Art

[0002] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0003] With the development of smart devices and Automatic Speech Recognition (ASR) technology, the application scenarios of speech recognition are increasing. However, speech recognition technology is challenged by various changing conditions in actual applications, such as environmental noise, different accents of different users, and pronunciation changes caused by emotional fluctuations of speakers. However, the speech recognition robustness of related technologies is very poor, and all of the above factors will affect the accuracy of speech recognition. Summary of the invention

[0004] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium and computer program product for training a speech recognition model, which can train a speech recognition model with strong robustness and improve the accuracy of speech recognition.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present application embodiment provides a method for training a speech recognition model, wherein the speech recognition model includes an acoustic sub-model and a conversion sub-model, and the method includes:

[0007] Acquire an original audio sample, where the original audio sample carries a first text label;

[0008] Performing generalization processing on the original audio sample to obtain a plurality of audio samples corresponding to the original audio sample and carrying the first text label;

[0009] Performing phoneme prediction on each of the audio samples using the acoustic sub-model to obtain a phoneme sequence corresponding to each of the audio samples;

[0010] Performing text conversion on each of the phoneme sequences respectively through the conversion sub-model to obtain a conversion text corresponding to each of the audio samples;

[0011] The errors between each of the converted texts and the first text label are respectively obtained, and the model parameters of the speech recognition model are updated based on the obtained errors.

[0012] The embodiment of the present application provides a training device for a speech recognition model, wherein the speech recognition model includes an acoustic sub-model and a conversion sub-model, and the device includes:

[0013] An acquisition module, configured to acquire an original audio sample, wherein the original audio sample carries a first text label;

[0014] A generalization module, configured to perform generalization processing on the original audio sample to obtain a plurality of audio samples corresponding to the original audio sample and carrying the first text label;

[0015] A phoneme prediction module, used to perform phoneme prediction on each of the audio samples using the acoustic sub-model to obtain a phoneme sequence corresponding to each of the audio samples;

[0016] A text conversion module, used to perform text conversion on each of the phoneme sequences through the conversion sub-model to obtain a conversion text corresponding to each of the audio samples;

[0017] An updating module is used to respectively obtain the error between each of the converted texts and the first text label, and update the model parameters of the speech recognition model based on the obtained error.

[0018] In the above scheme, the training device of the speech recognition model also includes: a pre-training module, which is used to perform phoneme conversion on the second text label based on the mapping relationship between text and phonemes to obtain a standard phoneme sequence corresponding to the second text label; perform text conversion on the standard phoneme sequence through the conversion sub-model to obtain a corresponding target conversion text; based on the error between the target conversion text and the second text label, update the model parameters of the conversion sub-model to obtain an updated conversion sub-model; accordingly, the text conversion module is also used to perform text conversion on each of the phoneme sequences respectively through the updated conversion sub-model.

[0019] In the above scheme, the pre-training module is also used to generalize the standard phoneme sequence to obtain a corresponding plurality of deviation phoneme sequences; correspondingly, the text conversion module is also used to perform text conversion on each of the phoneme sequences and each of the deviation phoneme sequences respectively through the updated conversion sub-model.

[0020] In the above solution, the generalization module is further used to obtain multiple interference information; and based on each interference information, perform the following processing: adding the interference information to the original audio sample.

[0021] In the above scheme, the generalization module is also used to perform the following processing on the original audio sample multiple times: changing at least one frame of speech signal of the original audio sample, each frame of the speech signal corresponds to a phoneme; wherein the change includes at least one of the following: phoneme deletion, phoneme insertion and phoneme replacement.

[0022] In the above scheme, the text conversion module is also used to perform the following processing on each of the phoneme sequences through the conversion sub-model: extract semantic features from the phoneme sequence to obtain corresponding semantic features; based on the semantic features, perform text conversion on the phoneme sequence to obtain multiple candidate word sequences corresponding to the phoneme sequence, and scores corresponding to each of the candidate word sequences; and select the candidate word sequence with the highest score from the multiple candidate word sequences as the converted text corresponding to the audio sample.

[0023] In the above scheme, the text conversion module is also used to perform text conversion on the phoneme sequence based on the semantic features to obtain multiple candidate word sequences corresponding to the phoneme sequence and the conditional probability of each candidate word in each of the candidate word sequences; based on the conditional probability of each candidate word in the candidate word sequence, the score corresponding to each of the candidate word sequences is determined respectively.

[0024] In the above scheme, the training device of the speech recognition model also includes: a speech recognition module, used to obtain the audio to be recognized; perform phoneme prediction on the audio to be recognized through the acoustic sub-model to obtain a target phoneme sequence corresponding to the audio to be recognized; perform text conversion on the target phoneme sequence through the conversion sub-model to obtain a converted text corresponding to the audio to be recognized.

[0025] In the above scheme, the conversion sub-model includes a pre-conversion sub-model and a retyping molecular model, and the speech recognition module is also used to perform text conversion on the target phoneme sequence through the pre-conversion sub-model to obtain multiple candidate texts corresponding to the audio to be recognized and a first score for each candidate text; through the retyping molecular model, score prediction is performed on each candidate text to obtain a corresponding second score; based on the first score and the second score, the target score of each candidate text is determined; and the candidate text with the highest target score is selected from the multiple candidate texts as the conversion text corresponding to the audio to be recognized.

[0026] An embodiment of the present application provides an electronic device, including:

[0027] A memory for storing executable instructions;

[0028] The processor is used to implement the training method of the speech recognition model provided in the embodiment of the present application when executing the executable instructions stored in the memory.

[0029] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to execute and implement the training method of the speech recognition model provided in the embodiment of the present application.

[0030] An embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the training method of the speech recognition model provided in the embodiment of the present application.

[0031] The embodiments of the present application have the following beneficial effects:

[0032] In an embodiment of the present application, the original audio sample is generalized to obtain a corresponding plurality of audio samples carrying the first text label, and each of the audio samples is predicted by an acoustic sub-model to obtain a phoneme sequence corresponding to each audio sample, and then each phoneme sequence is converted into text by a conversion sub-model to obtain a converted text corresponding to each audio sample, and the model parameters of the speech recognition model are updated based on the error between each converted text and the first text label. By training the model with the generalized audio samples, the model can have a certain error correction capability, thereby making the model more robust, overcoming the defect of low speech recognition accuracy in related technologies and improving the accuracy of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is an optional structural diagram of a training system for a speech recognition model provided in an embodiment of the present application;

[0034] Figure 2 is an optional structural diagram of an electronic device provided in an embodiment of the present application;

[0035] Figure 3 It is an optional structural diagram of the speech recognition model provided in the embodiment of the present application;

[0036] Figure 4 It is an optional flowchart of the training method of the speech recognition model provided in the embodiment of the present application;

[0037] Figure 5 It is an optional flowchart of the training method of the speech recognition model provided in the embodiment of the present application;

[0038] Figure 6It is an optional flowchart of the training method of the speech recognition model provided in the embodiment of the present application;

[0039] Figure 7 It is an optional flowchart of the training method of the speech recognition model provided in the embodiment of the present application;

[0040] Figure 8 It is an optional structural diagram of the speech recognition model provided in the embodiment of the present application;

[0041] Fig. 9 It is an optional flowchart of the training method of the speech recognition model provided in the embodiment of the present application;

[0042] Fig.10 It is an optional flowchart of the training method of the speech recognition model provided in the embodiment of the present application;

[0043] Fig.11 This is an optional flowchart of the speech recognition process provided by the embodiment of the present application;

[0044] Fig.12 It is an optional structural diagram of a training model of the speech recognition model provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.

[0046] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0047] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0049] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0050] 1) Acoustic model (AM), which represents the differentiated knowledge of acoustics, phonetics, environmental variables, speaker gender, and accent, including acoustic models based on hidden Markov model (HMM), such as mixed Gaussian-hidden Markov model (GMM-HMM) and deep neural network-hidden Markov model (DNN-HMM). In addition, acoustic models also include end-to-end acoustic models, such as connectionist temporal classification-long short-term memory (CTC-LSTM) model and attention model.

[0051] Each state of the acoustic model represents the probability distribution of the speech features of a speech unit (such as a word, syllable, or phoneme) in that state, and is connected into an ordered sequence of states through transitions between states, thus obtaining a sequence of speech units represented by a speech signal.

[0052] It should be understood that the acoustic sub-model in the embodiment of the present application is the acoustic model.

[0053] 2) Language Model (LM) is the knowledge representation of language structure. Here, language structure can include the rules between words and sentences, such as grammar, common collocation of words, etc.

[0054] For a text sequence, the task of the language model is to calculate the probability distribution of the sequence, which can be generally explained as determining whether a language sequence is a normal sentence.

[0055] It should be noted that the transformer sub-model in the embodiment of the present application is a language model, which can convert text in combination with the context information of the phoneme sequence, and convert the phoneme sequence into a converted text that conforms to the language logic.

[0056] 3) Pronunciation dictionary, which records the mapping relationship between text and phonemes.

[0057] 4) Phoneme is the smallest unit of speech divided according to the natural properties of speech. It is analyzed based on the pronunciation actions in the syllable, and one action constitutes a phoneme.

[0058] 5) Phoneme sequence is a sequence of multiple phonemes arranged in a certain order.

[0059] For example, for the word "I", it includes two phonemes, "w" and "o3", and the phoneme series obtained by sorting them according to their pronunciation is "w o3".

[0060] 6) Standard phoneme sequence refers to the phoneme sequence corresponding to the correct pronunciation of a specific phrase.

[0061] For example, for the specific phrase “I am in Guiyang”, the corresponding standard phoneme sequence is “w o3 zai4 g ui4 y ang2”.

[0062] 7) Deviation phoneme sequence refers to the phoneme sequence corresponding to the incorrect pronunciation of a specific phrase.

[0063] For example, for the specific phrase "I am in Guiyang", the corresponding deviation phoneme sequence may be "w o3 zai4 g ui4 l v2", etc. It should be noted that since there are many forms of incorrect pronunciation, the deviation phoneme sequence for a specific phrase also has many deviation forms, and only one of them is listed here.

[0064] 8) Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, mechatronics, and other technologies. Artificial intelligence software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0065] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0066] Based on this, the embodiments of the present application provide a method, device, electronic device, computer-readable storage medium and computer program product for training a speech recognition model, which can provide robustness of the speech recognition model.

[0067] First, the training system of the speech recognition model provided in the embodiment of the present application is described. Figure 1 , Figure 1It is an optional architecture diagram of the training system 100 of the speech recognition model provided in the embodiment of the present application. The terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, and a wireless link is used to realize data transmission. In some embodiments, the terminal 400 can be a laptop, a tablet computer, a desktop computer, a smart phone, a dedicated messaging device, a portable gaming device, a smart speaker, a smart watch, etc., but is not limited to this. The server 200 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The network 300 can be a wide area network or a local area network, or a combination of the two. The terminal 400 and the server 200 can be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiment of the present application.

[0068] The terminal 400 is used to obtain the original audio sample, generate a model training instruction carrying the original audio sample based on the original audio sample, and send the model training instruction to the server 200.

[0069] The server 200 is used to perform generalization processing on the original audio sample to obtain multiple audio samples carrying the first text label corresponding to the original audio sample; perform phoneme prediction on each of the audio samples through the acoustic sub-model to obtain a phoneme sequence corresponding to each of the audio samples; perform text conversion on each of the phoneme sequences through the conversion sub-model to obtain a converted text corresponding to each of the audio samples; obtain the error between each of the converted texts and the first text label, and update the model parameters of the speech recognition model based on the obtained error to obtain a trained speech recognition model; and send the trained speech recognition model to the terminal 400.

[0070] The terminal 400 is also used to obtain the audio to be recognized, perform speech recognition on the audio to be recognized through the trained speech recognition model, obtain the converted text corresponding to the audio to be recognized, and output the converted text.

[0071] See also Figure 2 , Figure 2 is an optional structural diagram of an electronic device 500 provided in an embodiment of the present application. In practical applications, the electronic device 500 can be implemented as Figure 1 The terminal 400 or the server 200 in the embodiment of the present invention is an electronic device. Figure 1 Taking the server 200 shown as an example, an electronic device for implementing the training method of the speech recognition model of the embodiment of the present application is described. Figure 2 The electronic device 500 shown includes: at least one processor 510, a memory 550, at least one network interface 520 and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It can be understood that the bus system 540 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 540 is not described in detail. Figure 2 Various buses are labeled as bus system 540 .

[0072] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0073] The user interface 530 includes one or more output devices 531 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0074] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 510.

[0075] The memory 550 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.

[0076] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.

[0077] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0078] A network communication module 552, for reaching other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB);

[0079] a presentation module 553 for enabling presentation of information via one or more output devices 531 (e.g., display screen, speaker, etc.) associated with the user interface 530 (e.g., a user interface for operating peripherals and displaying content and information);

[0080] The input processing module 554 is used to detect one or more user inputs or interactions from one of the one or more input devices 532 and translate the detected inputs or interactions.

[0081] In some embodiments, the training device for the speech recognition model provided in the embodiments of the present application can be implemented in software. Figure 2 The training device 555 of the speech recognition model stored in the memory 550 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: an acquisition module 5551, a generalization module 5552, a phoneme prediction module 5553, a text conversion module 5554 and an update module 5555. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.

[0082] In other embodiments, the training device of the speech recognition model provided in the embodiments of the present application can be implemented in hardware. As an example, the training device of the speech recognition model provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the speech recognition model provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.

[0083] See also Figure 3 , Figure 3It is an optional structural diagram of the speech recognition model provided in the embodiment of the present application. The speech recognition model provided in the embodiment of the present application includes an acoustic sub-model and a conversion sub-model. Among them, the acoustic sub-model is used to predict the phonemes of the input audio and output the corresponding predicted phoneme sequence; the conversion sub-model is used to convert the input phoneme sequence into text and output the corresponding converted text.

[0084] The following will describe the training method of the speech recognition model provided by the embodiment of the present application in combination with the exemplary application and implementation of the terminal provided by the embodiment of the present application. Figure 4 , Figure 4 is an optional flow chart of the training method of the speech recognition model provided in the embodiment of the present application, which will be combined with Figure 4 The steps shown are explained.

[0085] Step 101: The server obtains an original audio sample, where the original audio sample carries a first text tag.

[0086] Here, the original audio sample can be obtained from a web page by accessing the web page. Specifically, the server can access the relevant audio library web page and download the audio with the first text tag in the audio library of the web page. The original audio sample can also be obtained by manually recording the first text tag. The embodiment of the present application does not specifically limit the source of the original audio sample.

[0087] Step 102: generalize the original audio sample to obtain a plurality of audio samples corresponding to the original audio sample and carrying the first text label.

[0088] In actual implementation, the server performs generalization processing on the original audio sample carrying the first text label to obtain multiple audio samples. Here, the phoneme sequence corresponding to the original audio sample is recorded as the original phoneme sequence, and the text corresponding to the original phoneme sequence is the text corresponding to the first text label. It should be noted that the phoneme sequence corresponding to the audio sample obtained after the generalization processing of the original audio sample may not match the original phoneme sequence, that is, the phoneme sequence corresponding to the audio sample may have missing phonemes, wrong phonemes, or increased phonemes compared to the original phoneme sequence. The text label carried by the audio sample is still the first text label corresponding to the original phoneme sequence. There may also be interference information in the audio sample that does not exist in the original audio sample, such as noise. By generalizing the original audio sample, a plurality of audio samples that do not completely match the text content corresponding to the first text label are obtained, and these audio samples are used to train the speech recognition model, which can improve the robustness of the speech recognition model.

[0089] In some embodiments, based on Figure 4Step 102 may also be implemented in the following manner: the server obtains multiple interference information; and based on each interference information, performs the following processing: adding the interference information to the original audio sample.

[0090] Here, the interference information can be different types of noise, different types of sound effects, etc. Here, different types of noise can include but are not limited to acoustic noise and electrical noise. Among them, acoustic noise is an irregular and chaotic combination of sounds of different frequencies and sound intensities, such as various types of environmental noise. Electrical noise is the noise of the electronic circuits of the audio-visual equipment itself, the interference sound of the power supply AC signal, the interference of the stray electromagnetic field in the space, etc., such as white noise. Different types of sound effects include but are not limited to action sound effects and environmental sound effects. Here, environmental sound effects can be, for example, far-field sound effects, and audio samples with far-field sound effects can be obtained by far-field synthesis of the original audio samples.

[0091] In actual implementation, the server obtains multiple interference information and adds each interference information to the original audio sample to perform a corresponding generalization process on the original audio sample to obtain a generalized audio sample. It should be understood that the generalized audio sample carries the first text label.

[0092] In some embodiments, based on Figure 4 Step 102 can also be implemented in the following manner: the server performs the following processing on the original audio sample multiple times: changing at least one frame of speech signal of the original audio sample, each frame of the speech signal corresponds to a phoneme; wherein the change includes at least one of the following: phoneme deletion, phoneme insertion and phoneme replacement.

[0093] In actual implementation, the server generalizes the original audio sample by changing the phonemes in the original audio sample. In actual implementation, the server changes the phonemes of the original audio sample by changing the voice signal of the corresponding phoneme. Here, one frame of voice signal corresponds to one phoneme. Specifically, the server can delete, insert or replace the voice signal at any position of the original audio sample to perform corresponding phoneme deletion, phoneme insertion or phoneme replacement. It should be noted that here, if the number of frames of the voice signal corresponding to the original audio sample is recorded as the original number of frames, the frame ratio of the number of voice signal frames changed to the original audio sample to the original number of frames is set to a suitable range, such as within 10%, so as to avoid changing too many voice signal frames of the original audio sample and changing the original audio sample to other content that is irrelevant to the first text label, thereby interfering with the speech recognition accuracy of the speech recognition model trained based on the audio sample carrying the first text label.

[0094] Exemplarily, if the first text tag is "I am in Guiyang," and the original phoneme sequence corresponding to the original audio sample is "w o3 z ai4 g ui4 y ang2," then the audio sample obtained after performing voice signal deletion on the original audio sample to perform phoneme deletion may be "w o3 g ui4 y ang1," and the audio sample obtained after performing voice signal insertion on the original audio sample to perform phoneme insertion may be "w o3 z ai4 z ao4 yi n1 g ui4 y ang2," and the audio sample obtained after performing voice signal replacement on the original audio sample to perform phoneme replacement may be "w o3 z ai4 g ui4 lv2," and so on.

[0095] In an embodiment of the present application, the original audio sample is generalized by deleting, inserting or replacing phonemes in the original audio sample to obtain an audio sample with some phonemes changed, so that the robustness of the speech recognition model trained based on the audio sample is significantly improved.

[0096] In some embodiments, the server can also perform generalization processing on the original audio sample by adding interference information and changing the voice signal at the same time, so as to obtain an audio sample carrying interference information and an audio sample after the phoneme is changed. The server can also perform addition of interference information and change of the voice signal on the original audio sample at the same time, so as to obtain an audio sample carrying interference information and a changed phoneme. By generalizing the original audio sample in different ways, the audio samples obtained after generalization are more diverse, so that the trained speech recognition model is more robust.

[0097] Step 103: perform phoneme prediction on each of the audio samples using the acoustic sub-model to obtain a phoneme sequence corresponding to each of the audio samples.

[0098] In actual implementation, the server uses the acoustic sub-model to perform phoneme prediction for each audio sample and obtains the phoneme sequence corresponding to each audio sample. It should be understood that since the audio sample is generalized compared to the original audio sample, such as carrying interference information or the phoneme sequence does not match the original phoneme sequence, it can be understood that if the audio sample carries interference information, the interference information will interfere with the phoneme prediction of the acoustic sub-model, and the acoustic sub-model will predict the audio sample to obtain interference phonemes that do not match the original phoneme sequence; if the audio sample is obtained by changing the speech signal frame of the original audio sample, the phoneme sequence actually corresponding to the audio sample does not match the original phoneme sequence, and the acoustic sub-model will predict a phoneme sequence that does not match the original phoneme. In other words, the phoneme sequence obtained by the acoustic sub-model's phoneme prediction of the audio sample may not be completely consistent with the original phoneme sequence corresponding to the original audio sample.

[0099] Step 104: Perform text conversion on each of the phoneme sequences using the conversion sub-model to obtain a converted text corresponding to each of the audio samples.

[0100] In actual implementation, the server inputs the phoneme sequence output by the acoustic sub-model into the conversion sub-model, and performs text conversion on the phoneme sequence through the conversion sub-model to obtain the converted text of the corresponding audio sample. Here, the conversion sub-model performs text conversion on multiple phoneme sequences output by the acoustic sub-model respectively to obtain the converted text corresponding to each audio sample. In the embodiment of the present application, the server can perform text conversion on each phoneme sequence in sequence through the conversion sub-model in a serial manner, and can also perform text conversion on each phoneme sequence simultaneously in a parallel manner.

[0101] In some embodiments, reference Figure 5 , Figure 5 FIG. 1 is an optional flow chart of the method for training a speech recognition model provided in an embodiment of the present application. Figure 4 Before step 104, you may also execute:

[0102] Step 201: The server performs phoneme conversion on the second text tag based on the mapping relationship between text and phonemes to obtain a standard phoneme sequence corresponding to the second text tag.

[0103] It should be noted that Figure 5 Steps 201 to 203 are shown as being executed before step 101. In the embodiment of the present application, steps 201 to 203 may also be executed between any steps before step 104, or executed in parallel with steps 101 to 103, etc. Figure 5 This is just an example of one possible order of execution.

[0104] Here, the second text label is used to pre-train the conversion sub-model, which can be the same as the first text label or different from the first text label. In the embodiment of the present application, the mapping relationship between text and phoneme is the mapping relationship between word and phoneme. For example, for the word "I", the phoneme with which it is mapped is "w o3". In some embodiments, the server obtains a pronunciation dictionary that records the mapping relationship between text and phoneme, takes the second text label as an index, obtains the phonemes corresponding to each word in the second text label by querying the pronunciation dictionary, and then sorts each phoneme according to the order of each word in the second text label to obtain a standard phoneme sequence corresponding to the second text label. Specifically, the server performs word segmentation processing on the second text label to obtain multiple words corresponding to the second text label, and queries the pronunciation dictionary for each word respectively to obtain the phonemes corresponding to each word, and then sorts the phonemes corresponding to each word according to the order of each word in the second text label to obtain the corresponding standard phoneme sequence.

[0105] For example, if the second text tag is "I am in Guiyang," the server performs word segmentation processing on it to obtain three words, namely "I," "in" and "Guiyang." Then, the server queries the mapping relationship between words and phonemes, and obtains the phonemes "w o3," "z ai4," and "g ui4 y ang2" corresponding to these three words respectively. Then, the server sorts and combines these phonemes according to the order of the three words in the second text tag to obtain the corresponding standard phoneme sequence "w o3 z ai4 g ui4 yang2."

[0106] Step 202: Perform text conversion on the standard phoneme sequence through the conversion sub-model to obtain a corresponding target conversion text.

[0107] In actual implementation, since the conversion sub-model is a language model, when performing text conversion, it is converted into a target conversion text with contextual semantics by combining the context information of the standard phoneme sequence.

[0108] Step 203: based on the error between the target conversion text and the second text label, update the model parameters of the conversion sub-model to obtain an updated conversion sub-model.

[0109] In actual implementation, the server updates the model parameters of the conversion sub-model through multiple iterative training until the convergence condition of the conversion sub-model is reached or the number of iterations reaches the iteration threshold, stops training the conversion sub-model, and obtains an updated conversion sub-model. It should be noted that the multiple iterative training of the conversion sub-model here is based on the standard phoneme sequence.

[0110] In actual implementation, the server can train the conversion sub-model in the following ways:

[0111] The server determines the error between the target conversion text and the second text label by calculating the value of the loss function. When the value of the loss function reaches a threshold, the server determines a corresponding error signal based on the loss function, back-propagates the error signal in the conversion sub-model, and updates the model parameters of each layer of the conversion sub-model during the propagation process.

[0112] Here we explain the back propagation. The training samples are input into the input layer of the neural network model, through the hidden layer, and finally reach the output layer and output the result. This is the forward propagation process of the neural network model. Since there is an error between the output result of the neural network model and the actual result, the error between the output result and the actual value is calculated, and the error is back-propagated from the output layer to the hidden layer until it propagates to the input layer. During the back propagation process, the value of the model parameter is adjusted according to the error; the above process is continuously iterated until convergence. Taking the loss function as an example, the server determines the error signal based on the loss function. The error signal is back-propagated from the output layer of the neural network model, and the error signal is back-propagated layer by layer. When the error signal reaches each layer, the gradient (that is, the partial derivative of the Loss function with respect to the parameters of this layer) is solved in combination with the conducted error signal, and the parameters of this layer are updated to the corresponding gradient value.

[0113] Correspondingly, step 104 may also be implemented in the following manner: the server performs text conversion on each of the phoneme sequences using the updated conversion sub-model to obtain a conversion text corresponding to each of the audio samples.

[0114] In actual implementation, the conversion sub-model is trained for the first round based on the standard phoneme sequence to obtain an updated conversion sub-model, so that the conversion sub-model learns the mapping relationship between the standard text and the phoneme sequence to obtain the updated conversion sub-model. Then, the updated conversion sub-model is trained for the second round based on the phoneme sequences output by the acoustic sub-model. Here, each phoneme sequence output by the acoustic sub-model has a certain deviation compared with the original phoneme sequence corresponding to the original audio sample. The second round of training of the conversion sub-model using these deviated phoneme sequences carrying the first text label can improve the robustness of the conversion sub-model, so that it can convert the deviated phoneme sequence into correct text.

[0115] In some embodiments, reference Figure 6 , Figure 6 is an optional flow chart of the training method of the speech recognition model provided in the embodiment of the present application, based on Figure 5 Before step 104, the following may also be performed: step 301, the server generalizes the standard phoneme sequence to obtain a corresponding plurality of deviation phoneme sequences.

[0116] In actual implementation, the server may generalize the standard phoneme sequence by randomly removing at least one phoneme in the standard phoneme sequence, randomly replacing at least one phoneme in the standard phoneme sequence, or randomly inserting at least one phoneme sequence into the standard phoneme sequence to obtain a corresponding deviation phoneme sequence. Exemplarily, if the standard phoneme sequence is "w o3 z ai4 g ui4 y ang2," the deviation phoneme sequence obtained after generalization may be "w o3 g ui4 y ang1," "w o3 z ai4 g ui4 l v2," or "w o3 z ai4 z ao4 y in1 g ui4 y ang2," etc. In actual implementation, the server may also perform a variety of generalization processes on the same standard phoneme sequence, such as simultaneously performing phoneme deletion and phoneme replacement to obtain a deviation phoneme sequence "w o3 g ui4 l v2," etc. It should be noted that generalization processes performed by any generalization method and a combination of a variety of generalization methods are all within the protection scope of the embodiments of the present application.

[0117] Correspondingly, step 104 may also be implemented in the following manner: the server performs text conversion on each of the phoneme sequences and each of the deviation phoneme sequences respectively through the updated conversion sub-model.

[0118] It should be noted that the multiple phoneme sequences output by the acoustic sub-model and the multiple deviation phoneme sequences obtained after generalization of the standard phoneme sequence all include phoneme sequences with deviations, and of course, also include standard phoneme sequences that match the first text label. The server constructs a corresponding training sample set based on the multiple phoneme sequences output by the acoustic sub-model and the multiple deviation phoneme sequences obtained after generalization of the standard phoneme sequence, and the phoneme sequences in the training sample set all carry the first text label. Then, the server inputs the training sample set into the updated conversion sub-model and performs a second round of training on the updated conversion sub-model.

[0119] In an embodiment of the present application, a second round of training is performed on the updated conversion sub-model in combination with a phoneme sequence with a deviation. After the conversion sub-model learns the mapping relationship between the standard phoneme sequence and the text, the robustness of the conversion sub-model is improved so that it can convert the phoneme sequence with a deviation into the correct text.

[0120] In some embodiments, based on Figure 4, step 104 can also be implemented in the following manner: the server performs the following processing on each of the phoneme sequences through the conversion sub-model: extracting semantic features from the phoneme sequence to obtain corresponding semantic features; based on the semantic features, performing text conversion on the phoneme sequence to obtain multiple candidate word sequences corresponding to the phoneme sequence and scores corresponding to each of the candidate word sequences; and selecting the candidate word sequence with the highest score from the multiple candidate word sequences as the converted text corresponding to the audio sample.

[0121] In actual implementation, when the conversion sub-model converts the phoneme sequence into text, it first extracts the semantic features of the phoneme sequence to obtain the corresponding semantic features. Specifically, the server encodes the phoneme sequence to obtain the phoneme vector of the phoneme sequence, and then extracts the semantic features of the phoneme sequence based on the phoneme vector. Here, the semantic features are represented by vectors. In actual implementation, the server converts the phoneme sequence into text based on the semantic features of the phoneme sequence to obtain multiple candidate word sequences corresponding to the phoneme sequence and the scores of each candidate word sequence. It should be noted that the scores of the candidate word sequences are obtained by the conversion sub-model performing probability calculations on the candidate word sequences.

[0122] In some embodiments, the text conversion of the phoneme sequence is performed based on the semantic feature to obtain multiple candidate word sequences corresponding to the phoneme sequence and scores corresponding to each of the candidate word sequences, including: based on the semantic feature, the text conversion of the phoneme sequence is performed to obtain multiple candidate word sequences corresponding to the phoneme sequence and the conditional probability of each candidate word in each of the candidate word sequences; based on the conditional probability of each candidate word in the candidate word sequence, the scores corresponding to each of the candidate word sequences are determined respectively.

[0123] First, the conversion sub-model combines the context of each candidate word in the candidate word sequence to determine the conditional probability of each candidate word appearing in the candidate word sequence, and determines the fluency of the candidate word sequence based on the conditional probability of each candidate word. For example, for a candidate word sequence such as "I am in Guiyang", the server can determine its fluency based on the following calculation formula (1):

[0124] P=P(T 我 )P(T 在 |T 我 )P(T 贵阳 |T 我 ,T 在 ) (1)

[0125] Among them, P(T 我 ) is the conditional probability of the word “I”, P(T 在 |T 我) is the conditional probability that the word "in" appears in the candidate word sequence containing the word "I", P(T 贵阳 |T 我 ,T 在 ) is the conditional probability that the word "Guiyang" appears in the candidate word sequence containing the two words "I" and "in", and P is the smoothness of the candidate word sequence "I am in Guiyang". In actual implementation, the server takes the product of the conditional probabilities of each candidate word as the smoothness of the candidate word sequence.

[0126] Next, the conversion sub-model can use the smoothness as the score of the candidate word sequence, and can also determine the corresponding score by combining the smoothness and the scoring parameter. Here, the scoring parameter can be a preset constant used to correct the smoothness. Then, the server selects the candidate word sequence with the highest score from multiple candidate word sequences as the conversion text corresponding to the phoneme sequence, that is, the conversion text corresponding to the corresponding audio sample.

[0127] In the embodiments of the present application, by combining the semantic features of the phoneme sequence to perform text conversion on the phoneme sequence, and scoring multiple candidate word sequences after conversion, and taking the candidate word sequence with the highest score as the conversion text corresponding to the phoneme sequence, it is possible to perform text conversion by combining the semantic features of the phoneme sequence, so that the converted text is more logically coherent, avoiding the deviation text that does not conform to the language logic obtained by conversion, and improving the accuracy and robustness of speech recognition.

[0128] Step 105: Obtain the errors between each of the conversion texts and the first text label respectively, and update the model parameters of the speech recognition model based on the obtained errors.

[0129] In actual implementation, the server determines the errors between each conversion text and the first text label for each conversion text, and updates the model parameters of the speech recognition model based on the determined errors. Here, the server can determine the errors between the conversion text and the first text label by calculating the word error rate. In the embodiments of the present application, the server updates the model parameters of the conversion sub-model based on the errors between the two. Only the conversion sub-model is trained. It should be understood that when only the conversion sub-model needs to be trained, the acoustic sub-model is a pre-trained acoustic model used to assist the training of the conversion sub-model. In some embodiments, the server can also update the model parameters of the conversion sub-model and the acoustic sub-model based on the errors between the two, and train these two models simultaneously.

[0130] In an embodiment of the present application, the original audio samples are generalized to obtain a corresponding plurality of audio samples carrying the first text label, and each of the audio samples is predicted by an acoustic sub-model to obtain a phoneme sequence corresponding to each audio sample, and then each phoneme sequence is converted into text by a conversion sub-model to obtain a converted text corresponding to each audio sample, and the model parameters of the speech recognition model are updated based on the error between each converted text and the first text label. By training the model with the generalized audio samples, the model can have a certain error correction capability, thereby making the model more robust and improving the accuracy of speech recognition.

[0131] In some embodiments, reference Figure 7 , Figure 7 is an optional flow chart of the training method of the speech recognition model provided in the embodiment of the present application, based on Figure 4 , you can also execute:

[0132] Step 401: The server obtains audio to be recognized.

[0133] In actual implementation, the audio to be recognized can be audio in any scenario, such as call voice on a customer service platform, or conversation voice recorded by users in social tools, etc.

[0134] Step 402: perform phoneme prediction on the audio to be recognized by using the acoustic sub-model to obtain a target phoneme sequence corresponding to the audio to be recognized.

[0135] Here, after obtaining the audio to be recognized, the server inputs the audio to be recognized into the acoustic sub-model, performs phoneme prediction on the audio to be recognized through the acoustic sub-model, and obtains the corresponding target phoneme sequence.

[0136] Step 403: Perform text conversion on the target phoneme sequence through the conversion sub-model to obtain a converted text corresponding to the audio to be recognized.

[0137] Next, the server inputs the target phoneme sequence output by the acoustic sub-model into the conversion sub-model, and converts the target phoneme sequence into text through the conversion sub-model to obtain the converted text corresponding to the audio to be recognized, thereby completing the speech recognition of the audio to be recognized. Here, since there may be interference noise or the lack of some semantic phonemes in the audio to be recognized and cannot form a complete sentence, the target phoneme sequence output by the acoustic sub-model is a phoneme sequence with deviations. The conversion sub-model converts it into a complete converted text based on the semantic features of the target phoneme sequence, so as to obtain accurate text through speech recognition.

[0138] In some embodiments, see Figure 8 , Figure 8It is an optional structural diagram of the speech recognition model provided in the embodiment of the present application. The speech recognition model provided in the embodiment of the present application includes an acoustic sub-model and a conversion sub-model. Among them, the conversion sub-model includes a pre-conversion sub-model and a retyping molecular model. Here, the retyping molecular model is a language model, which can be trained based on language samples in various fields, or it can be trained based on language samples in a specific field. For example, if the speech recognition model is used in the financial field, the retyping molecular model can be trained with samples from the financial field to more accurately locate the speech recognition model in the field, so that the recognized text conforms to the language habits of the financial field.

[0139] See also Fig. 9 , Fig. 9 is an optional flow chart of the training method of the speech recognition model provided in the embodiment of the present application, based on Figure 7 , step 403 can also be implemented in the following manner:

[0140] Step 501: The server performs text conversion on the target phoneme sequence through the pre-conversion sub-model to obtain multiple candidate texts corresponding to the audio to be recognized and a first score for each candidate text.

[0141] In actual implementation, the server inputs the audio to be recognized into the acoustic sub-model for phoneme prediction to obtain the corresponding target phoneme sequence, and then inputs the target phoneme sequence into the pre-conversion sub-model, and performs text conversion on the target phoneme sequence through the pre-conversion sub-model to obtain the corresponding multiple candidate texts and the first score of each candidate text. It should be noted that the candidate text here consists of multiple words, and the candidate text is also the candidate word sequence.

[0142] Step 502: Use the retyping molecular model to predict the score of each candidate text to obtain a corresponding second score.

[0143] In actual implementation, the server inputs the multiple candidate texts output by the pre-conversion sub-model into the re-typing molecular model, and scores each candidate text through the re-typing molecular model to obtain a second score corresponding to each candidate text. Here, the second score can be obtained by calculating the fluency of the candidate text.

[0144] Step 503: Determine a target score for each of the candidate texts based on the first score and the second score.

[0145] Next, the server determines a target score for the candidate text based on the first score and the second score corresponding to the candidate text. Here, the target score can be the product of the first score and the second score, or the weighted sum of the first score and the second score, and the weight of the first score and the weight of the second score can be obtained according to the attribute value of the pre-converted sub-model and the attribute value of the re-typed molecular model.

[0146] In some embodiments, step 503 can also be implemented in the following manner: the server obtains the attribute value of the pre-conversion sub-model and the attribute value of the re-typed molecular model; normalizes the attribute value of the pre-conversion sub-model and the attribute value of the re-typed molecular model respectively, determines the normalized result of the attribute value of the pre-conversion sub-model as the weight of the pre-conversion sub-model, and determines the normalized result of the attribute value of the re-typed molecular model as the weight of the re-typed molecular model; based on the weight of the pre-conversion sub-model and the weight of the re-typed molecular model, weights the first score and the second score of each candidate text, and determines the weighted score as the target score of each candidate text.

[0147] It should be noted that the weighted processing includes linear weighted processing, and different weights are determined according to the contribution of different models to the candidate text score. Exemplarily, the first score Scorea of ​​the pre-conversion sub-model for the candidate text and the second score Scorel of the re-typing sub-model for the candidate recognition text are obtained, and the target score Scoren of the candidate text is obtained by linear weighted processing, so the target score of the candidate text can be determined by formula (2):

[0148] Scoren=δ*Scorea+λ*Scorel (2)

[0149] Here, δ represents the weight of the first score, and λ represents the weight of the second score.

[0150] In some embodiments, obtaining the attribute value of the pre-conversion sub-model and the attribute value of the re-typed molecular model can be achieved by: obtaining the training indicators of the pre-conversion sub-model as the attribute value of the pre-conversion sub-model, and obtaining the training indicators of the re-typed molecular model as the attribute value of the re-typed molecular model; wherein the training indicators include at least one of the following: the number of training samples, the number of training times, and the timeliness of training.

[0151] In actual implementation, the number of training samples of the pre-conversion sub-model can be obtained, that is, the number of samples of the phoneme sequence used to train the pre-conversion sub-model, and the number of training samples of the re-typing molecular model can be obtained, that is, the number of samples of the language database. If the re-typing molecular model is a model in a specific field, the training samples are language samples in the specific field. Then, the server uses the number of samples of the two models as the attribute value of the model, and assigns different weights to the first score output by the pre-conversion sub-model and the second score corresponding to the re-typing molecular model based on the attribute value; the weight can be positively correlated with the number of samples. If the number of samples is larger, the assigned weight is higher, indicating that the corresponding model has a higher contribution to the score of the candidate text.

[0152] In actual implementation, the number of iterations of model training can also be obtained as the attribute value of the model, and the weight can be determined according to the number of model iterations. The weight can be positively correlated with the number of model iterations. If the number of model iterations is more, the assigned weight is higher, which indicates that the corresponding model has a higher contribution to the score of the candidate text.

[0153] In actual implementation, the model training timeliness can also be obtained as the attribute value of the model. The model training timeliness can include the inverse of the model's non-updated time (which can be understood as the time interval between the current time and the last update) or the average update cycle. The weight can be negatively correlated with the number of model training times. If the non-updated time or the average update cycle is longer, the assigned weight is lower, indicating that the corresponding model's contribution to the speech recognition candidate score is lower.

[0154] The training indicators of the model can reflect the training degree of the model. The different training degrees of the pre-conversion sub-model and the re-typing molecular model affect the model function and model effect. The contribution degree of the pre-conversion sub-model and the re-typing molecular model to the first score and the second score obtained by the pre-conversion sub-model and the re-typing molecular model for scoring the candidate text will be different. Through the linkage between the training indicators of the model and the weight setting in the weighted processing, the importance of the pre-conversion sub-model and the re-typing molecular model to the scoring of the candidate text is fully referred to, so that the target score of the candidate text obtained is more accurate and reasonable.

[0155] In other embodiments, obtaining the property value of the model can also be achieved in the following ways: obtaining the performance indicators of the pre-conversion sub-model as the property value of the pre-conversion sub-model; obtaining the performance indicators of the re-typed molecular model as the property value of the re-typed molecular model; wherein the performance indicators include at least one of the following: time complexity, space complexity.

[0156] It should be noted that time complexity determines the training / prediction time of the model. If the time complexity is too high, it will take a lot of time to train and predict the model, and it will be impossible to quickly improve the model or make fast predictions. Space complexity determines the number of parameters of the model. Due to dimensionality limitations, if the space complexity is higher, the more parameters the model has, the more data is needed to train the model, and the model training is more likely to overfit.

[0157] In actual implementation, the time complexity or space complexity of the pre-conversion sub-model and the re-typing molecular model is obtained as the model attribute value, and different weights are assigned to the pre-conversion sub-model score and the re-typing molecular model score based on the attribute value; illustratively, the weight may be negatively correlated with the time complexity or space complexity. If the time complexity of the model (computational amount / FLOPS, i.e., the number of operations of the model) is higher, the assigned weight is lower, and if the space complexity of the model (memory access amount / Bytes, i.e., the number of parameters of the model) is higher, the assigned weight is lower.

[0158] The performance index of the model is used to evaluate the quality of the model. Different performance indexes often give different results for evaluating the model. The performance indexes of the pre-conversion sub-model and the language model are different, and the contribution of the acoustic score and language score obtained for scoring the speech recognition candidate results will be different. Through the linkage between the performance index of the model and the weight setting in the weighted processing, the importance of the pre-conversion sub-model and the re-typing sub-model to the candidate text scoring is fully referenced, so that the target score of the candidate text obtained is more accurate and reasonable.

[0159] Step 504: Select a candidate text with the highest target score from the multiple candidate texts as a conversion text corresponding to the audio to be recognized.

[0160] In actual implementation, the server selects the candidate text with the highest target score from multiple candidate texts as the converted text after speech recognition of the audio to be recognized, and finally determines the target score of the candidate text by combining the first score of the candidate text after text conversion of the phoneme sequence by the pre-conversion sub-model and the second score of the candidate text by the retyping molecular model. This makes the scoring of the candidate text more accurate and more in line with the language field learned by the retyping molecular model, and can identify more targeted and accurate texts for specific fields.

[0161] Next, the training method of the speech recognition model provided in the embodiment of the present application is introduced. The training method of the speech recognition model provided in the embodiment of the present application is implemented by the terminal and the server in collaboration. Fig.10 , Fig.10 : is an optional flow chart of a method for training a speech recognition model provided in an embodiment of the present application. The method for training a speech recognition model provided in an embodiment of the present application includes:

[0162] Step 601: The terminal obtains original audio samples.

[0163] Here, the original audio sample can be collected by the terminal through a microphone connected to the terminal, or can be obtained from an audio library, or can be crawled from a web page. It should be noted that the original audio sample carries a first text label.

[0164] Step 602: The terminal sends a model training instruction carrying the original audio sample to the server.

[0165] In actual implementation, the terminal can generate corresponding model training instructions based on the original audio sample when obtaining the original audio sample. The terminal can also present a model training interface, present a model training function item in the model training interface, and generate corresponding model training instructions based on the original audio sample in response to a trigger operation for the model training function item. In addition, the terminal can also receive model training instructions sent by other devices, encapsulate the original audio sample into the model training instruction, obtain the model training instruction carrying the original audio sample and send it to the server.

[0166] Step 603: The server performs generalization processing on the original audio sample to obtain a plurality of audio samples corresponding to the original audio sample and carrying the first text label.

[0167] Step 604: The server performs phoneme prediction on each of the audio samples using the acoustic sub-model to obtain a phoneme sequence corresponding to each of the audio samples.

[0168] Step 605: The server performs text conversion on each of the phoneme sequences through the conversion sub-model to obtain a converted text corresponding to each of the audio samples.

[0169] Step 606: The server obtains the error between each of the converted texts and the first text label, and updates the model parameters of the speech recognition model based on the obtained error.

[0170] Step 607: The terminal obtains the audio to be recognized.

[0171] Here, the audio to be recognized can be collected by the microphone of the terminal through a voice application. For example, the voice application can be a social application, and the user starts the voice recording function by clicking a related function item in the voice application interface. After the voice recording function is started, the terminal collects audio through the microphone and uses the collected audio as the audio to be recognized.

[0172] Step 608: The terminal sends the audio to be recognized to the server.

[0173] In step 609, the server inputs the audio to be recognized into the acoustic sub-model, performs phoneme prediction on the audio to be recognized through the acoustic sub-model, and obtains a target phoneme sequence corresponding to the audio to be recognized.

[0174] Step 610: The server performs text conversion on the target phoneme sequence through the conversion sub-model to obtain a converted text corresponding to the audio to be recognized.

[0175] Step 611: The server sends the converted text to the terminal.

[0176] Step 612: The terminal outputs the converted text.

[0177] In actual implementation, after receiving the converted text sent by the server, the terminal outputs the converted text for the user to browse. For example, if the text to be recognized is collected by a social application, the terminal presents the converted text obtained after voice recognition of the audio to be recognized in the relevant area of ​​the social application interface.

[0178] In an embodiment of the present application, multiple audio samples carrying a first text label are obtained by generalizing the original audio samples, and then the speech recognition model is trained using the audio samples, so that the speech recognition model has a certain error correction capability, thereby improving the robustness of the speech recognition model.

[0179] The following is an explanation of an exemplary application of the embodiments of the present application in a practical application scenario.

[0180] The server obtains the original audio sample, which carries a text label. Exemplarily, the voice content of the original audio sample can be "I am in Guiyang", and the text label it carries is the text "I am in Guiyang". In actual implementation, the server adds a variety of noises to the original audio sample to obtain a plurality of audio samples, and then inputs these audio samples into the acoustic sub-model, performs phoneme prediction on the audio sample through the acoustic sub-model, and obtains the phoneme sequence corresponding to each audio sample. It should be understood that since the audio sample is obtained after adding noise based on the original audio sample, the acoustic sub-model may not be accurate when predicting the phoneme on it. In the embodiment of the present application, the server uses these phoneme sequences output by the acoustic sub-model as training samples for the conversion sub-model.

[0181] In some embodiments, the server also queries the standard phoneme sequence corresponding to the text label through the dictionary, and generalizes the standard phoneme sequence to obtain the corresponding deviation phoneme sequence. Specifically, the server changes at least one phoneme in the standard phoneme sequence, and the change includes at least one of the following: phoneme deletion, phoneme insertion and phoneme replacement. Exemplarily, for the label text "I am in Guiyang", the corresponding standard phoneme sequence is "w o3 z ai4 g ui4 y ang2," and the server can insert random phonemes into it, for example, simulating the background sound being misrecognized, inserting the background noise phoneme into it; or, simulating some sounds being missed, randomly deleting some of the phonemes; or, simulating some sounds being misrecognized, randomly replacing some of the phonemes. Through the above means, the standard phoneme sequence is data enhanced, the deviation phoneme sequence is constructed by simulating the possible errors of the acoustic sub-model, and the deviation phoneme sequence is used to train the conversion sub-model, so that the conversion sub-model has stronger robustness.

[0182] In actual implementation, the server first uses the standard phoneme sequence to perform the first round of training on the conversion sub-model until the model reaches the convergence condition, and obtains the trained conversion sub-model M1. Then, the server uses the deviation phoneme sequence obtained in the above manner to perform the second round of training on the trained conversion sub-model M1, and obtains the trained conversion sub-model M2. Specifically, the server inputs the standard phoneme sequence into the conversion sub-model, performs text conversion on the standard phoneme sequence through the conversion sub-model, obtains the converted text corresponding to the standard phoneme sequence, and updates the model parameters of the conversion sub-model based on the error between the converted text and the text label. The model parameters of the model are continuously updated through continuous iterative training until the convergence condition is reached, and the iterative training of the model is stopped to obtain the trained conversion sub-model M1. Then, the server inputs the deviation phoneme sequence into the trained conversion sub-model M1, and performs text conversion on the deviation phoneme sequence through the conversion sub-model M1 to obtain the conversion text corresponding to the deviation phoneme sequence, and then updates the trained conversion sub-model M1 based on the error between the conversion text and the text label. When the convergence condition is reached, the training of the model M1 is stopped to obtain the conversion sub-model M2 obtained by two rounds of training. It should be noted that here, the conversion sub-model can extract the contextual semantic features of the phoneme sequence and perform text conversion based on the contextual semantic features, so as to obtain a conversion text that is more in line with the semantics of the phoneme sequence.

[0183] After obtaining the trained acoustic sub-model and conversion sub-model, the server can use the speech recognition model composed of the two to perform corresponding speech recognition. Specifically, the server obtains the audio to be recognized, performs phoneme prediction on the audio to be recognized through the acoustic sub-model, obtains the target phoneme sequence corresponding to the audio to be recognized, and performs text conversion on the target phoneme sequence through the conversion sub-model to obtain the converted text corresponding to the audio to be recognized. For example, see Fig.11 , Fig.11 This is an optional flow chart of the speech recognition process provided in an embodiment of the present application. If the audio to be recognized is "I am in Guiyang" with an accent or interference information, the server inputs the audio to be recognized into the acoustic sub-model, performs phoneme prediction on the audio to be recognized through the acoustic sub-model, and outputs a phoneme sequence with a deviation "w o3 z ai4 g ui4 y v2," and then inputs the phoneme sequence with a deviation into the conversion sub-model, performs text conversion on the phoneme sequence through the conversion sub-model, corrects the deviated phoneme sequence, and outputs the correct text "I am in Guiyang".

[0184] In an embodiment of the present application, the server first performs a round of training on the conversion sub-model through a standard phoneme sequence, so that the model learns the mapping relationship between standard text and phonemes, and obtains a trained conversion sub-model. Then, the server performs a second round of training on the trained conversion sub-model based on the deviation phoneme sequence, so that the model can combine the contextual information of the phonemes, correct the deviation phoneme sequence, and obtain a standard conversion text, thereby improving the robustness of the model and enabling more accurate speech recognition of audio containing noise in reality.

[0185] The following is a description of an exemplary structure of a speech recognition model training device 555 provided in an embodiment of the present application implemented as a software module. In some embodiments, for example, Fig.12 As shown, Fig.12 is an optional structural diagram of a training model of a speech recognition model provided in an embodiment of the present application. The software modules stored in the training device 555 of the speech recognition model in the memory 550 may include:

[0186] An acquisition module 5551 is used to acquire an original audio sample, where the original audio sample carries a first text label;

[0187] A generalization module 5552 is used to perform generalization processing on the original audio sample to obtain a plurality of audio samples corresponding to the original audio sample and carrying the first text label;

[0188] A phoneme prediction module 5553, configured to perform phoneme prediction on each of the audio samples using the acoustic sub-model to obtain a phoneme sequence corresponding to each of the audio samples;

[0189] A text conversion module 5554 is used to perform text conversion on each of the phoneme sequences through the conversion sub-model to obtain a conversion text corresponding to each of the audio samples;

[0190] The updating module 5555 is used to respectively obtain the errors between each of the converted texts and the first text label, and update the model parameters of the speech recognition model based on the obtained errors.

[0191] In some embodiments, the training device of the speech recognition model also includes: a pre-training module, which is used to perform phoneme conversion on the second text label based on the mapping relationship between text and phonemes to obtain a standard phoneme sequence corresponding to the second text label; perform text conversion on the standard phoneme sequence through the conversion sub-model to obtain a corresponding target converted text; based on the error between the target converted text and the second text label, update the model parameters of the conversion sub-model to obtain an updated conversion sub-model; accordingly, the text conversion module is also used to perform text conversion on each of the phoneme sequences respectively through the updated conversion sub-model.

[0192] In some embodiments, the pre-training module is also used to generalize the standard phoneme sequence to obtain a corresponding plurality of deviation phoneme sequences; correspondingly, the text conversion module is also used to perform text conversion on each of the phoneme sequences and each of the deviation phoneme sequences respectively through the updated conversion sub-model.

[0193] In some embodiments, the generalization module is further used to obtain multiple interference information; and based on each interference information, perform the following processing: adding the interference information to the original audio sample.

[0194] In some embodiments, the generalization module is further used to perform the following processing on the original audio sample multiple times: modify at least one frame of speech signal of the original audio sample, each frame of the speech signal corresponds to a phoneme; wherein the modification includes at least one of the following: phoneme deletion, phoneme insertion and phoneme replacement.

[0195] In some embodiments, the text conversion module is further used to perform the following processing on each of the phoneme sequences through the conversion sub-model: extracting semantic features of the phoneme sequence to obtain corresponding semantic features; based on the semantic features, performing text conversion on the phoneme sequence to obtain multiple candidate word sequences corresponding to the phoneme sequence, and scores corresponding to each of the candidate word sequences; and selecting the candidate word sequence with the highest score from the multiple candidate word sequences as the converted text corresponding to the audio sample.

[0196] In some embodiments, the text conversion module is also used to perform text conversion on the phoneme sequence based on the semantic features to obtain multiple candidate word sequences corresponding to the phoneme sequence and the conditional probability of each candidate word in each of the candidate word sequences; based on the conditional probability of each candidate word in the candidate word sequence, determine the score corresponding to each of the candidate word sequences.

[0197] In some embodiments, the training device of the speech recognition model also includes: a speech recognition module for acquiring audio to be recognized; performing phoneme prediction on the audio to be recognized through the acoustic sub-model to obtain a target phoneme sequence corresponding to the audio to be recognized; and performing text conversion on the target phoneme sequence through the conversion sub-model to obtain a converted text corresponding to the audio to be recognized.

[0198] In some embodiments, the conversion sub-model includes a pre-conversion sub-model and a retyping molecular model, and the speech recognition module is further used to perform text conversion on the target phoneme sequence through the pre-conversion sub-model to obtain multiple candidate texts corresponding to the audio to be recognized and a first score for each of the candidate texts; to predict the score of each of the candidate texts through the retyping molecular model to obtain a corresponding second score; to determine the target score of each of the candidate texts based on the first score and the second score; and to select the candidate text with the highest target score from the multiple candidate texts as the conversion text corresponding to the audio to be recognized.

[0199] It should be noted that the description of the device of the embodiment of the present application is similar to the description of the above-mentioned method embodiment, and has similar beneficial effects as the method embodiment, so it will not be repeated.

[0200] An embodiment of the present application provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the training method of the speech recognition model provided in the embodiment of the present application is implemented.

[0201] The present application embodiment provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will be caused to execute the method provided by the present application embodiment, for example, Figure 4 The training method of the speech recognition model is shown.

[0202] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0203] In some embodiments, executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.

[0204] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).

[0205] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices located at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.

[0206] In summary, the embodiments of the present application can train a highly robust speech recognition model, thereby improving the accuracy of speech recognition.

[0207] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A method for training a speech recognition model, characterized in that: The speech recognition model includes an acoustic sub-model and a conversion sub-model, and the method includes: Acquire an original audio sample, where the original audio sample carries a first text label; Performing generalization processing on the original audio sample to obtain a plurality of audio samples corresponding to the original audio sample and carrying the first text label; Performing phoneme prediction on each of the audio samples using the acoustic sub-model to obtain a phoneme sequence corresponding to each of the audio samples; Performing text conversion on each of the phoneme sequences respectively through the conversion sub-model to obtain a conversion text corresponding to each of the audio samples; Respectively obtaining an error between each of the converted texts and the first text label, and updating a model parameter of the speech recognition model based on the obtained error; The step of performing text conversion on each of the phoneme sequences by using the conversion sub-model to obtain a conversion text corresponding to each of the audio samples includes: The following processing is performed on each of the phoneme sequences through the conversion sub-model: Extracting semantic features from the phoneme sequence to obtain corresponding semantic features; Based on the semantic features, the phoneme sequence is converted into text to obtain a plurality of candidate word sequences corresponding to the phoneme sequence and a conditional probability of each candidate word in each of the candidate word sequences; Based on the conditional probability of each candidate word in the candidate word sequence, respectively determine the score corresponding to each candidate word sequence; From the multiple candidate word sequences, select the candidate word sequence with the highest score as the converted text corresponding to the audio sample.

2. The method according to claim 1, characterized in that Before performing text conversion on each of the phoneme sequences by using the conversion submodel to obtain the converted text corresponding to each of the audio samples, the method further includes: Based on the mapping relationship between text and phonemes, performing phoneme conversion on the second text label to obtain a standard phoneme sequence corresponding to the second text label; Performing text conversion on the standard phoneme sequence through the conversion sub-model to obtain a corresponding target conversion text; Based on the error between the target conversion text and the second text label, updating the model parameters of the conversion sub-model to obtain an updated conversion sub-model; Correspondingly, performing text conversion on each of the phoneme sequences by using the conversion sub-model includes: The updated conversion sub-model is used to perform text conversion on each of the phoneme sequences.

3. The method according to claim 2, characterized in that The method further comprises: Performing generalization processing on the standard phoneme sequence to obtain a corresponding plurality of deviation phoneme sequences; Correspondingly, the text conversion is performed on each of the phoneme sequences using the updated conversion sub-model, including: The updated conversion sub-model is used to perform text conversion on each of the phoneme sequences and each of the deviation phoneme sequences.

4. The method according to claim 1, characterized in that: The generalizing process of the original audio sample comprises: Obtain multiple interference information; Based on each interference information, the following processing is performed: The interference information is added to the original audio samples.

5. The method according to claim 1, characterized in that The generalizing process of the original audio sample comprises: The following processing is performed multiple times on the raw audio sample: Modifying at least one frame of speech signal on the original audio sample, each frame of the speech signal corresponding to a phoneme; The modification includes at least one of the following: phoneme deletion, phoneme insertion and phoneme replacement.

6. The method according to claim 1, characterized in that The method further comprises: Get the audio to be recognized; Performing phoneme prediction on the audio to be recognized by using the acoustic sub-model to obtain a target phoneme sequence corresponding to the audio to be recognized; The target phoneme sequence is converted into text by using the conversion sub-model to obtain a converted text corresponding to the audio to be recognized.

7. The method according to claim 6, characterized in that The conversion sub-model includes a pre-conversion sub-model and a re-typing sub-model. The text conversion of the target phoneme sequence is performed by the conversion sub-model to obtain a conversion text corresponding to the audio to be recognized, including: Performing text conversion on the target phoneme sequence by using the pre-conversion sub-model to obtain a plurality of candidate texts corresponding to the audio to be recognized and a first score for each of the candidate texts; Using the retyping molecular model, respectively predicting the scores of the candidate texts, and obtaining corresponding second scores; Determining a target score for each of the candidate texts based on the first score and the second score; A candidate text with the highest target score is selected from the multiple candidate texts as the converted text corresponding to the audio to be recognized.

8. A training device for a speech recognition model, characterized in that: The speech recognition model includes an acoustic sub-model and a conversion sub-model, and the device includes: An acquisition module, configured to acquire an original audio sample, wherein the original audio sample carries a first text label; A generalization module, configured to perform generalization processing on the original audio sample to obtain a plurality of audio samples corresponding to the original audio sample and carrying the first text label; A phoneme prediction module, used to perform phoneme prediction on each of the audio samples using the acoustic sub-model to obtain a phoneme sequence corresponding to each of the audio samples; A text conversion module is used to perform text conversion on each of the phoneme sequences through the conversion submodel to obtain a conversion text corresponding to each of the audio samples; the text conversion module is also used to perform the following processing on each of the phoneme sequences through the conversion submodel: extract semantic features from the phoneme sequence to obtain corresponding semantic features; based on the semantic features, perform text conversion on the phoneme sequence to obtain multiple candidate word sequences corresponding to the phoneme sequence and the conditional probability of each candidate word in each of the candidate word sequences; based on the conditional probability of each candidate word in the candidate word sequence, determine the score corresponding to each of the candidate word sequences; from the multiple candidate word sequences, select the candidate word sequence with the highest score as the conversion text corresponding to the audio sample; An updating module is used to respectively obtain the error between each of the converted texts and the first text label, and update the model parameters of the speech recognition model based on the obtained error.

9. An electronic device, characterized in that: include: A memory for storing executable instructions; A processor, used to implement the training method of the speech recognition model described in any one of claims 1 to 7 when executing the executable instructions stored in the memory.

10. A computer-readable storage medium, characterized in that: Executable instructions are stored, which are used to implement the training method of the speech recognition model described in any one of claims 1 to 7 when executed by a processor.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the training method of the speech recognition model described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Audio recognition method, system and machinery device

    CN109859743A

  • Method and apparatus for speech recognition

    US20160155436A1