Speech recognition method and device, electronic equipment and storage medium
By introducing the hybrid expert model MoE module into the speech recognition model, and using routing networks and expert network groups to process mixed speech, the problem of low accuracy of multilingual mixed speech recognition is solved, and higher recognition accuracy and more accurate speech representation are achieved.
Patent Information
- Application Number
- CN202510012055.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-03
AI Technical Summary
Existing speech recognition models have low recognition accuracy when facing mixed speeches including multiple language types, especially in mixed speech scenarios in Chinese and English.
A hybrid speech recognition model is adopted, which includes a target encoding module, a hybrid expert model MoE module and a target decoding module. The MoE module processes the initial speech representation through a routing network and multiple expert network groups. The routing network determines the target expert network group based on the initial speech representation and processes the speech representation through this group.
It improves the recognition accuracy of mixed speeches including multiple language types, reduces confusion and misrecognition of language types, and enhances the accuracy of target speech representation.
Smart Images

Figure CN119943050A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a speech recognition method, device, electronic device and storage medium. Background Art
[0002] Automatic speech recognition is a technology that automatically transcribes speech into corresponding text. It can provide support for speech content understanding and human-computer interaction, and is widely used in various basic businesses, such as search, voice assistants, automatic subtitles, etc. The current speech recognition model has a usable accuracy rate for speech containing one language, but in many scenarios, speech data may include multiple languages, such as English professional terms that may appear when communicating in Chinese. When faced with speech containing multiple languages, such as speech containing Chinese and English, the accuracy of speech recognition is low. Summary of the invention
[0003] The embodiments of the present application disclose a speech recognition method, device, electronic device and storage medium, which improve the recognition accuracy of mixed speech containing multiple languages.
[0004] The present application discloses a speech recognition method, including:
[0005] Get the speech to be recognized;
[0006] The speech to be recognized is recognized by a hybrid speech recognition model to obtain a speech recognition result; the hybrid speech recognition model includes a target encoding module, a hybrid expert model MoE module and a target decoding module, the target encoding module is used to encode the speech to be recognized to obtain an initial speech representation corresponding to the speech to be recognized; the MoE module includes a routing network and a plurality of expert network groups corresponding to a plurality of language types, the routing network is used to determine a target expert network group corresponding to the initial speech representation from the plurality of expert network groups according to the initial speech representation, and the initial speech representation is processed by the target expert network group to obtain a target speech representation; the target decoding module is used to perform decoding processing according to the target speech representation to obtain the speech recognition result.
[0007] In one embodiment, each of the expert network groups includes a routing unit and multiple expert networks; the routing unit is used to determine one or more target expert networks from the multiple expert networks based on the initial speech representation, and process the initial speech representation through the one or more target expert networks to obtain a target speech representation.
[0008] In one embodiment, after acquiring the speech to be recognized, the method further includes:
[0009] Performing frame processing on the speech to be recognized to obtain multiple frames of sub-speech data;
[0010] The step of recognizing the speech to be recognized by using the hybrid speech recognition model to obtain a speech recognition result includes:
[0011] Encoding the multiple frames of sub-speech data by the target encoding module to obtain initial speech representations corresponding to the multiple frames of sub-speech respectively;
[0012] Determining, by the routing network, a target expert network group corresponding to the target frame sub-speech data from the plurality of expert network groups according to the initial speech representation corresponding to the target frame sub-speech data; the target frame sub-speech data is any frame sub-speech data;
[0013] Determining one or more target expert networks from the target expert network group by a routing unit in the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data;
[0014] Processing the initial speech representation corresponding to the target frame sub-speech data through the one or more target expert networks to obtain a fused speech representation corresponding to the target frame sub-speech data;
[0015] Obtaining a target speech representation according to the fused speech representations corresponding to the multiple frames of sub-speech data;
[0016] The target speech representation is decoded by the target decoding module to obtain a speech recognition result.
[0017] In one embodiment, determining, by the routing network according to the initial speech representation corresponding to the target frame sub-speech data, from the plurality of expert network groups a target expert network group corresponding to the target frame sub-speech data comprises:
[0018] The routing network determines the target language type corresponding to the target frame sub-speech data according to the initial speech representation corresponding to the target frame sub-speech data, and determines the expert network group corresponding to the target language type as the target expert network group corresponding to the target frame sub-speech data.
[0019] In one embodiment, determining one or more target expert networks from the target expert network group by a routing unit in the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data includes:
[0020] Determining, by a routing unit in the target expert network group, one or more target expert networks and network weights corresponding to each of the target expert networks from the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data;
[0021] The processing of the initial speech representation corresponding to the target frame sub-speech data by the one or more target expert networks to obtain the fused speech representation corresponding to the target frame sub-speech data includes:
[0022] Processing the initial speech representations corresponding to the target frame sub-speech data respectively through the one or more target expert networks to obtain processing results corresponding to each of the target expert networks;
[0023] Based on the network weights corresponding to the target expert networks, the processing results corresponding to the target expert networks are weightedly fused to obtain the fused speech representation corresponding to the target frame sub-speech data.
[0024] In one embodiment, the target encoding module is obtained by self-supervised training based on the first data set and supervised training based on the second data set, and the MoE module and the target decoding module are obtained by supervised training based on the second data set;
[0025] The first data set includes a plurality of unlabeled first sample speech, each of which corresponds to a language type; the second data set includes a plurality of second sample speech and labeling information corresponding to each of the second sample speech, each of which corresponds to one or more language types.
[0026] In one embodiment, the method further comprises:
[0027] Based on the first data set, self-supervised training is performed on the initial encoding module to obtain a first encoding module, wherein the first encoding module includes multiple encoding layers;
[0028] The second encoding module is supervisedly trained based on the second data set to obtain a target encoding module; the second encoding module is obtained by adding a conditional decoding layer between every two adjacent encoding layers in the first encoding module; the conditional decoding layer is used to calculate the posterior probability of the connected previous encoding layer.
[0029] In one embodiment, the method further comprises:
[0030] Normalizing the first speech representation extracted by the current coding layer to obtain a second speech representation;
[0031] Calculating, by means of a conditional decoding layer connected to the current encoding layer, a posterior probability corresponding to the second speech representation;
[0032] Performing a linear projection on the posterior probability through a linear projection layer so that the dimension of the posterior probability is aligned with the dimension of the second speech representation;
[0033] The mapped posterior probability is residually connected with the second speech representation to obtain a third speech representation, and the third speech representation is input to a coding layer next to the current coding layer.
[0034] The present application discloses a speech recognition device, comprising:
[0035] A speech acquisition module, used to acquire the speech to be recognized;
[0036] A speech recognition module is used to recognize the speech to be recognized through a hybrid speech recognition model to obtain a speech recognition result; the hybrid speech recognition model includes a target encoding module, a hybrid expert model MoE module and a target decoding module, the target encoding module is used to encode the speech to be recognized to obtain an initial speech representation corresponding to the speech to be recognized; the MoE module includes a routing network and a plurality of expert network groups corresponding to a plurality of language types, the routing network is used to determine a target expert network group corresponding to the initial speech representation from the plurality of expert network groups according to the initial speech representation, and process the initial speech representation through the target expert network group to obtain a target speech representation; the target decoding module is used to perform decoding processing according to the target speech representation to obtain the speech recognition result.
[0037] In one embodiment, each of the expert network groups includes a routing unit and multiple expert networks; the routing unit is used to determine one or more target expert networks from the multiple expert networks based on the initial speech representation, and process the initial speech representation through the one or more target expert networks to obtain a target speech representation.
[0038] In one embodiment, the speech acquisition module is further used to perform frame processing on the speech to be recognized to obtain multiple frames of sub-speech data; the speech recognition module is further used to encode the multiple frames of sub-speech data through the target encoding module to obtain initial speech representations corresponding to the multiple frames of sub-speech; determine the target expert network group corresponding to the target frame sub-speech data from the multiple expert network groups according to the initial speech representation corresponding to the target frame sub-speech data through the routing network; the target frame sub-speech data is any frame sub-speech data; determine one or more target expert networks from the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data through the routing unit in the target expert network group; process the initial speech representation corresponding to the target frame sub-speech data through the one or more target expert networks to obtain a fused speech representation corresponding to the target frame sub-speech data; obtain a target speech representation according to the fused speech representations corresponding to the multiple frames of sub-speech data; decode the target speech representation through the target decoding module to obtain a speech recognition result.
[0039] In one embodiment, the speech recognition module is also used to determine the target language type corresponding to the target frame sub-speech data based on the initial speech representation corresponding to the target frame sub-speech data through the routing network, and determine the expert network group corresponding to the target language type as the target expert network group corresponding to the target frame sub-speech data.
[0040] In one embodiment, the speech recognition module is also used to determine one or more target expert networks and the network weights corresponding to each target expert network from the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data through the routing unit in the target expert network group; process the initial speech representation corresponding to the target frame sub-speech data respectively through the one or more target expert networks to obtain the processing results corresponding to each target expert network; and perform weighted fusion on the processing results corresponding to each target expert network based on the network weights corresponding to each target expert network to obtain the fused speech representation corresponding to the target frame sub-speech data.
[0041] In one embodiment, the target encoding module is obtained by self-supervised training based on a first data set and supervised training based on a second data set, and the MoE module and the target decoding module are obtained by supervised training based on the second data set; the first data set includes a plurality of unlabeled first sample speech, each of the first sample speech corresponds to a language type, and the second data set includes a plurality of second sample speech and labeling information corresponding to each of the second sample speech, each of the second sample speech corresponds to one or more language types.
[0042] In one embodiment, the speech recognition device may also include a model training module, which is used to perform self-supervised training on the initial coding module based on the first data set to obtain a first coding module, wherein the first coding module includes multiple coding layers; and perform supervised training on the second coding module based on the second data set to obtain a target coding module; the second coding module is obtained by adding a conditional decoding layer between every two adjacent coding layers in the first coding module; the conditional decoding layer is used to calculate the posterior probability of the connected previous coding layer.
[0043] In one embodiment, the model training module is also used to normalize the first speech representation extracted by the current coding layer to obtain a second speech representation; calculate the posterior probability corresponding to the second speech representation through the conditional decoding layer connected to the current coding layer; linearly project the posterior probability through the linear projection layer so that the dimension of the posterior probability is aligned with the dimension of the second speech representation; perform a residual connection between the mapped posterior probability and the second speech representation to obtain a third speech representation, and input the third speech representation into the next coding layer of the current coding layer.
[0044] The present application discloses an electronic device, including:
[0045] A memory storing executable program code;
[0046] a processor coupled to the memory;
[0047] The processor calls the executable program code stored in the memory to execute the method described in any one of the above embodiments.
[0048] An embodiment of the present application discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor executes the method described in any one of the above embodiments.
[0049] Through the speech recognition method, device, electronic device and storage medium disclosed in the embodiments of the present application, the electronic device can obtain the speech to be recognized, and recognize the speech through a hybrid speech recognition model to obtain a speech recognition result, wherein the hybrid speech recognition model includes a target encoding module, an MoE module and a target decoding module, the target encoding module can be used to encode the speech to be recognized to obtain an initial speech representation corresponding to the speech to be recognized, the MoE module includes a routing network and multiple expert network groups corresponding to multiple languages; the routing network can be used to determine the target expert network group corresponding to the initial speech representation from multiple expert network groups according to the initial speech representation, and process the initial speech representation through the target expert network group to obtain a target speech representation; the target decoding module is used to perform decoding processing according to the target speech representation to obtain a speech recognition result. Through this embodiment, multiple expert network groups can be obtained in MoE by grouping according to language types. The routing module in MoE can route the initial speech representation corresponding to the speech to be recognized to the corresponding target expert network group. The language type contained in the speech to be recognized can correspond to the language type corresponding to the target expert network. The initial speech representation is then processed by the target expert network group, which can accurately capture the speech features of different language types to obtain the target speech representation, reduce confusion and misrecognition of language types, and improve the accuracy of the target speech representation, thereby improving the recognition accuracy of mixed speech containing multiple language types. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0051] Figure 1-A It is a structural diagram of a speech recognition model disclosed in the related art;
[0052] Figure 1-B is a structural diagram of another speech recognition model disclosed in the related art;
[0053] Figure 2 It is a schematic diagram of an application scenario of a differential upgrade method disclosed in an embodiment of the present application;
[0054] Figure 3 It is a flowchart of a speech recognition method disclosed in an embodiment of the present application;
[0055] Figure 4 It is a structural schematic diagram of a MoE module disclosed in an embodiment of the present application;
[0056] Figure 5 It is a flowchart of a speech recognition method disclosed in an embodiment of the present application;
[0057] Figure 6 It is a flowchart of a training process of a target coding module disclosed in an embodiment of the present application;
[0058] Figure 7 is a schematic diagram of a target encoding module disclosed in an embodiment of the present application;
[0059] Figure 8 It is a modular schematic diagram of a speech recognition device disclosed in an embodiment of the present application;
[0060] Fig. 9 It is a structural block diagram of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0061] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0062] It should be noted that the terms "including" and "having" in the embodiments of the present application and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0063] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of this application, a first data set may be referred to as a second data set, and similarly, a second data set may be referred to as a first data set. Both the first data set and the second data set are data sets, but they are not the same data set.
[0064] For mixed speech containing multiple languages, since speech of different languages usually has similar pronunciations, such as "I" in English and "love" in Chinese, it is difficult for the speech recognition model to recognize such words in mixed speech. In the early days, a cascade system was usually used to recognize mixed speech. The cascade system may include a speech boundary detection system, a language recognition system, and speech recognition systems corresponding to multiple languages. First, the speech is segmented by the language boundary detection system to obtain multiple speech segments, and then the multiple speech segments are classified into language types by the language recognition system. Finally, the corresponding speech recognition system is used to perform speech recognition on each speech segment. However, the cascade system is prone to cumulative propagation of errors, resulting in low accuracy of speech recognition, and the cascade system has a high system delay, which greatly affects the experience of human-computer interaction.
[0065] Therefore, end-to-end neural network models have gradually become the mainstream solution for identifying mixed speech. The speech recognition model in the relevant technology may include encoders corresponding to multiple language types. Each encoder can learn the ability to distinguish speech of different languages during the training process, so as to extract the speech representations corresponding to the multiple language types respectively, and then perform weighted fusion on the speech representations output by multiple encoders to obtain the final speech representation, thereby improving the accuracy of speech recognition.
[0066] like Figure 1-A As shown, Figure 1-A It is a structural diagram of a speech recognition model disclosed in the related art, in which the Chinese encoder can extract and process the semantic information in the speech to be recognized according to Chinese rules and generate the corresponding Chinese speech representation The English encoder can extract and process the semantic information in the speech to be recognized according to English rules and generate the corresponding English speech representation. The gated network can represent the Chinese speech generated by the Chinese encoder and English speech representations generated by the English encoder Perform weighted fusion to obtain a fused speech representation. Among them, the gating network can determine the weight corresponding to the Chinese encoder based on the Softmax function: And determine the weight corresponding to the English encoder as So according to the weight corresponding to the Chinese encoder and the corresponding weights for the English encoder Representation of Chinese Phonology and English phonetic representation Perform weighted fusion to obtain fused speech representation However, the training process of the speech recognition model is complex and difficult. Usually, the encoder corresponding to each language type needs to be pre-trained, and the speech recognition model has a large computational complexity, and the speech recognition efficiency is low.
[0067] like Figure 1-B As shown, Figure 1-B This is a structural diagram of another speech recognition model disclosed in the related art, in which the routing network can calculate the mixed speech through the Softmax function to obtain the probability distribution corresponding to each expert network, and select some expert networks from n expert networks according to the probability distribution corresponding to each expert network, so as to activate the expert network to perform speech recognition on the mixed speech, and fuse the recognition results of the expert network to obtain the final recognition result, thereby effectively reducing the computational complexity while keeping the task complexity unchanged. However, Figure 1-B The speech recognition model in the system requires a large amount of manually annotated data for training in order to improve the recognition accuracy of the speech recognition model. Moreover, the accuracy for mixed speech containing multiple languages is still not high.
[0068] The embodiments of the present application disclose a speech recognition method, apparatus, electronic device and storage medium. A plurality of expert network groups are obtained by grouping according to language types, thereby determining a target expert network group corresponding to the language type corresponding to the speech to be recognized. The initial speech representation corresponding to the speech to be recognized is processed by the target expert network group, thereby obtaining a more accurate target speech representation, which can improve the recognition accuracy of mixed speech containing multiple language types.
[0069] The following is a detailed description with reference to the accompanying drawings.
[0070] like Figure 2 As shown, Figure 2 It is a schematic diagram of an application scenario of a differential upgrade method disclosed in an embodiment of the present application. The application scenario may include an electronic device 210, which may include but is not limited to a mobile phone, a tablet computer, a wearable device, a laptop computer, a PC (Personal Computer), etc.
[0071] In one embodiment, the electronic device 210 may include a hybrid speech recognition model, and the electronic device 210 may obtain a speech to be recognized, and recognize the speech to be recognized through the hybrid speech recognition model to obtain a speech recognition result. Figure 2 In the example, the speech to be recognized can be input into the electronic device 210, and the electronic device 210 recognizes the speech to be recognized through the hybrid speech recognition model to obtain a speech recognition result "good morning, good morning everyone".
[0072] In view of the feature that the electronic device 210 can perform voice recognition, a voice recognition function can be set in some application software to facilitate users to complete specific operations through voice. For example, in the scenario where the user uses a mobile phone voice assistant, the user can send control instructions to the mobile phone through voice, and the mobile phone can detect the user's voice and convert the user's voice into target text through a voice recognition model to determine the user's control instruction and respond; in the scenario of meeting recording, the meeting voice of one or more users can be detected, converted into meeting text and recorded.
[0073] like Figure 3 As shown, Figure 3 : is a flow chart of a speech recognition method disclosed in an embodiment of the present application. The speech recognition method can be applied to the electronic device in the above embodiment. The speech recognition method may include the following steps:
[0074] Step 310: Acquire speech to be recognized.
[0075] The speech to be recognized may refer to the original speech signal or speech data to be recognized. In different scenarios, the speech to be recognized may be obtained in different ways. Optionally, in the scenario of real-time speech recognition, the speech to be recognized may be collected in real time by the electronic device through a speech collection device such as a microphone. In the scenario of offline speech processing, the speech to be recognized may be pre-stored in the electronic device or imported into the electronic device.
[0076] When an electronic device detects a speech to be recognized in real time, the electronic device may include a speech detection device and a speech processing device. The speech detection device may be used to collect sound signals outside the electronic device, such as a microphone, etc. The speech processing device may be used to perform analog-to-digital conversion on the collected sound signal to obtain the speech to be recognized, such as an analog-to-digital converter. It can be understood that the collected sound signal may be an analog signal, and the speech to be recognized may be a digital signal.
[0077] Step 320: Recognize the speech to be recognized by using the hybrid speech recognition model to obtain a speech recognition result.
[0078] The hybrid speech recognition model may include a target encoding module, a hybrid expert model MoE module and a target decoding module. The target encoding module is used to encode the speech to be recognized to obtain an initial speech representation corresponding to the speech to be recognized; the MoE module includes a routing network and multiple expert network groups corresponding to multiple language types. The routing network is used to determine the target expert network group corresponding to the initial speech representation according to the initial speech representation, and process the initial speech representation through the target expert network group to obtain a target speech representation; the target decoding module is used to perform decoding processing according to the target speech representation to obtain a speech recognition result.
[0079] Optionally, the target encoding module may refer to an encoding module specifically used for speech recognition, which may include a feature extraction network and an encoder, such as a HuBERT model. The feature extraction network may include a convolutional downsampling network, such as a pre-trained CNN (Convolutional Neural Networks), and the encoder may include multiple encoding layers, such as multiple transformer layers.
[0080] Among them, the encoding process of the target encoding module for the speech to be recognized may include signal preprocessing, feature extraction, feature mapping, etc. Signal preprocessing may refer to performing noise reduction processing on the speech to be recognized to obtain a preprocessed signal, such as filtering out high-frequency signals or low-frequency signals to eliminate noise interference. Feature extraction may refer to extracting signal features of the preprocessed signal through a feature extraction network, and the signal features may include spectral features or time domain features. Feature mapping may refer to performing multi-level feature extraction and transformation on signal features through multiple coding layers to obtain high-dimensional features and then reducing the dimensionality of the high-dimensional features and mapping them to a hidden state sequence of a low-dimensional vector space as a speech representation. It is understandable that in order to improve the robustness of speech recognition, noise can be added to the preprocessed signal.
[0081] It can be understood that speech representation is an abstract expression extracted by deep neural networks. Compared with the signal characteristics that directly describe the physical or statistical properties of the speech to be recognized, speech representation can better represent the semantic information of the speech to be recognized, is more suitable for use in the field of speech recognition, and can improve the accuracy of speech recognition.
[0082] Optionally, the target decoding module may refer to a decoding module specifically used for speech recognition, which is used to convert speech representation into recognition results, such as a CTC (Connectionist Temporal Classification) decoder, which can convert speech representation into a text sequence.
[0083] The multiple expert network groups in the MoE (Mixture of Experts) module correspond to multiple languages one by one. Each expert network group can be used to process the speech representation of the language type. For example, the expert network group corresponding to Chinese can process the speech representation of Chinese. Each expert network group can include one or more expert networks. The expert networks in the same expert network group all process the speech representation of the same language type, but the network functions corresponding to each expert network in the same expert network group can be different. For example, in the expert network group corresponding to Chinese, the network function corresponding to one expert network can be semantic analysis of Chinese speech representation, and the network function corresponding to another expert network can be sentiment analysis of Chinese speech representation.
[0084] The routing network is responsible for determining a suitable target expert network group from a plurality of expert network groups corresponding to a plurality of languages to process the initial speech representation so that the processed speech representation can more prominently represent specific speech information, such as speech text, speech emotion, etc. Among them, the routing network in the MoE module can determine the target expert network group corresponding to the initial speech representation based on the initial speech representation, and activate the expert network in the target expert network group so that the activated expert network processes the initial speech representation. Optionally, the routing network can include a preset specific function, such as a Softmax function, so as to calculate the score or weight of the initial speech representation in each target expert network group according to the specific function, and then determine the target expert network group corresponding to the initial speech representation according to the score or weight of the initial speech representation in each target expert network group.
[0085] As an example, assuming that there are 4 expert network groups in the MoE module, when the routing network receives the initial speech representation, the routing network can generate weight vectors corresponding to the 4 expert network groups through linear transformation and Softmax function. The 4 weights included in the weight vector correspond to the 4 expert network groups, such as [0.1, 0.7, 0.15, 0.05]. Then, a target expert network group with the largest weight (the expert network group corresponding to 0.7) or two target expert network groups (the target expert network groups corresponding to 0.7 and 0.15) can be selected, and the initial speech representation can be sent to the target expert network group.
[0086] In an embodiment of the present application, an electronic device can obtain a speech to be recognized, and recognize the speech through a hybrid speech recognition model to obtain a speech recognition result, wherein the hybrid speech recognition model includes a target encoding module, an MoE module and a target decoding module, the target encoding module can be used to encode the speech to be recognized to obtain an initial speech representation corresponding to the speech to be recognized, the MoE module includes a routing network and multiple expert network groups corresponding to multiple languages; the routing network can be used to determine a target expert network group corresponding to the initial speech representation from multiple expert network groups according to the initial speech representation, and process the initial speech representation through the target expert network group to obtain a target speech representation; the target decoding module is used to perform decoding processing according to the target speech representation to obtain a speech recognition result. Through this embodiment, multiple expert network groups can be obtained in MoE by grouping according to language types. The routing module in MoE can route the initial speech representation corresponding to the speech to be recognized to the corresponding target expert network group. The language type contained in the speech to be recognized can correspond to the language type corresponding to the target expert network. The initial speech representation is then processed by the target expert network group, which can accurately capture the speech features of different language types to obtain the target speech representation, reduce confusion and misrecognition of language types, and improve the accuracy of the target speech representation, thereby improving the recognition accuracy of mixed speech containing multiple language types.
[0087] In order to further select a suitable expert network to process the initial speech representation, in one embodiment, each expert network group may include a routing unit and multiple expert networks. The routing unit is used to determine one or more target expert networks from multiple expert networks according to the initial speech representation, and process the initial speech representation through the one or more target expert networks to obtain the target speech representation.
[0088] Among them, the routing unit in the target expert network group can determine one or more target expert networks corresponding to the initial speech representation from the multiple expert networks included in the target expert network group based on the initial speech representation, and activate the one or more expert networks, so that the activated one or more expert networks process the initial speech representation. Optionally, the routing unit can also include a preset specific function, and the specific function included in the routing unit can be the same as the specific function included in the routing network, such as the Softmax function, so as to calculate the score or weight of the initial speech representation in each expert network of the target expert network group according to the specific function, and then determine the one or more expert networks corresponding to the initial speech representation according to the score or weight of the initial speech representation in each expert network of the target expert network group.
[0089] By implementing this embodiment, a suitable target expert network can be selected from the target expert network group through the routing unit. By routing the determined target expert network twice through the routing network and the routing unit, the accuracy of the determined target speech representation can be further improved to improve the accuracy of speech recognition.
[0090] like Figure 4 As shown, Figure 4 is a schematic diagram of the structure of a MoE module disclosed in an embodiment of the present application, wherein the routing network can represent the input initial speech H pre , split into the initial phonetic representation of Chinese and the initial phonological representation of English Then the initial phonetic representation of Chinese Routed to the Chinese expert network group and the initial English speech representation Routed to the English Expert Network Group.
[0091] Among them, the input to the initial speech representation H pre The initial speech representation H corresponding to each frame of sub-speech data may be included. The MoE module may determine the language type corresponding to each frame of sub-speech data, thereby converting the input initial speech representation H pre , split into the initial phonetic representation of Chinese and the initial phonological representation of English The multi-frame sub-speech data may be obtained after the speech to be recognized is subjected to frame processing.
[0092] The routing unit in the Chinese expert network group then uses the initial phonetic representation of the Chinese From the n expert network groups of the Chinese expert network group, one or more target expert networks of the Chinese expert network group are determined, so that the initial speech representation of the Chinese is performed according to the one or more target expert networks of the Chinese expert network group. After processing, weighted fusion is performed to obtain the output data of the Chinese expert network group, that is, the fused speech representation corresponding to the initial speech representation of Chinese
[0093] The routing unit in the English expert network group then uses the initial speech representation of the English From the n expert network groups of the English expert network group, one or more target expert networks of the English expert network group are determined, so that the initial speech representation of English is performed according to the one or more target expert networks of the English expert network group. After processing, weighted fusion is performed to obtain the output data of the English expert network group, that is, the fused speech representation corresponding to the initial speech representation of English
[0094] The MoE module then merges the output data of the Chinese expert network group with the output data of the English expert network group. and By connecting them, we can get the target speech representation H corresponding to the speech to be recognized output by the MoE module out .
[0095] In order to explain the specific process of speech recognition more clearly, Figure 5 As shown, Figure 5 : is a flow chart of a speech recognition method disclosed in an embodiment of the present application. The speech recognition method can be applied to the electronic device in the above embodiment. The speech recognition method may include the following steps:
[0096] Step 510: Acquire speech to be recognized.
[0097] Step 520, dividing the speech to be recognized into frames to obtain multiple frames of sub-speech data.
[0098] Optionally, the framing process may be performed according to one of the following methods: fixed-length framing, adaptive framing, and windowed framing. The lengths of the obtained multi-frame sub-voice data may be the same or different, and there is no limitation on this.
[0099] In the fixed-length framing method, the electronic device can divide the speech to be recognized into multiple frames according to a fixed time length, and there can be overlapping speech data between two adjacent frames of sub-speech data to ensure the continuity and information integrity between frames. As an example, if the speech to be recognized is 1 second, the fixed time length can be set to 25 milliseconds, and the time length of the overlapping speech data can be 10 milliseconds. Then the first frame of sub-speech data can include speech data from the 0th millisecond to the 25th millisecond, and the second frame of sub-speech data can include speech data from the 10th millisecond to the 35th millisecond, and so on, until the division of the speech data to be recognized is completed. The fixed-length framing method is simple to calculate, efficient, and has a wide range of applicability.
[0100] In the adaptive framing method, the electronic device can dynamically adjust the time length of each frame of sub-speech data according to the acoustic characteristics of the speech to be recognized. For example, vowels and consonants can use frames of different time lengths, thereby improving the uniformity of features contained in a frame of sub-speech data.
[0101] In the windowing and framing method, the electronic device can divide the speech to be recognized into multiple frames according to a fixed time length, and process each frame of speech using a specific windowing function, such as a Hamming window, to obtain multiple frames of sub-speech data, thereby reducing the signal edge effect caused by the framing process and reducing the spectrum leakage phenomenon.
[0102] Step 530: encode the multiple frames of sub-speech data through a target encoding module to obtain initial speech representations corresponding to the multiple frames of sub-speech data.
[0103] As an example, the target encoding module may include CNN and transformer. CNN may perform feature extraction on multiple frames of sub-speech data respectively, extract speech features of each frame of sub-speech data through multi-layer convolution operations, and then input the speech features corresponding to each frame of sub-speech data into the transformer. The transformer may utilize the self-attention mechanism to process the speech features corresponding to multiple frames of sub-speech data respectively, and generate initial speech representations corresponding to the multiple frames of sub-speech respectively.
[0104] Optionally, the feature extraction network may also include one of CNN, long short-term memory network, and ResNet feature extraction networks, and the encoder may also include one of transformer and autoencoder. The types of feature extraction networks and encoders in the embodiments of the present application are not limited.
[0105] Step 540, determining a target expert network group corresponding to the target frame sub-speech data from a plurality of expert network groups through a routing network according to the initial speech representation corresponding to the target frame sub-speech data; the target frame sub-speech data is any frame sub-speech data.
[0106] Optionally, the electronic device may input the initial voice representation corresponding to each frame of sub-voice data into the MoE module in chronological order, so that the MoE module processes the initial voice representation corresponding to each frame of sub-voice data in sequence. Optionally, the electronic device may also input the initial voice representation corresponding to multiple frames of sub-voice data into the MoE module together, and the MoE module may process the initial voice representation corresponding to multiple frames of sub-voice data in chronological order or randomly. This embodiment of the application is not limited to this.
[0107] In one embodiment, the electronic device can determine the target language type corresponding to the target frame sub-speech data based on the initial speech representation corresponding to the target frame sub-speech data through a routing network, and determine the expert network group corresponding to the target language type as the target expert network group corresponding to the target frame sub-speech data.
[0108] The routing network is a deep neural network, which can learn the language classification ability of the initial speech representation during the training process, so that in the speech recognition scenario, the target language corresponding to the target frame sub-speech data can be determined according to the initial speech representation corresponding to the target frame sub-speech data. Optionally, the routing network can calculate the probability that the initial speech representation corresponding to the target frame sub-speech data belongs to various language categories, and select the language category corresponding to the largest probability as the target language category.
[0109] By implementing this embodiment, the electronic device can determine the target language type corresponding to the target sub-speech data, and then determine the expert network group corresponding to the target language type as the target expert network group corresponding to the target frame sub-speech data, which can improve the accuracy of the determined target expert network group.
[0110] Step 550 : determining one or more target expert networks from the target expert network group through a routing unit in the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data.
[0111] Optionally, the target expert network group may be preset with a preset number of expert networks. The electronic device may determine a preset number of target expert networks from the target expert network group through a routing unit in the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data. The electronic device may also determine one or more target expert networks from the target expert network group through a routing unit in the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data, and the number of one or more target expert networks is not greater than a preset number. By implementing this embodiment, an appropriate number of expert networks can be selected to process the initial speech representation, which can reduce computational complexity and improve processing efficiency of the initial speech representation.
[0112] Step 560: Process the initial speech representation corresponding to the target frame sub-speech data through one or more target expert networks to obtain a fused speech representation corresponding to the target frame sub-speech data.
[0113] Among them, one or more target expert networks respectively correspond to network weights, and the network weights can be determined by the routing unit of the target expert network group or can be preset. Each target expert network can process the initial speech representation corresponding to the target frame sub-speech data to obtain the processing results corresponding to each target expert network, and then weightedly fuse the processing results corresponding to each target expert network to obtain the fused speech representation corresponding to the target frame sub-speech data.
[0114] In one embodiment, step 550 may include: determining one or more target expert networks and network weights corresponding to each target expert network from the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data through the routing unit in the target expert network group, and step 560 may include: processing the initial speech representation corresponding to the target frame sub-speech data through one or more target expert networks respectively to obtain processing results corresponding to each target expert network; based on the network weights corresponding to each target expert network, weighted fusion of the processing results corresponding to each target expert network to obtain a fused speech representation corresponding to the target frame sub-speech data.
[0115] Among them, the routing unit can also be a deep neural network, which can learn the representation processing capabilities of each target expert network for speech representation during the training process, so that in the speech recognition scenario, one or more target expert networks and the network weights corresponding to each target expert network can be determined according to the initial speech representation corresponding to the target frame sub-speech data. The network weight can be used to characterize the representation processing effect of the target expert network on the initial speech representation. The larger the network weight, the better the representation processing capability of the target expert network for the representation processing effect of the initial speech representation.
[0116] For example, in the Chinese expert network group, two target expert networks can be included. The first target expert network has a stronger representation processing capability for the speech representation corresponding to the Cantonese accent, and the second target expert network has a stronger representation processing capability for the speech representation corresponding to the Hong Kong accent. If the routing unit analyzes that the initial speech representation is closer to the speech representation corresponding to the Cantonese accent, the network weight of the first target expert network is determined to be greater than the network weight of the second target expert network.
[0117] By implementing this embodiment, the electronic device can perform weighted fusion on the processing results corresponding to each target expert network based on the network weights corresponding to each target expert network determined by the routing unit, and obtain a fused speech representation corresponding to the target frame sub-speech data, thereby improving the accuracy of the fused speech representation. The comprehensive network performance of the target expert network group is improved through the dynamic weight method, thereby improving the accuracy of speech recognition.
[0118] Step 570: Obtain a target speech representation based on the fused speech representations corresponding to the multiple frames of sub-speech data.
[0119] After the MoE completes processing of the multiple frames of sub-speech data, the electronic device may connect the fused speech representations corresponding to the multiple frames of sub-speech data to obtain the target speech representation.
[0120] Step 580: decode the target speech representation through the target decoding module to obtain a speech recognition result.
[0121] In an embodiment of the present application, an electronic device can obtain a speech to be recognized, and perform frame processing on the speech to be recognized to obtain multiple frames of sub-speech data, thereby processing the initial speech representation corresponding to each frame of sub-speech data, and finally obtaining the target speech representation based on the fused speech representation corresponding to each frame of sub-speech data. Compared with generating a speech representation of a long segment of speech, the method of dividing the speech into frames can reduce the computational complexity. Moreover, for mixed speech containing multiple languages, the frame processing can make each frame of sub-speech data include only one language type, thereby improving the convenience and accuracy of speech recognition.
[0122] The above content explains the speech recognition process, and the following content explains the training process of the hybrid speech recognition model, wherein the hybrid speech recognition model can be obtained through one or more trainings.
[0123] In one embodiment, the target encoding module is obtained by self-supervised training based on the first data set and supervised training based on the second data set, and the MoE module and the target decoding module are obtained by supervised training based on the second data set. The first data set includes a plurality of unlabeled first sample speech, each of which corresponds to a language type, and the second data set includes a plurality of second sample speech and labeling information corresponding to each second sample speech, each of which corresponds to one or more language types.
[0124] In the process of self-supervised training, the target coding module can mine multiple unlabeled first sample voices included in the first data set through preset auxiliary tasks, generate discretized labels corresponding to each first sample voice, and then train the target coding module according to the discretized labels corresponding to each first sample voice to minimize the prediction error and improve the coding ability of the target coding module until the self-supervised training is completed. Among them, the target coding module can mask some features of the first sample voice, such as masking a frame of voice in the first sample voice, train the target coding module to predict the masked features, and then use the prediction results and convert them into discretized labels.
[0125] During the supervised training process, multiple second sample voices can be input into the hybrid voice recognition model, and each second sample voice can be recognized by the hybrid voice recognition model to obtain the sample recognition results corresponding to each second sample voice. According to the sample recognition results corresponding to each first sample voice and the corresponding sample information, the target loss corresponding to each first sample voice can be obtained, so that the parameters of the target encoding module, MoE module and target decoding module can be adjusted according to the target loss corresponding to each first sample voice until the supervised training is completed. Optionally, the annotation information can be the text corresponding to the second sample voice.
[0126] It is understandable that, due to the different training purposes, self-supervised training is to improve the feature extraction ability of the target encoding module, while supervised training is to adapt the hybrid speech recognition module to a multilingual environment. Therefore, the first sample speech can correspond to only one language type, while the second sample speech can correspond to one or more speech types. For example, two first sample speech can correspond to Chinese and English respectively, while one second sample speech can contain both Chinese and English. In addition, the number of second sample speech in the second data set can be much smaller than the number of first sample speech in the first data set, reducing the cost of manual annotation.
[0127] By implementing this embodiment, the target coding module can perform one self-supervised training and one supervised training. The target coding module can be capable of extracting speech features through self-supervised training, and can then be adapted to mixed speech recognition scenarios and improve speech recognition accuracy through supervised training.
[0128] like Figure 6 As shown, Figure 6 : is a flow chart of a training process of a target coding module disclosed in an embodiment of the present application, wherein:
[0129] Step 610: Based on the first data set, perform self-supervisory training on the initial encoding module to obtain a first encoding module, where the first encoding module includes multiple encoding layers.
[0130] Among them, the encoding layer may refer to a transformer layer, and the first encoding module library includes multiple layers of transformer layers.
[0131] Step 620, supervised training is performed on the second encoding module based on the second data set to obtain a target encoding module; the second encoding module is obtained by adding a conditional decoding layer between every two adjacent encoding layers in the first encoding module.
[0132] The purpose of self-supervised training is to improve the ability of multi-layer encoding layers to extract speech features. The decoder can be omitted at this stage to avoid overfitting the distribution of specific tasks or specific data. However, the purpose of supervised training is to allow the speech recognition model to learn the precise mapping relationship in the speech recognition task of mixed speech through a small amount of first sample speech and the annotation information corresponding to the first sample speech. At this time, introducing the decoder to optimize the framework of the encoding module will help improve the performance of the mixed speech recognition model.
[0133] The conditional decoding layer is used to calculate the posterior probability of the connected previous encoding layer. During the supervised training process, the posterior probability calculated by the conditional decoding can be connected with the speech representation extracted by the previous encoding layer, and the connected information is input into the next encoding layer.
[0134] In an embodiment of the present application, the electronic device can optimize the structure of the coding module after completing self-supervised training. By adding a conditional decoding layer between each two adjacent coding layers, the conditional independence of the target coding module can be alleviated to improve the performance of the speech recognition model.
[0135] In one embodiment, the electronic device can normalize the first speech representation extracted by the current coding layer to obtain a second speech representation, and calculate the posterior probability corresponding to the second speech representation through the conditional decoding layer connected to the current coding layer, and then linearly project the posterior probability through the linear projection layer so that the dimension of the posterior probability is aligned with the dimension of the first speech representation, and then perform a residual connection between the mapped posterior probability and the first speech representation to obtain a third speech representation, and input the third speech representation to the next coding layer of the current coding layer.
[0136] The current coding layer may refer to any layer except the last coding layer in the multi-layer coding layer. In this embodiment, the posterior probability after linear projection is residually connected with the first speech representation, so that each coding layer can obtain the original information of the previous coding layer, which helps the model learn cross-layer features and enhance context dependence.
[0137] like Figure 7 As shown, Figure 7 It is a schematic diagram of a target coding module disclosed in an embodiment of the present application, wherein the coding layer of the target coding module is a transformer layer, the layer normalization algorithm is in Layer Norm, the conditional decoding layer is a conditional CTC decoder, the current transformer layer can output a first speech representation to Layer Norm, Layer Norm normalizes the first speech representation to obtain a second speech representation, the second speech representation is then input into the conditional CTC decoder to obtain a posterior probability, the linear projection layer then maps the posterior probability to a latitudinal space corresponding to the second speech representation, and performs a residual connection between the mapped posterior probability and the first speech representation to obtain a third speech representation, and the third speech representation is input into the next transformer layer.
[0138] By implementing this embodiment, the optimized model structure can improve the recognition capability of mixed speech containing multiple languages and improve the accuracy of speech recognition.
[0139] like Figure 8 As shown, Figure 8 8 is a modular schematic diagram of a speech recognition device disclosed in an embodiment of the present application. The speech recognition device 800 can be applied to the electronic device in the above embodiment. The speech recognition device 800 may include a speech acquisition module 810 and a speech recognition module 820, wherein:
[0140] The speech acquisition module 810 is used to acquire the speech to be recognized;
[0141] The speech recognition module 820 is used to recognize the speech to be recognized through a hybrid speech recognition model to obtain a speech recognition result; the hybrid speech recognition model includes a target encoding module, a hybrid expert model MoE module and a target decoding module, the target encoding module is used to encode the speech to be recognized to obtain an initial speech representation corresponding to the speech to be recognized; the MoE module includes a routing network and multiple expert network groups corresponding to multiple language types, the routing network is used to determine the target expert network group corresponding to the initial speech representation from multiple expert network groups according to the initial speech representation, and process the initial speech representation through the target expert network group to obtain a target speech representation; the target decoding module is used to perform decoding processing according to the target speech representation to obtain a speech recognition result.
[0142] In one embodiment, each expert network group includes a routing unit and multiple expert networks; the routing unit is used to determine one or more target expert networks from the multiple expert networks based on the initial speech representation, and process the initial speech representation through the one or more target expert networks to obtain the target speech representation.
[0143] In one embodiment, the speech acquisition module 810 is also used to perform frame processing on the speech to be recognized to obtain multiple frames of sub-speech data; the speech recognition module 820 is also used to encode the multiple frames of sub-speech data through the target encoding module to obtain initial speech representations corresponding to the multiple frames of sub-speech; determine the target expert network group corresponding to the target frame sub-speech data from multiple expert network groups through the routing network according to the initial speech representation corresponding to the target frame sub-speech data; the target frame sub-speech data is any frame of sub-speech data; determine one or more target expert networks from the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data through the routing unit in the target expert network group; process the initial speech representation corresponding to the target frame sub-speech data through one or more target expert networks to obtain a fused speech representation corresponding to the target frame sub-speech data; obtain the target speech representation according to the fused speech representations corresponding to the multiple frames of sub-speech data; decode the target speech representation through the target decoding module to obtain a speech recognition result.
[0144] In one embodiment, the speech recognition module 820 is also used to determine the target language type corresponding to the target frame sub-speech data based on the initial speech representation corresponding to the target frame sub-speech data through the routing network, and determine the expert network group corresponding to the target language type as the target expert network group corresponding to the target frame sub-speech data.
[0145] In one embodiment, the speech recognition module 820 is also used to determine one or more target expert networks and the network weights corresponding to each target expert network from the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data through the routing unit in the target expert network group; process the initial speech representation corresponding to the target frame sub-speech data through one or more target expert networks respectively to obtain the processing results corresponding to each target expert network; based on the network weights corresponding to each target expert network, perform weighted fusion on the processing results corresponding to each target expert network to obtain the fused speech representation corresponding to the target frame sub-speech data.
[0146] In one embodiment, the target encoding module is obtained by self-supervised training based on the first data set and supervised training based on the second data set, and the MoE module and the target decoding module are obtained by supervised training based on the second data set; the first data set includes multiple unlabeled first sample speech, each first sample speech corresponds to a language type, and the second data set includes multiple second sample speech and labeling information corresponding to each second sample speech, each second sample speech corresponds to one or more language types.
[0147] In one embodiment, the speech recognition device may also include a model training module, which is used to perform self-supervised training on the initial coding module based on the first data set to obtain a first coding module, wherein the first coding module includes multiple coding layers; and perform supervised training on the second coding module based on the second data set to obtain a target coding module; the second coding module is obtained by adding a conditional decoding layer between each two adjacent coding layers in the first coding module; the conditional decoding layer is used to calculate the posterior probability of the connected previous coding layer.
[0148] In one embodiment, the model training module is also used to normalize the first speech representation extracted by the current coding layer to obtain a second speech representation; calculate the posterior probability corresponding to the second speech representation through the conditional decoding layer connected to the current coding layer; linearly project the posterior probability through the linear projection layer so that the dimension of the posterior probability is aligned with the dimension of the second speech representation; perform a residual connection between the mapped posterior probability and the second speech representation to obtain a third speech representation, and input the third speech representation to the next coding layer of the current coding layer.
[0149] In an embodiment of the present application, an electronic device can obtain a speech to be recognized, and recognize the speech through a hybrid speech recognition model to obtain a speech recognition result, wherein the hybrid speech recognition model includes a target encoding module, an MoE module and a target decoding module, the target encoding module can be used to encode the speech to be recognized to obtain an initial speech representation corresponding to the speech to be recognized, the MoE module includes a routing network and multiple expert network groups corresponding to multiple languages; the routing network can be used to determine a target expert network group corresponding to the initial speech representation from multiple expert network groups according to the initial speech representation, and process the initial speech representation through the target expert network group to obtain a target speech representation; the target decoding module is used to perform decoding processing according to the target speech representation to obtain a speech recognition result. Through this embodiment, multiple expert network groups can be obtained in MoE by grouping according to language types. The routing module in MoE can route the initial speech representation corresponding to the speech to be recognized to the corresponding target expert network group. The language type contained in the speech to be recognized can correspond to the language type corresponding to the target expert network. The initial speech representation is then processed by the target expert network group, which can accurately capture the speech features of different language types to obtain the target speech representation, reduce confusion and misrecognition of language types, and improve the accuracy of the target speech representation, thereby improving the recognition accuracy of mixed speech containing multiple language types.
[0150] like Fig. 9 As shown, in one embodiment, an electronic device is provided, which may include:
[0151] A memory 910 storing executable program codes;
[0152] a processor 920 coupled to the memory 910;
[0153] The processor 920 calls the executable program code stored in the memory 910 to implement the speech recognition method provided in the above embodiments.
[0154] The memory 910 may include a random access memory (RAM) or a read-only memory (ROM). The memory 910 may be used to store instructions, programs, codes, code sets or instruction sets. The memory 910 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc. The data storage area may also store data created by the electronic device during use, etc.
[0155] The processor 920 may include one or more processing cores. The processor 920 uses various interfaces and lines to connect various parts of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 910, and calling data stored in the memory 910. Optionally, the processor 920 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 920 can integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 920, but may be implemented separately through a communication chip.
[0156] It is understandable that the electronic device may include more or fewer structural elements than those in the above structural block diagram, for example, a power module, physical buttons, a WiFi (Wireless Fidelity) module, a speaker, a Bluetooth module, a sensor, etc., and no limitation is made here.
[0157] An embodiment of the present application discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the methods described in the above embodiments.
[0158] In addition, an embodiment of the present application further discloses a computer program product. When the computer program product is run on a computer, the computer can execute all or part of the steps in any one of the speech recognition methods described in the above embodiments.
[0159] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0160] The above is a detailed introduction to a speech recognition method, device, electronic device and storage medium disclosed in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A speech recognition method, characterized in that: include: Get the speech to be recognized; Recognize the speech to be recognized by using a hybrid speech recognition model to obtain a speech recognition result; The hybrid speech recognition model includes a target encoding module, a hybrid expert model MoE module and a target decoding module. The target encoding module is used to encode the speech to be recognized to obtain an initial speech representation corresponding to the speech to be recognized; the MoE module includes a routing network and multiple expert network groups corresponding to multiple language types. The routing network is used to determine the target expert network group corresponding to the initial speech representation from the multiple expert network groups according to the initial speech representation, and process the initial speech representation through the target expert network group to obtain a target speech representation; the target decoding module is used to perform decoding processing according to the target speech representation to obtain the speech recognition result.
2. The method according to claim 1, characterized in that Each of the expert network groups includes a routing unit and multiple expert networks; the routing unit is used to determine one or more target expert networks from the multiple expert networks based on the initial speech representation, and process the initial speech representation through the one or more target expert networks to obtain a target speech representation.
3. The method according to claim 2, characterized in that After acquiring the speech to be recognized, the method further includes: Performing frame processing on the speech to be recognized to obtain multiple frames of sub-speech data; The step of recognizing the speech to be recognized by using the hybrid speech recognition model to obtain a speech recognition result includes: Encoding the multiple frames of sub-speech data by the target encoding module to obtain initial speech representations corresponding to the multiple frames of sub-speech respectively; Determining, by the routing network, a target expert network group corresponding to the target frame sub-speech data from the plurality of expert network groups according to the initial speech representation corresponding to the target frame sub-speech data; the target frame sub-speech data is any frame sub-speech data; Determining one or more target expert networks from the target expert network group by a routing unit in the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data; Processing the initial speech representation corresponding to the target frame sub-speech data through the one or more target expert networks to obtain a fused speech representation corresponding to the target frame sub-speech data; Obtaining a target speech representation according to the fused speech representations respectively corresponding to the multiple frames of sub-speech data; The target speech representation is decoded by the target decoding module to obtain a speech recognition result.
4. The method according to claim 3, characterized in that The step of determining, by the routing network according to the initial speech representation corresponding to the target frame sub-speech data, from the plurality of expert network groups a target expert network group corresponding to the target frame sub-speech data comprises: The routing network determines the target language type corresponding to the target frame sub-speech data according to the initial speech representation corresponding to the target frame sub-speech data, and determines the expert network group corresponding to the target language type as the target expert network group corresponding to the target frame sub-speech data.
5. The method according to claim 3, characterized in that: The determining one or more target expert networks from the target expert network group by a routing unit in the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data comprises: Determining, by a routing unit in the target expert network group, one or more target expert networks and network weights corresponding to each of the target expert networks from the target expert network group according to the initial speech representation corresponding to the target frame sub-speech data; The processing of the initial speech representation corresponding to the target frame sub-speech data by the one or more target expert networks to obtain the fused speech representation corresponding to the target frame sub-speech data includes: Processing the initial speech representations corresponding to the target frame sub-speech data respectively through the one or more target expert networks to obtain processing results corresponding to each of the target expert networks; Based on the network weights corresponding to the target expert networks, the processing results corresponding to the target expert networks are weightedly fused to obtain the fused speech representation corresponding to the target frame sub-speech data.
6. The method according to claim 1, characterized in that The target encoding module is obtained by self-supervised training based on the first data set and supervised training based on the second data set, and the MoE module and the target decoding module are obtained by supervised training based on the second data set; The first data set includes a plurality of unlabeled first sample speech, each of which corresponds to a language type; the second data set includes a plurality of second sample speech and labeling information corresponding to each of the second sample speech, each of which corresponds to one or more language types.
7. The method according to claim 6, characterized in that The method further comprises: Based on the first data set, self-supervised training is performed on the initial encoding module to obtain a first encoding module, wherein the first encoding module includes multiple encoding layers; The second encoding module is supervisedly trained based on the second data set to obtain a target encoding module; the second encoding module is obtained by adding a conditional decoding layer between every two adjacent encoding layers in the first encoding module; the conditional decoding layer is used to calculate the posterior probability of the connected previous encoding layer.
8. The method according to claim 7, characterized in that The method further comprises: Normalizing the first speech representation extracted by the current coding layer to obtain a second speech representation; Calculating, by means of a conditional decoding layer connected to the current encoding layer, a posterior probability corresponding to the second speech representation; Performing a linear projection on the posterior probability through a linear projection layer so that the dimension of the posterior probability is aligned with the dimension of the second speech representation; The mapped posterior probability is residually connected with the second speech representation to obtain a third speech representation, and the third speech representation is input to a coding layer next to the current coding layer.
9. A speech recognition device, characterized in that: include: A speech acquisition module, used to acquire the speech to be recognized; A speech recognition module is used to recognize the speech to be recognized through a hybrid speech recognition model to obtain a speech recognition result; the hybrid speech recognition model includes a target encoding module, a hybrid expert model MoE module and a target decoding module, the target encoding module is used to encode the speech to be recognized to obtain an initial speech representation corresponding to the speech to be recognized; the MoE module includes a routing network and a plurality of expert network groups corresponding to a plurality of language types, the routing network is used to determine a target expert network group corresponding to the initial speech representation from the plurality of expert network groups according to the initial speech representation, and process the initial speech representation through the target expert network group to obtain a target speech representation; the target decoding module is used to perform decoding processing according to the target speech representation to obtain the speech recognition result.
10. An electronic device, characterized in that: include: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Generation method and device of mixed language speech recognition model
CN115064154A
End-to-end multi-language speech recognition method based on hybrid expert model
CN115457942A
Speech recognition method and device, speech recognition model training method and device, medium and equipment
CN116013257A
Low-resource voice keyword detection method based on unsupervised learning and transfer learning
CN116434742A
Speech recognition model training method and device, computer equipment and storage medium
CN116913254A
Cited By
Zip-MoE model grouping mixed expert layer-based Chinese and English speech recognition method and system
CN120126451A
Speech recognition method and device
CN120319222A
Multi-language speech recognition method and system based on edge computing power
CN120877712A