Multilingual speech processing

The multilingual processing model extends language coverage by incorporating a supplementary network trained on new languages, ensuring effective processing of both existing and new languages without compromising performance.

WO2025145407A1PCT designated stage expired Publication Date: 2025-07-10DOUYIN VISION CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/070695
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-04
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing multilingual speech processing models are limited to a small subset of languages and struggle to effectively process new languages without degrading performance on existing languages.

Method used

A multilingual processing model is enhanced by adding a supplementary network trained on new languages, allowing it to process both existing and new languages without updating parameters of the original network, thus preserving performance on existing languages.

Benefits of technology

The solution enables efficient extension of language coverage to new languages with minimal computational resources and without performance degradation on existing languages, reducing the time and data requirements for supporting new languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024070695_10072025_PF_FP_ABST
    Figure CN2024070695_10072025_PF_FP_ABST
Patent Text Reader

Abstract

There are proposed methods, devices, and computer program products for multilingual speech processing. In the method, in response to receiving speech data, a network corresponding to the speech data is identified from a first and a second network that are comprised in a multilingual processing model, the first network being associated with a first group of languages and the second network being associated with a second group of languages. An output that is determined by the identified network based on the speech data is provided as an output of the multilingual processing model. With these implementations, in addition to the first group of languages covered by the first network, language coverage of the multilingual processing model may be extended by the second group of languages covered by the second network.
Need to check novelty before this filing date? Find Prior Art

Description

MULTILINGUAL SPEECH PROCESSINGFIELD

[0001] The present disclosure generally relates to multilingual speech processing, and more specifically, to methods, devices, and computer program products for language coverage extension for multilingual tasks using a supplementary network associated with new language (s) .BACKGROUND

[0002] The machine learning technology is widely used in multilingual speech processing, and various machine learning models have been developed for various tasks such as the multilingual Automatic Speech Recognition (mASR) , language identification, emotion detection, and the like from the speech data. However, there are thousands of languages, while existing multilingual processing models are trained by only a small portion of existing languages among all the languages. Accordingly, these models can only process speech data represented in the existing languages but become useless when handling a new language other than the existing languages. At this point, it is desired to extending the coverage of the model to new language (s) .SUMMARY

[0003] In a first aspect of the present disclosure, there is provided a method for multilingual speech processing. In the method, in response to receiving speech data, a network corresponding to the speech data is identified from a first and a second network that are comprised in a multilingual processing model, the first network being associated with a first group of languages and the second network being associated with a second group of languages. An output that is determined by the identified network based on the speech data is provided as an output of the multilingual processing model.

[0004] In a second aspect of the present disclosure, there is provided an electronic device. The electronic device comprises: a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements a method according to the first aspect of the present disclosure.

[0005] In a third aspect of the present disclosure, there is provided a computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by an  electronic device to cause the electronic device to perform a method according to the first aspect of the present disclosure.

[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0007] BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS

[0008] Through the more detailed description of some implementations of the present disclosure in the accompanying drawings, the above and other objects, features and advantages of the present disclosure will become more apparent, wherein the same reference generally refers to the same components in the implementations of the present disclosure.

[0009] Fig. 1 illustrates an example environment for multilingual speech processing according to the machine learning technique;

[0010] Fig. 2 illustrates an example diagram for multilingual speech processing by using a supplementary network associated with new language (s) according to implementations of the present disclosure;

[0011] Fig. 3 illustrates an example diagram for a multilingual processing model according to implementations of the present disclosure;

[0012] Fig. 4 illustrates an example diagram for supplementary components in the multilingual processing model according to implementations of the present disclosure;

[0013] Fig. 5 illustrates an example diagram for results of multilingual automatic speech recognition according to implementations of the present disclosure;

[0014] Fig. 6 illustrates an example flowchart of a method for multilingual speech processing according to implementations of the present disclosure; and

[0015] Fig. 7 illustrates a block diagram of a computing device in which various implementations of the present disclosure can be implemented.DETAILED DESCRIPTION

[0016] Principle of the present disclosure will now be described with reference to some implementations. It is to be understood that these implementations are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below.

[0017] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0018] References in the present disclosure to “one implementation, ” “an implementation, ” “an example implementation, ” and the like indicate that the implementation described may include a particular feature, structure, or characteristic, but it is not necessary that every implementation includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same implementation. Further, when a particular feature, structure, or characteristic is described in connection with an example implementation, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other implementations whether or not explicitly described.

[0019] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example implementations. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0020] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of example implementations. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.

[0021] Principle of the present disclosure will now be described with reference to some implementations. It is to be understood that these implementations are described only for the purpose of illustration and help those skilled in the art to understand and implement the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below. In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0022] It may be understood that data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with requirements of corresponding laws and regulations and relevant rules.

[0023] It may be understood that, before using the technical solutions disclosed in various implementation of the present disclosure, the user should be informed of the type, scope of use, and use scenario of the personal information involved in the present disclosure in an appropriate manner in accordance with relevant laws and regulations, and the user’s authorization should be obtained.

[0024] For example, in response to receiving an active request from the user, prompt information is sent to the user to explicitly inform the user that the requested operation will need to acquire and use the user’s personal information. Therefore, the user may independently choose, according to the prompt information, whether to provide the personal information to software or hardware such as electronic devices, applications, servers, or storage media that perform operations of the technical solutions of the present disclosure.

[0025] As an optional but non-limiting implementation, in response to receiving an active request from the user, the way of sending prompt information to the user, for example, may include a pop-up window, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose “agree” or “disagree” to provide the personal information to the electronic device.

[0026] It may be understood that the above process of notifying and obtaining the user authorization is only illustrative and does not limit the implementation of the present disclosure. Other methods that satisfy relevant laws and regulations are also applicable to the implementation of the present disclosure.

[0027] Nowadays, the machine learning technology is widely used in the multilingual speech processing, and various machine learning models have been developed for various tasks. Referring to Fig. 1 for a brief of the multilingual speech processing, here Fig. 1 illustrates an example environment 100 for the multilingual speech processing according to the machine learning technique. As illustrated in Fig. 1, a multilingual processing model 130 may be trained for processing speech data 110, and then providing output 120 based on the inputted speech data 110. Here, the speech data 110 may be a voice input including the speech in various natural languages such as English, French, German, and the like. Herein, the speech data 110 may be represented in various formats, such as an audio file, an audio stream, an audio component in a video file  / stream, and the like.

[0028] Here, the multilingual processing model 130 may include an encoder 132 and a decoder 134. The encoder 132 may convert the speech data 110 into an internal representation (such as an embedding) , and then the decoder 134 may convert the internal representation into the output 120. Based on a task of the multilingual processing model 130, the output 120 may comprise various types, such as a text (i.e., an identified transcript) , a language identification, an emotion, a domain, a topic of the speech data 110, and the like.

[0029] For a purpose of description, an mASR task may be taken as an example of the tasks implemented by the multilingual processing model 130. In the mASR task, a transcript in the text format may be identified from the speech data 110. Although existing mASR models have been scaled to encompass over a thousand languages, the challenge of extending them to new languages under the following constraints remains unresolved. For example, training data for existing languages is limited or even unavailable, there is no prior knowledge of the input audio, the models need to be parameter-efficient and performance of the models should not degrade on existing languages.

[0030] Various solutions are developed for extending these models to new languages, for example, attachable language-specific adapters require prior knowledge of input audio (i.e., language ID) to deploy the corresponding adapter. On the other hand, lifelong learning approaches, where new language data is combined with existing data to continue model training, assume the availability of training data for existing languages. Moreover, these techniques are considered parameter-inefficient since the entire model needs to be updated. Particularly, for models with over a billion parameters, it leads to huge computational cost to update each time a new language is requested to be supported by the models. Other techniques based on cross-lingual transfer learning, where the existing mASR model is fine-tuned (fully or partially) using new language data, usually result in performance degradation of existing languages due to catastrophic forgetting. At this point, it is desired to extend the model to new language (s) in a more easy and effective way.

[0031] In view of the above, the present disclosure proposes a multilingual speech processing solution by using a new network that is trained by new languages. Generally, with respect to an existing model implemented by a machine learning network (trained by a dataset represented in existing language (s) (i.e., a first group of languages) ) , an additional network may be added for processing speech data represented in new language (s) (i.e., a second group of languages) that are different from the existing language (s) . Here, an encoder and a decoder included in the additional network may be trained by a dataset represented in the new  language (s) , and then the language coverage of the model may be extended to include both of the existing and new languages.

[0032] Referring to Fig. 2 for more details about the proposed multilingual speech processing solution. Fig. 2 illustrates an example diagram 200 for multilingual speech processing by using a supplementary network associated with new language (s) according to implementations of the present disclosure. As illustrated in Fig. 2, a multilingual processing model 230 is proposed in the present disclosure, where the multilingual processing model 230 may include a first network 210 and a second network 220. Specifically, the first network 210 may include an encoder 212 and a decoder 214, both of them are trained by a first dataset that is represented in the first group of languages. Here, the first network 210 may be implemented by an existing mASR model and it may work as a pipeline for converting the speech data 240 in the existing languages into a text 242.

[0033] Further, the second network 220 is added into the multilingual processing model 230 as an additional pipeline for processing the speech data 240. Specifically, the second network 220 may convert the speech data 240 represented in the new languages into a text 244. At this point, the multilingual processing model 230 may process the speech data represented in both of the existing languages as well as the new languages. Once the multilingual processing model 230 is trained, the speech data 240 may be received during the inference stage, and a network may be identified from the first and second networks 210 and 220. Then, an output that is determined by the identified network based on the speech data may be provided as an output of the multilingual processing model 230. In other words, one of the texts 242 and 244 may be selected as the final output of the multilingual processing model 230.

[0034] With these implementations of the present disclosure, the proposed solution does not necessitate the availability of training data for the supported languages, which may be fully or partially unavailable due to various reasons, such as expired data licenses, modifications made by the data owner (e.g., deleted video / audio, changed access rights to private content, etc. ) , data loss, etc. Compared with some existing solutions that require prior knowledge of the input audio, the proposed solution does not require any prior knowledge of the input audio and the model may automatically convert the speech data into an appropriate output. Further, the proposed solution may minimize the impact on the performance of existing languages by keeping the corresponding parameters unchanged, and thus the proposed solution may alleviate the problem of catastrophic forgetting in existing languages, and circumvent the modeling capacity limit by dedicating additional parameters for new languages. Meanwhile, only the second network is updated by the dataset represented in the new languages, and parameters in  the first network are fixed. Therefore, a size of the to-be-trained parameters is limited to be smaller than mASR. As a result, it requires less computational time and resources.

[0035] Having provide a brief description of the proposed solution, the following paragraphs will provide more details about the second network 220. Here, the second network 220 includes an encoder 222 and a decoder 224 for the new languages. To seamlessly extend the mASR model (i.e., the first network 210) to include new languages, the second network 220 adds an additional pipeline for processing the speech data 240, where a portion of the encoder 212 is shared among the multiple encoders 212 and 222.

[0036] In implementations of the present disclosure, one or more supplementary networks may be added into the multilingual processing model 230. For simplicity, only one supplementary network is described as an example encoder-decoder pipeline for the new languages. However, multiple supplementary encoder-decoder pipelines can be added, and these pipelines can be used in parallel or attached / detached depending on the specific working environment.

[0037] In implementations of the present disclosure, first parameters of the first network are trained by a first dataset represented in the first group of languages, and second parameters of the second network are trained by a second dataset represented in the second group of languages. For example, an existing mASR model may be selected as the first network, where the first parameters are already trained. Different from the existing solution, only the second network 220 is trained by the new training data represented in the new languages, while parameters of the first network 210 are fixed and not affected by the new training data during a training procedure of the second parameters of the second network. The second group of languages may be new languages excluded from the first group of languages. With these implementations, parameters of the first network 210 are not updated by the training data represented in the new languages, and thus performance of the first network 210 for processing the existing languages may remain unchanged and will not degrade due to catastrophic forgetting.

[0038] In implementations of the present disclosure, the first network comprises a first encoder (such as the encoder 212) , the second network comprises a second encoder (such as the encoder 222) . Referring to Fig. 3 for more details about the encoders, here Fig. 3 illustrates an example diagram 300 for a multilingual processing model according to implementations of the present disclosure. Here, any pre-trained mASR architecture (such as attention-based encoder-decoder (AED) , connectionist temporal classification (CTC) , RNN-T, and others) may be used for implementing the first and second networks 210 and 220. For the  decoder architecture, the new decoder 224 may be modeled using any network architecture, such as Long Short-Term Memory (LSTM) , Transformers, and so on. The size of the decoder may vary depending on the number of layers / blocks, hidden dimensions, and other hyperparameters. Furthermore, the decoder component may be omitted, as in the CTC architecture. The new decoder may be initialized randomly or transferred from other pre-trained models. In addition to the identified text, the output format of the decoder may be designed to predict other information related to the input test audio, such as language ID, an emotion of the speaker, and the like.

[0039] Similar to the new decoder, the new encoder component can be modeled using any network architecture with varying sizes. It can also be initialized either randomly or from any other pre-trained model. The input feature for the new encoder is sourced from the output or other internal layers / blocks of the original, fixed mASR encoder. This input feature may be preprocessed by other modules as needed. In Fig. 3, the encoder 212 may include n layers: layer 1 (indicated by a block 310) , …, layer n-k (indicated by a block 312) , …, and layer n (indicated by a block 314) . Here, n and k represent integer numbers and 1 < k < n.

[0040] In implementations of the present disclosure, the second encoder shares at least one layer in a plurality of layers that are comprised in the first encoder, and a total number of layers comprised in the first encoder is equal to a total number of layers comprised in the second encoder. As illustrated in Fig. 3, the output of the layer n-k in the encoder 212 is connected to the input of the layer 1 in the encoder 222, and thus the encoder 222 also includes n layers: layer 1 (indicated by a block 310) , …, layer n-k (indicated by a block 312) , …, layer 1 (indicated by a block 320) , …, and layer k (indicated by a block 322) . Specifically, the encoder 212 and the encoder 222 share the following shaded layers in Fig. 3: layer 1 (indicated by a block 310) , …, layer n-k (indicated by a block 312) .

[0041] Although the new languages and the existing languages are different, natural languages often share common features to some extent. With these implementations, the second network 220 may reuse knowledges that are obtained from the existing languages for processing the new languages. Therefore, the accuracy level of the second network 220 may be increased even if the training data represented in the new languages is limited.

[0042] The Transformer layers within the encoder 222 may be initialized from the fixed encoder. For example, to initialize a four block encoder component, the last four blocks of the fixed encoder may be employed. The new encoder is integrated in a manner that preserves the total encoder depth at the same level (i.e., 24 blocks) . In the aforementioned  example, the four block encoder component is connected to the output of the 20th block of the fixed encoder.

[0043] In implementations of the present disclosure, a mapping component may be added between the first and second networks in response to a first data space of the first network being different from a second space of the second network. With these implementations, the present disclosure allows the first and second networks to have different dimensions in their corresponding data spaces, and thus more network architectures may be selected in designing the multilingual processing model 230 in a more flexible way.

[0044] Referring to Fig. 4, Fig. 4 illustrates an example diagram 400 for supplementary components in the multilingual processing model according to implementations of the present disclosure. As illustrated in Fig. 4, a component 410 may be added between the layer n-k in the first network 210 and the layer 1 in the second network 220. This component 410 may be dimensionally adapted from the encoder 212 to the encoder 222. Here, the addition of the component 410 is optional; in other words, the new decoder can be designed to consume the input feature directly from the original encoder or other preprocessing modules. In implementations of the present disclosure, the output feature from the encoder 222 may be further modified (such as, by a component 420) before being sent to the decoder 224. For example, this feature may go through normalization modules (e.g., layer-wise, or batch-wise) , residual connection modules, attention modules, and the like, for fitting to the downstream tasks implemented by the decoder 224.

[0045] The primary limitation of multi-head decoder approach is that the fixed encoder, trained on existing languages, lacks exposure to new languages, leading to suboptimal performance. To address this problem, the additional encoder component is incorporated for the new languages, enabling the learning of relevant acoustic representations. During training, the parameters of the first network 210 are kept fixed to preserve the performance of existing languages. The parameters of the second network 220 are updated using the training data of new languages. Overall, the training procedure follows the standard steps used in deep learning, namely, forward-backward propagation with parameter updates following the learning schedule. In cases where the training data of existing languages is only partially available, the parameters of the mASR may also be updated fully or partially (e.g., encoder only) .

[0046] Once the multilingual processing model 230 is trained, it may output the text corresponding to the inputted speech data 240 in the interference stage. Here, the experiments may involve three scenarios: 1) group-aware, 2) language-aware, and 3) language-agnostic.  In group-aware and the language-aware scenarios, the language classification of the input speech data 240 is known in advance. Here, the language classification indicates a group of languages in the first and second groups to which the speech data belongs. Then, the network may be selected from the first and second networks based on the language classification directly.

[0047] Specifically, in the group-aware scenario, whether the language for the speech data 110 is new or existing is known in advance. At this point, a network corresponding to the language classification may be selected for converting the speech data into a corresponding text. For example, if the speech data is represented in an existing language, the text 242 outputted from the first network 210 works as the final output; else if the speech data is not represented in any existing language, the text 244 outputted from the second network 220 works as the final output. In the language-aware scenario, the exact language identification (such as the name of the language) is known in advance, at this point the speech data 240 may be processed by the first network 210 or by the second network 220, and thus a corresponding text may also be obtained. With these implementations, output from an appropriate network may be selected as the final output of the multilingual processing model 230 based on the prior knowledge about the speech data, and thus the accuracy level of the multilingual processing model 230 may be increased effectively.

[0048] In the language-agnostic scenario, no prior knowledge about the speech data is inputted into the multilingual processing model. At this point, the first and the second networks 210 and 220 may work in parallel and then two texts 242 and 244 may be outputted from the first and the second networks 210 and 220, respectively. Therefore, the challenge of using multiple pipelines is selecting an output from an appropriate network as the final output of the multilingual processing model 230. In implementations of the present disclosure, in order to select the network, a selection strategy is provided for enabling a fully language-agnostic mode. Essentially, a first score may be obtained from the speech data 240 by the first network 210, and a second score may be obtained from the speech data 240 by the second network 220. Then, the network may be selected from the first and second networks 210 and 220 based on a difference between the first and second scores. With these implementations, the problem of selecting an appropriate network is converted into a mathematics problem and thus the network may be selected in an easy and effective way.

[0049] In implementations of the present disclosure, the first score is determined based on a first probability score of at least one first language tag that is identified by the first network from the speech data, and the second score is determined based on a second probability score  of at least one second language tag that is identified by the second network from the speech data. During the inference stage, each network may identify a plurality of language tags (for example, words represented in a corresponding language) from the speech data 240. Then, the log-probability scores may be determined for the identified language tags (for example, a portion of all the identified language tags) . Specifically, first log-probability scores may be obtained from the first network 210, and second log-probability scores may be obtained from the second network 210.

[0050] The difference may be determined by a comparison of the first and second log-probability scores. If the difference is above a predetermined threshold (τ) , it shows that the difference is reliable for selecting the appropriate network, and thus the network corresponding to a greater score in the first and second scores may be selected from the first and second networks. Here, the predetermined threshold τ may be set to a specific value between 0 and 1 according to the working environment. For example, τ may be set to 0.5 or another value. With these implementations, the speech data 240 may be processed in a fast and effective way without waiting for all the language tags are identified from the speech data 240.

[0051] In implementations of the present disclosure, if the difference is below the predetermined threshold (τ) , it shows that the first and second scores are close to each other and the difference that relates to a portion of the identified language tags is not enough for selecting the appropriate network. At this point, an average probability score may be determined for each network. Specifically, a first average probability score is determined for a first plurality of first tags that are identified by the first network from the speech data, and a second average probability score is determined for a second plurality of second tags that are obtained by the second network from the speech data, respectively. In other words, it requires that the whole speech data 240 is processed by both of the first and second networks and all the language tags are identified for determining the average probability score.

[0052] Further, the network may be selected from the first and second networks, and the selected network corresponds to a greater average probability score in the first and second average probability scores. With these implementations, although the decoding speed is relatively slower, it may ensure that the selection is made in a more accurate way and thus the speech data 110 may be converted into the text in a higher accuracy level.

[0053] In implementations of the present disclosure, adjusting the predetermined threshold allows to manage the decoding speed. For instance, setting a smaller threshold enables the decision without calculating average log-probability scores for the remaining tokens using both  decoders. Additionally, a bias score (β) may be added to the average log-probability score of the new decoder, enabling the prioritization of one decoder over the other.

[0054] For a given input test audio, the proposed solution will generate output sequences from each decoder, i.e., the decoders 214 and 224. The output sequence with a higher average log-likelihood score may be selected. Depending on the deployed decoder architecture, the scores might require additional adjustments to match the score range, such as scaling and normalization. Note that the output sequence is not limited to the transcript. In implementations of the present disclosure, the output of the multilingual processing model comprises any of a text, a language identification, an emotion, a domain, or a topic associated with the speech data.

[0055] Based on the above architecture, the multilingual process model 230 may achieve various purposes. Although the above paragraphs describe the multilingual process model 230 by taking the mASR task as an example, the multilingual process model 230 may be trained for identifying a language identification of the speech data. At this point, the language identification (for example, the name of the language) may be outputted as English, French, and the like. Alternatively and / or in addition, the multilingual process model 230 may be trained for identifying an emotion of the speaker, and the like. With these implementations, the multilingual process model 230 may be adjusted for implementing various tasks by expending the language coverage with the supplementary network associated with the new languages.

[0056] The following paragraphs will describe more details about implementations and experiment results of the proposed solution. In some implementations of the present disclosure, the mASR model may employ an encoder-decoder Transformer architecture, and the model parameters may be initialized according to any existing method. Then fine-tune may be implemented based on a predetermined dataset for 500k steps. For example, the final mASR model may include 427M parameters and is structured as follows: The encoder comprises two convolution layers with a filter width of three and a stride of two. Following these convolution layers, there are 24 Transformer blocks with 1, 024 hidden states, 16 attention heads, and a feed-forward dimension of 4, 096 using the activation function. The decoder consists of four Transformer blocks with analogous hidden states, attention heads, and feed-forward dimensions as the encoder.

[0057] The 39 languages covered by the first network may include: Arabic, Bengali, Bulgarian, Burmese, Czech, Dutch, English, Filipino, Finnish, French, German, Greek, Hebrew, Hindi, Hungarian, Indonesian, Italian, Japanese, Kannada, Khmer, Korean, Malay,  Malayalam, Marathi, Nepali, Pashto, Polish, Portuguese, Punjabi, Romanian, Russian, Spanish, Swedish, Tamil, Telugu, Thai, Turkish, Urdu, and Vietnamese.

[0058] Further, 19 languages may be covered by the second network: Asturian, Cebuano, Fula, Ganda, Igbo, Irish, Kabuverdianu, Kamba, Kyrgyz, Luo, Northern Sotho, Nyanja, Oriya, Oromo, Sorani Kurdish, Umbundu, Wolof, Xhosa, and Zulu. These languages represent six different language families, with approximately 10 hours of training data available for each language. Importantly, none of these 19 languages were previously encountered by either the 39-language mASR model used for its initialization. The output vocabulary for these languages is constructed from the unified text using the byte-level BPE algorithm with a size set to 2,000, and no text normalization is applied. During the fine-tuning, the second network is fine tuned for 10,000 steps. The fine-tuned model is tested in three scenarios: 1) group-aware, 2) language aware, and 3) language-agnostic, and results are shown in Table 1 as below.

[0059] Table 1 Word Error Rate results for mASR

[0060] Table 1 shows the Word Error Rate (WER) results for both new and existing languages: WER results for 19 new and 39 existing languages across language-aware, group-aware, and language-agnostic scenarios. In the language-agnostic scenario, the language tag threshold (τ) was set to 0.5, and the bias score (β) was set to 0.1. Results in Table 1 show that although the training dataset for the 19 new languages are relatively small, the WER for all the languages may reach nearly 27%, which is much better than the existing solutions.

[0061] Fig. 5 illustrates an example diagram 500 for effects of multilingual automatic speech recognition according to implementations of the present disclosure. As illustrated in Fig. 5, the horizontal axis indicates the number of the parameters, and the vertical axis indicates the average Character Error Rate (CER) of the experiment result. A curve 510 corresponds to the experiment result of a model including a decoder with the LSTM network of 128 hidden states, a curve 510 corresponds to the experiment result of a model including a decoder with the LSTM network of 512 hidden states, and a curve 530 corresponds to the  experiment result of a model including a decoder and an encoder setup. The curve 530 shows that more parameters lead to lower average CER.

[0062] With the proposed solution, large and powerful ASR models can be extended to new languages without compromising the performance of existing languages, and without the need for corresponding data. For instance, existing ASR models may be extended by using supplementary network (s) for processing new languages without negatively affecting the performance of existing ones poses a significant challenge.

[0063] The proposed solution reduces the time, data requirements, and computational resources necessary to support new languages. For example, in certain scenarios, businesses might urgently require support for a new language to address immediate needs. However, supporting a new language typically involves a time-consuming process, often taking several years. Most of this time is dedicated to data collection, which typically requires hundreds to thousands of hours of labeled data and training processes. With the proposed solution, only a few hours of data are needed to support a new language with acceptable performance.

[0064] The above paragraphs have described details for multilingual speech processing. According to implementations of the present disclosure, a method is provided for multilingual speech processing. Reference will be made to Fig. 6 for more details about the method, where Fig. 6 illustrates an example flowchart of a method 600 for multilingual speech processing according to implementations of the present disclosure. At a block 610, it is determined whether speech data is received. If the speech data is received, the method 600 proceeds to a block 620. At the block 620, a network corresponding to the speech data is identified from a first and a second network that are comprised in a multilingual processing model, the first network being associated with a first group of languages and the second network being associated with a second group of languages. At a block 630, an output that is determined by the identified network based on the speech data is provided as an output of the multilingual processing model.

[0065] In implementations of the present disclosure, identifying the network comprises: obtaining a first score of the speech data by the first network and obtaining a second score of the speech data by the second network, respectively; and selecting the network from the first and second networks based on a difference between the first and second scores.

[0066] In implementations of the present disclosure, the first score is determined based on a first probability score of at least one first language tag that is identified by the first network from the speech data, and the second score is determined based on a second probability score  of at least one second language tag that is identified by the second network from the speech data.

[0067] In implementations of the present disclosure, selecting the network based on the difference between the first and second scores comprises: in response to a determination that the difference is above a predetermined threshold, selecting, from the first and second networks, the network corresponding to a greater score in the first and second scores.

[0068] In implementations of the present disclosure, selecting the network based on the difference between the first and second scores comprises: in response to a determination that the difference is below a predetermined threshold, determining a first average probability score of a first plurality of first tags that are identified by the first network from the speech data, and determining a second average probability score of a second plurality of second tags that are obtained by the second network from the speech data, respectively; and selecting, from the first and second networks, the network corresponding to a greater average probability score in the first and second average probability scores.

[0069] In implementations of the present disclosure, determining the network comprises: obtaining a language classification of the speech data, the language classification indicating a group of languages in the first and second groups to which the speech data belongs; and selecting the network from the first and second networks based on the language classification.

[0070] In implementations of the present disclosure, first parameters of the first network are trained by a first dataset represented in the first group of languages, second parameters of the second network are trained by a second dataset represented in the second group of languages, the first parameters are fixed during a training procedure of the second parameters of the second network, and the second group of languages are new languages excluded from the first group of languages.

[0071] In implementations of the present disclosure, the first network comprises a first encoder, the second network comprises a second encoder, the second encoder shares at least one layer in a plurality of layers that are comprised in the first encoder, and a total number of layers comprised in the first encoder is equal to a total number of layers comprised in the second encoder.

[0072] In implementations of the present disclosure, a mapping component is added between the first and second networks in response to a first data space of the first network being different from a second space of the second network.

[0073] In implementations of the present disclosure, the output of the multilingual processing model comprises any of a text, a language identification, an emotion, a domain, or a topic associated with the speech data.

[0074] According to implementations of the present disclosure, an apparatus is provided for multilingual speech processing. The apparatus comprises: an identifying module, configured for, in response to receiving speech data, identifying, from a first and a second network that are comprised in a multilingual processing model, a network corresponding to the speech data, the first network being associated with a first group of languages and the second network being associated with a second group of languages; and a providing module, configured for providing an output that is determined by the identified network based on the speech data as an output of the multilingual processing model. The apparatus further comprise other modules being configured for implementing other steps in the above method.

[0075] According to implementations of the present disclosure, an electronic device is provided for implementing the method 600. The electronic device comprises: a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed by the computer processor implements a method for multilingual speech processing. The method comprises: in response to receiving speech data, identifying, from a first and a second network that are comprised in a multilingual processing model, a network corresponding to the speech data, the first network being associated with a first group of languages and the second network being associated with a second group of languages; and providing an output that is determined by the identified network based on the speech data as an output of the multilingual processing model.

[0076] In implementations of the present disclosure, identifying the network comprises: obtaining a first score of the speech data by the first network and obtaining a second score of the speech data by the second network, respectively; and selecting the network from the first and second networks based on a difference between the first and second scores.

[0077] In implementations of the present disclosure, the first score is determined based on a first probability score of at least one first language tag that is identified by the first network from the speech data, and the second score is determined based on a second probability score of at least one second language tag that is identified by the second network from the speech data.

[0078] In implementations of the present disclosure, selecting the network based on the difference between the first and second scores comprises: in response to a determination that  the difference is above a predetermined threshold, selecting, from the first and second networks, the network corresponding to a greater score in the first and second scores.

[0079] In implementations of the present disclosure, selecting the network based on the difference between the first and second scores comprises: in response to a determination that the difference is below a predetermined threshold, determining a first average probability score of a first plurality of first tags that are identified by the first network from the speech data, and determining a second average probability score of a second plurality of second tags that are obtained by the second network from the speech data, respectively; and selecting, from the first and second networks, the network corresponding to a greater average probability score in the first and second average probability scores.

[0080] In implementations of the present disclosure, determining the network comprises: obtaining a language classification of the speech data, the language classification indicating a group of languages in the first and second groups to which the speech data belongs; and selecting the network from the first and second networks based on the language classification.

[0081] In implementations of the present disclosure, first parameters of the first network are trained by a first dataset represented in the first group of languages, second parameters of the second network are trained by a second dataset represented in the second group of languages, the first parameters are fixed during a training procedure of the second parameters of the second network, and the second group of languages are new languages excluded from the first group of languages.

[0082] In implementations of the present disclosure, the first network comprises a first encoder, the second network comprises a second encoder, the second encoder shares at least one layer in a plurality of layers that are comprised in the first encoder, and a total number of layers comprised in the first encoder is equal to a total number of layers comprised in the second encoder.

[0083] In implementations of the present disclosure, a mapping component is added between the first and second networks in response to a first data space of the first network being different from a second space of the second network.

[0084] In implementations of the present disclosure, the output of the multilingual processing model comprises any of a text, a language identification, an emotion, a domain, or a topic associated with the speech data.

[0085] According to implementations of the present disclosure, a computer program product, the computer program product comprising a computer readable storage medium  having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform the method 600.

[0086] Fig. 7 illustrates a block diagram of a computing device 700 in which various implementations of the present disclosure can be implemented. It would be appreciated that the computing device 700 shown in Fig. 7 is merely for purpose of illustration, without suggesting any limitation to the functions and scopes of the present disclosure in any manner. The computing device 700 may be used to implement the above method in implementations of the present disclosure. As shown in Fig. 7, the computing device 700 may be a general-purpose computing device. The computing device 700 may at least comprise one or more processors or processing units 710, a memory 720, a storage unit 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760.

[0087] The processing unit 710 may be a physical or virtual processor and can implement various processes based on programs stored in the memory 720. In a multi-processor system, multiple processing units execute computer executable instructions in parallel so as to improve the parallel processing capability of the computing device 700. The processing unit 710 may also be referred to as a central processing unit (CPU) , a microprocessor, a controller, or a microcontroller.

[0088] The computing device 700 typically includes various computer storage medium. Such medium can be any medium accessible by the computing device 700, including, but not limited to, volatile and non-volatile medium, or detachable and non-detachable medium. The memory 720 can be a volatile memory (for example, a register, cache, Random Access Memory (RAM) ) , a non-volatile memory (such as a Read-Only Memory (ROM) , Electrically Erasable Programmable Read-Only Memory (EEPROM) , or a flash memory) , or any combination thereof. The storage unit 730 may be any detachable or non-detachable medium and may include a machine-readable medium such as a memory, flash memory drive, magnetic disk, or another other media, which can be used for storing information and / or data and can be accessed in the computing device 700.

[0089] The computing device 700 may further include additional detachable / non-detachable, volatile / non-volatile memory medium. Although not shown in Fig. 7, it is possible to provide a magnetic disk drive for reading from and / or writing into a detachable and non-volatile magnetic disk and an optical disk drive for reading from and / or writing into a detachable non-volatile optical disk. In such cases, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0090] The communication unit 740 communicates with a further computing device via the communication medium. In addition, the functions of the components in the computing device 700 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 700 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs) or further general network nodes.

[0091] The input device 750 may be one or more of a variety of input devices, such as a mouse, keyboard, tracking ball, voice-input device, and the like. The output device 760 may be one or more of a variety of output devices, such as a display, loudspeaker, printer, and the like. By means of the communication unit 740, the computing device 700 can further communicate with one or more external devices (not shown) such as the storage devices and display device, with one or more devices enabling the user to interact with the computing device 700, or any devices (such as a network card, a modem, and the like) enabling the computing device 700 to communicate with one or more other computing devices, if required. Such communication can be performed via input / output (I / O) interfaces (not shown) .

[0092] In some implementations, instead of being integrated in a single device, some, or all components of the computing device 1000 may also be arranged in cloud computing architecture. In the cloud computing architecture, the components may be provided remotely and work together to implement the functionalities described in the present disclosure. In some implementations, cloud computing provides computing, software, data access and storage service, which will not require end users to be aware of the physical locations or configurations of the systems or hardware providing these services. In various implementations, the cloud computing provides the services via a wide area network (such as Internet) using suitable protocols. For example, a cloud computing provider provides applications over the wide area network, which can be accessed through a web browser or any other computing components. The software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote position. The computing resources in the cloud computing environment may be merged or distributed at locations in a remote data center. Cloud computing infrastructures may provide the services through a shared data center, though they behave as a single access point for the users. Therefore, the cloud computing architectures may be used to provide the components and functionalities described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.

[0093] The functionalities described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-Programmable Gate Arrays (FPGAs) , Application-specific Integrated Circuits (ASICs) , Application-specific Standard Products (ASSPs) , System-on-a-chip systems (SOCs) , Complex Programmable Logic Devices (CPLDs) , and the like.

[0094] Program code for carrying out the methods of the subject matter described herein may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose computer, special purpose computer, or other programmable data processing apparatus such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely or partly on a machine, executed as a stand-alone software package partly on the machine, partly on a remote machine, or entirely on the remote machine or server.

[0095] In the context of this disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0096] Further, while operations are illustrated in a particular order, this should not be understood as requiring that such operations are performed in the particular order shown or in sequential order, or that all illustrated operations are performed to achieve the desired results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the subject matter described herein, but rather as descriptions of features that may be specific to particular implementations. Certain features that are described in the context of separate implementations may also be  implemented in combination in a single implementation. Rather, various features described in a single implementation may also be implemented in multiple implementations separately or in any suitable sub-combination.

[0097] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter specified in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0098] From the foregoing, it will be appreciated that specific implementations of the presently disclosed technology have been described herein for purposes of illustration, but that various modifications may be made without deviating from the scope of the disclosure. Accordingly, the presently disclosed technology is not limited except as by the appended claims.

[0099] Implementations of the subject matter and the functional operations described in the present disclosure can be implemented in various systems, digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term “data processing unit” or “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0100] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing  environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document) , in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code) . A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.

[0101] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0102] It is intended that the specification, together with the drawings, be considered exemplary only, where exemplary means an example. As used herein, the use of “or” is intended to include “and / or” , unless the context clearly indicates otherwise.

[0103] While the present disclosure contains many specifics, these should not be construed as limitations on the scope of any disclosure or of what may be claimed, but rather as descriptions of features that may be specific to particular implementations of particular disclosures. Certain features that are described in the present disclosure in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0104] Similarly, while operations are illustrated in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the implementations described in the present disclosure should not be understood as requiring such separation in all implementations. Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in the present disclosure.

Claims

1.A method for multilingual speech processing, comprising:in response to receiving speech data, identifying, from a first and a second network that are comprised in a multilingual processing model, a network corresponding to the speech data, the first network being associated with a first group of languages and the second network being associated with a second group of languages; andproviding an output that is determined by the identified network based on the speech data as an output of the multilingual processing model.2.The method of claim 1, wherein identifying the network comprises:obtaining a first score of the speech data by the first network and obtaining a second score of the speech data by the second network, respectively; andselecting the network from the first and second networks based on a difference between the first and second scores.3.The method of claim 2, wherein the first score is determined based on a first probability score of at least one first language tag that is identified by the first network from the speech data, and the second score is determined based on a second probability score of at least one second language tag that is identified by the second network from the speech data.4.The method of claim 3, wherein selecting the network based on the difference between the first and second scores comprises: in response to a determination that the difference is above a predetermined threshold,selecting, from the first and second networks, the network corresponding to a greater score in the first and second scores.5.The method of claim 3, wherein selecting the network based on the difference between the first and second scores comprises: in response to a determination that the difference is below a predetermined threshold,determining a first average probability score of a first plurality of first tags that are identified by the first network from the speech data, and determining a second average probability score of a second plurality of second tags that are obtained by the second network from the speech data, respectively; andselecting, from the first and second networks, the network corresponding to a greater average probability score in the first and second average probability scores.6.The method of claim 1, wherein determining the network comprises:obtaining a language classification of the speech data, the language classification indicating a group of languages in the first and second groups to which the speech data belongs; andselecting the network from the first and second networks based on the language classification.7.The method of claim 1, wherein first parameters of the first network are trained by a first dataset represented in the first group of languages, second parameters of the second network are trained by a second dataset represented in the second group of languages, the first parameters are fixed during a training procedure of the second parameters of the second network, and the second group of languages are new languages excluded from the first group of languages.8.The method of claim 1, wherein the first network comprises a first encoder, the second network comprises a second encoder, the second encoder shares at least one layer in a plurality of layers that are comprised in the first encoder, and a total number of layers comprised in the first encoder is equal to a total number of layers comprised in the second encoder.9.The method of claim 1, wherein a mapping component is added between the first and second networks in response to a first data space of the first network being different from a second space of the second network.10.The method of claim 1, wherein the output of the multilingual processing model comprises any of a text, a language identification, an emotion, a domain, or a topic associated with the speech data.11.An electronic device, comprising a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that when executed  by the computer processor implements a method for multilingual speech processing, the method comprising:in response to receiving speech data, identifying, from a first and a second network that are comprised in a multilingual processing model, a network corresponding to the speech data, the first network being associated with a first group of languages and the second network being associated with a second group of languages; andproviding an output that is determined by the identified network based on the speech data as an output of the multilingual processing model.12.The device of claim 11, wherein identifying the network comprises:obtaining a first score of the speech data by the first network and obtaining a second score of the speech data by the second network, respectively; andselecting the network from the first and second networks based on a difference between the first and second scores.13.The device of claim 12, wherein the first score is determined based on a first probability score of at least one first language tag that is identified by the first network from the speech data, and the second score is determined based on a second probability score of at least one second language tag that is identified by the second network from the speech data.14.The device of claim 13, wherein selecting the network based on the difference between the first and second scores comprises: in response to a determination that the difference is above a predetermined threshold,selecting, from the first and second networks, the network corresponding to a greater score in the first and second scores.15.The device of claim 13, wherein selecting the network based on the difference between the first and second scores comprises: in response to a determination that the difference is below a predetermined threshold,determining a first average probability score of a first plurality of first tags that are identified by the first network from the speech data, and determining a second average probability score of a second plurality of second tags that are obtained by the second network from the speech data, respectively; andselecting, from the first and second networks, the network corresponding to a greater average probability score in the first and second average probability scores.16.The device of claim 11, wherein determining the network comprises:obtaining a language classification of the speech data, the language classification indicating a group of languages in the first and second groups to which the speech data belongs; andselecting the network from the first and second networks based on the language classification.17.The device of claim 11, wherein first parameters of the first network are trained by a first dataset represented in the first group of languages, second parameters of the second network are trained by a second dataset represented in the second group of languages, the first parameters are fixed during a training procedure of the second parameters of the second network, and the second group of languages are new languages excluded from the first group of languages.18.The device of claim 11, wherein the first network comprises a first encoder, the second network comprises a second encoder, the second encoder shares at least one layer in a plurality of layers that are comprised in the first encoder, and a total number of layers comprised in the first encoder is equal to a total number of layers comprised in the second encoder.19.The device of claim 11, wherein a mapping component is added between the first and second networks in response to a first data space of the first network being different from a second space of the second network, and the output of the multilingual processing model comprises any of a text, a language identification, an emotion, a domain, or a topic associated with the speech data.20.A non-transitory computer program product, the non-transitory computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by an electronic device to cause the electronic device to perform a method for multilingual speech processing, the method comprising:in response to receiving speech data, identifying, from a first and a second network that are comprised in a multilingual processing model, a network corresponding to the speech data, the first network being associated with a first group of languages and the second network being associated with a second group of languages; andproviding an output that is determined by the identified network based on the speech data as an output of the multilingual processing model.

Citation Information

Patent Citations

  • Training method for language recognition model and language recognition method

    CN105280181A

  • Method and device for multilingual hybrid model establishment and data acquisition, and electronic equipment

    CN108711420A

  • Text processing method based on multilingual branch model and related device

    CN113705240A

  • Speech recognition method and server

    CN115132176A

  • Language model adaptation using result selection

    US20140365218A1