Multi-language speech processing

By adding a supplementary network associated with the new language to the multilingual speech processing model and using the new language data to train the second network, the problems of performance degradation and high computational cost when the existing model is expanded to a new language are solved, and efficient and accurate multilingual processing is achieved.

CN120660135APending Publication Date: 2025-09-16DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004134.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-04
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing multilingual speech processing models can only process a limited number of existing languages, suffer from performance degradation and high computational cost when expanding to new languages, and require prior knowledge of the input audio.

Method used

By adding a supplementary network associated with the new language to the existing multilingual processing model, the second network is trained using the new language data, the existing network parameters are kept unchanged, and some encoder layers are shared to achieve processing of the new language.

Benefits of technology

Effectively expand the model to cover new languages, reduce computing resources and time requirements, avoid performance degradation of existing languages, eliminate the need for prior audio knowledge, and improve the accuracy and efficiency of multi-language processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120660135A_ABST
    Figure CN120660135A_ABST
Patent Text Reader

Abstract

The present disclosure presents a method, an apparatus, and a computer program product for multilingual speech processing. In the method, in response to receiving voice data, a network corresponding to the voice data is identified from a first network and a second network included in a multi-language processing model, the first network being associated with a first set of languages, and the second network being associated with a second set of languages. An output determined by the identified network based on the speech data is provided as an output of the multilingual processing model. With these implementations, in addition to a first set of languages covered by a first network, language coverage of a multi-language processing model may be extended by a second set of languages covered by a second network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to multilingual speech processing and, more particularly, to methods, apparatus, and computer program products for language coverage extension of multilingual tasks using supplemental networks associated with new languages. Background Art

[0002] Machine learning technology is widely used in multilingual speech processing, and various machine learning models have been developed for various tasks, such as multilingual automatic speech recognition (mASR), language identification, and emotion detection based on speech data. However, there are thousands of languages, and existing multilingual processing models are only trained on a small subset of existing languages. As a result, these models can only process speech data represented in existing languages ​​and become useless when processing new languages ​​in addition to existing languages. At this time, it is desirable to expand the coverage of models to include new languages. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for multilingual speech processing is provided. In this method, in response to receiving speech data, a network corresponding to the speech data is identified from a first network and a second network included in a multilingual processing model, the first network being associated with a first set of languages, and the second network being associated with a second set of languages. An output determined by the identified network based on the speech data is provided as an output of the multilingual processing model.

[0004] In a second aspect of the present disclosure, an electronic device is provided, comprising: a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions that, when executed by the computer processor, implement the method according to the first aspect of the present disclosure.

[0005] In a third aspect of the present disclosure, a computer program product is provided, comprising a computer-readable storage medium having program instructions embodied therewith, the program instructions being executable by an electronic device to cause the electronic device to perform the method according to the first aspect of the present disclosure.

[0006] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The foregoing and other objects, features, and advantages of the present disclosure will become more apparent through a more particular description of some implementations of the present disclosure in the accompanying drawings, wherein like reference numerals generally refer to like parts throughout the implementations of the present disclosure.

[0008] Figure 1 An example environment for multilingual speech processing based on machine learning techniques is shown;

[0009] Figure 2 An example diagram illustrating multilingual speech processing by using a supplementary network associated with a new language according to implementations of the present disclosure is shown;

[0010] Figure 3 An example diagram illustrating a multi-language processing model according to an implementation of the present disclosure is shown;

[0011] Figure 4 An example diagram illustrating supplementary components in a multilingual processing model according to an implementation of the present disclosure;

[0012] Figure 5 An example diagram showing results for multilingual automatic speech recognition according to an implementation of the present disclosure is shown;

[0013] Figure 6 An example flow chart illustrating a method for multilingual speech processing according to an implementation of the present disclosure; and

[0014] Figure 7 A block diagram is shown of a computing device in which various implementations of the present disclosure may be implemented. DETAILED DESCRIPTION

[0015] The principles of the present disclosure will now be described with reference to some implementations. It should be understood that these implementations are intended for illustrative purposes only and are intended to assist those skilled in the art in understanding and implementing the present disclosure without implying any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in a variety of ways different from the methods described below.

[0016] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0017] References in this disclosure to "one implementation," "implementation," "example implementation," etc., indicate that the described implementation may include a particular feature, structure, or characteristic, but not every implementation necessarily includes the particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same implementation. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example implementation, it is claimed that it is within the knowledge of those skilled in the art to effect such feature, structure, or characteristic in conjunction with other implementations, whether or not explicitly described.

[0018] It should be understood that although the terms "first" and "second" etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example implementations. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0019] The terms used herein are for the purpose of describing specific implementations only and are not intended to limit the example implementations. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprises," "including," "having," "including," and / or "comprising" when used herein specify the presence of the features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0020] The principles of the present disclosure will now be described with reference to some implementations. It should be understood that these implementations are for illustrative purposes only and are intended to assist those skilled in the art in understanding and implementing the present disclosure without implying any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways different from those described below. In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present disclosure belongs.

[0021] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws and regulations and relevant rules.

[0022] It is understandable that before using the technical solutions disclosed in the embodiments of the present invention, the user should be notified of the type, scope of use, and usage scenarios of the personal information involved in this application in an appropriate manner in accordance with relevant laws and regulations, and the user's authorization should be obtained.

[0023] For example, in response to receiving an activity request from a user, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. Therefore, the user can independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the technical solution of this application based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to receiving a user's active request, the method for sending a prompt to the user may include, for example, a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "Agree" or "Disagree" to provide personal information to the electronic device.

[0025] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation of the present disclosure. Other methods that meet relevant laws and regulations are also applicable to the implementation of the present disclosure.

[0026] Currently, machine learning technology is widely used in multilingual speech processing, and various machine learning models have been developed for various tasks. Figure 1 ,here, Figure 1 An example environment 100 for multilingual speech processing based on machine learning techniques is shown. Figure 1 As shown, the multi-language processing model 130 can be trained to process speech data 110 and then provide output 120 based on the input speech data 110. Here, the speech data 110 can be speech input including speech in various natural languages ​​(such as English, French, German, etc.). Here, the speech data 110 can be represented in various formats, such as an audio file, an audio stream, an audio component in a video file / stream, etc.

[0027] Here, the multilingual processing model 130 may include an encoder 132 and a decoder 134. The encoder 132 may convert the speech data 110 into an internal representation (such as an embedding), and then the decoder 134 may convert the internal representation into an output 120. Based on the task of the multilingual processing model 130, the output 120 may include various types, such as text (i.e., recognized transcript), language identification, emotion, domain, topic of the speech data 110, etc.

[0028] For descriptive purposes, the mASR task can be used as an example of a task implemented by multilingual processing model 130. In the mASR task, textual transcriptions can be identified from speech data 110. Although existing mASR models have been scaled to cover over a thousand languages, extending them to new languages ​​under the following constraints remains an unresolved challenge. For example, training data for existing languages ​​is limited or even unavailable, prior knowledge of the input audio is absent, the model needs to be parameter-efficient, and the model's performance should not degrade on existing languages.

[0029] Various solutions have been developed for extending these models to new languages, e.g., attachable language-specific adapters that require prior knowledge of the input audio (i.e., language ID) to deploy the corresponding adapter. On the other hand, lifelong learning approaches, in which new language data is combined with existing data to continue model training, assume the availability of training data for existing languages. Furthermore, these techniques are considered parameter-inefficient because the entire model needs to be updated. In particular, for models with over a billion parameters, updating each time the model is requested to support a new language results in a huge computational cost. Other techniques based on cross-lingual transfer learning, in which an existing mASR model is fine-tuned (completely or partially) using new language data, typically result in performance degradation for the existing language due to catastrophic forgetting. At this point, it is desirable to extend the model to new languages ​​in an easier and efficient way.

[0030] In view of the above, the present disclosure proposes a multilingual speech processing solution using a new network trained on a new language. Generally, with respect to an existing model implemented by a machine learning network (trained on a dataset represented in (multiple) existing languages ​​(i.e., a first set of languages)), an additional network can be added to process speech data represented in (multiple) new languages ​​(i.e., a second set of languages) different from the existing languages. Here, the encoder and decoder included in the additional network can be trained on a dataset represented in the new language, and then the language of the model can be expanded to include both existing and new languages.

[0031] For more details about the proposed multilingual speech processing scheme, refer to Figure 2 . Figure 2 An example diagram 200 of multilingual speech processing using a supplementary network associated with a new language(s) is shown in accordance with an implementation of the present disclosure. Figure 2 As shown, the present disclosure proposes a multilingual processing model 230, which may include a first network 210 and a second network 220. Specifically, the first network 210 may include an encoder 212 and a decoder 214, both of which are trained using a first dataset represented in a first set of languages. Here, the first network 210 may be implemented by an existing mASR model and may function as a pipeline for converting speech data 240 in an existing language into text 242.

[0032] Furthermore, a second network 220 is added to the multilingual processing model 230 as an additional pipeline for processing speech data 240. Specifically, the second network 220 can convert speech data 240 in the new language into text 244. At this point, the multilingual processing model 230 can process speech data in both existing languages ​​and the new language. Once the multilingual processing model 230 is trained, speech data 240 can be received during the inference phase, and networks can be identified from the first network 210 and the second network 220. The output determined by the identified network based on the speech data can then be provided as the output of the multilingual processing model 230. In other words, one of the text 242 and the text 244 can be selected as the final output of the multilingual processing model 230.

[0033] Utilizing these implementations of the present disclosure, the proposed solution does not require the availability of training data for the supported languages, which may be completely or partially unavailable due to various reasons, such as expired data licenses, modifications made by the data owner (e.g., deleted video / audio, changed access rights to private content, etc.), data loss, etc. Compared to some existing solutions that require prior knowledge of the input audio, the proposed solution does not require any prior knowledge of the input audio, and the model can automatically convert the speech data into appropriate output. In addition, the proposed solution can minimize the impact on the performance of existing languages ​​by keeping the corresponding parameters unchanged, and thus the proposed solution can alleviate the problem of catastrophic forgetting in existing languages ​​and circumvent modeling capacity limitations through additional parameters dedicated to the new language. At the same time, only the second network is updated by the dataset representing the new language, and the parameters in the first network are fixed. Therefore, the size of the parameters to be trained is limited to be smaller than mASR. As a result, it requires less computation time and resources.

[0034] Having provided a brief description of the proposed solution, the following paragraphs will provide more details about the second network 220. Here, the second network 220 includes an encoder 222 and a decoder 224 for a new language. In order to seamlessly extend the mASR model (i.e., the first network 210) to include the new language, the second network 220 adds an additional pipeline for processing speech data 240, where a portion of the encoder 212 is shared between multiple encoders 212 and 222.

[0035] In implementations of the present disclosure, one or more supplementary networks can be added to the multi-language processing model 230. For simplicity, only one supplementary network is described as an example encoder-decoder pipeline for a new language. However, multiple supplementary encoder-decoder pipelines can be added, and these pipelines can be run in parallel or attached / detached depending on the specific working environment.

[0036] In implementations of the present disclosure, the first parameters of the first network are trained using a first dataset represented by a first set of languages, and the second parameters of the second network are trained using a second dataset represented by a second set of languages. For example, an existing mASR model can be selected as the first network, in which the first parameters have already been trained. Unlike existing solutions, the second network 220 is trained only using new training data represented by a new language, while the parameters of the first network 210 are fixed and are not affected by the new training data during the training process of the second parameters of the second network. The second set of languages ​​can be new languages ​​that are excluded from the first set of languages. With these implementations, the parameters of the first network 210 are not updated by training data represented by the new language, and therefore the performance of the first network 210 for processing existing languages ​​can remain unchanged and will not degrade due to catastrophic forgetting.

[0037] In an implementation of the present disclosure, the first network includes a first encoder (such as encoder 212), and the second network includes a second encoder (such as encoder 222). For more details about the encoders, see Figure 3 ,here, Figure 3 An example diagram 300 of an implementation for a multilingual processing model according to the present disclosure is shown. Here, the first network 210 and the second network 220 can be implemented using any pre-trained mASR architecture (such as attention-based encoder-decoder (AED), connection temporal classification (CTC), RNN-T, etc.). For the decoder architecture, the new decoder 224 can be modeled using any network architecture, such as long short-term memory (LSTM), transformer, etc. The size of the decoder can vary depending on the number of layers / blocks, hidden dimensions, and other hyperparameters. In addition, the decoder component can be omitted as in the CTC architecture. The new decoder can be randomly initialized or transferred from other pre-trained models. In addition to the recognized text, the output format of the decoder can also be designed to predict other information related to the input test audio, such as language ID, speaker emotion, etc.

[0038] Similar to the new decoder, the new encoder component can be modeled using any network architecture with different sizes. It can also be initialized randomly or from any other pre-trained model. The input features for the new encoder are derived from the output or other internal layers / blocks of the original fixed mASR encoder. The input features can be pre-processed by other modules as needed. Figure 3 In the example, the encoder 212 may include n layers: layer 1 (indicated by block 310), ..., layer nk (indicated by block 312), ..., and layer n (indicated by block 314). Here, n and k represent integers, 1 <k<n。

[0039] In an implementation of the present disclosure, the second encoder shares at least one layer among the plurality of layers included in the first encoder, and the total number of layers included in the first encoder is equal to the total number of layers included in the second encoder. Figure 3 As shown, the output of layer nk in encoder 212 is connected to the input of layer 1 in encoder 222, and thus encoder 222 also includes n layers: layer 1 (indicated by block 310), ..., layer nk (indicated by block 312), ..., layer 1 (indicated by block 320), ..., and layer k (indicated by block 322). Specifically, encoders 212 and 222 share Figure 3 The following shaded layers in : layer 1 (indicated by box 310), ..., layer nk (indicated by box 312).

[0040] Although the new language is different from existing languages, natural languages ​​often share common features to some extent. Using these implementations, the second network 220 can reuse knowledge acquired from existing languages ​​to process the new language. Therefore, even if training data in the new language is limited, the accuracy level of the second network 220 can be increased.

[0041] The Transformer layer within encoder 222 can be initialized from a fixed encoder. For example, to initialize a four-block encoder component, the last four blocks of the fixed encoder can be used. The new encoder is integrated in a manner that maintains the same overall encoder depth at the same level (i.e., 24 blocks). In the above example, the four-block encoder component is connected to the output of the 20th block of the fixed encoder.

[0042] In implementations of the present disclosure, in response to the first data space of the first network being different from the second data space of the second network, a mapping component may be added between the first network and the second network. Utilizing these implementations, the present disclosure allows the first network and the second network to have different dimensions in their corresponding data spaces, thereby enabling a more flexible selection of network architectures when designing the multilingual processing model 230.

[0043] refer to Figure 4 , Figure 4 An example diagram 400 of a supplementary component for use in a multilingual processing model according to an implementation of the present disclosure is shown. Figure 4As shown, a component 410 can be added between layer nk in the first network 210 and layer 1 in the second network 220. The component 410 can be adapted in size from the encoder 212 to the encoder 222. Here, the addition of component 410 is optional; in other words, the new decoder can be designed to consume input features directly from the original encoder or other preprocessing modules. In an implementation of the present disclosure, the output features from the encoder 222 can be further modified (such as by component 420) before being sent to the decoder 224. For example, the features can be passed through a normalization module (e.g., layer by layer or batch by batch), a residual connection module, an attention module, etc., for fitting to the downstream tasks implemented by the decoder 224.

[0044] The main limitation of the multi-head decoder approach is that the fixed encoder trained on the existing language lacks exposure to the new language, resulting in suboptimal performance. To address this issue, an additional encoder component is incorporated for the new language, enabling the learning of relevant acoustic representations. During training, the parameters of the first network 210 remain fixed to maintain the performance of the existing language. The parameters of the second network 220 are updated using the training data of the new language. In general, the training process follows the standard steps used in deep learning, i.e., forward-backward propagation with parameter updates after the learning schedule. In cases where training data for the existing language is only partially available, the parameters of the mASR can also be updated completely or partially (e.g., only the encoder).

[0045] Once the multilingual processing model 230 is trained, it can output text corresponding to the input speech data 240 during the inference phase. Experiments can involve three scenarios: 1) group-aware, 2) language-aware, and 3) language-agnostic. In the group-aware and language-aware scenarios, the language classification of the input speech data 240 is known in advance. Here, the language classification indicates the language group of the first and second groups to which the speech data belongs. A network can then be selected from the first and second networks based directly on the language classification.

[0046] Specifically, in a group-aware scenario, it is known in advance whether the language of the speech data 110 is new or existing. In this case, a network corresponding to the language classification can be selected to convert the speech data into corresponding text. For example, if the speech data is represented in an existing language, the text 242 output from the first network 210 is used as the final output; otherwise, if the speech data is not represented in any existing language, the text 244 output from the second network 220 is used as the final output. In a language-aware scenario, the exact language identifier (such as the name of the language) is known in advance. In this case, the speech data 240 can be processed by the first network 210 or the second network 220, and the corresponding text can also be obtained. Using these implementations, the output from the appropriate network can be selected as the final output of the multilingual processing model 230 based on the prior knowledge of the speech data, thereby effectively increasing the accuracy level of the multilingual processing model 230.

[0047] In a language-agnostic scenario, no prior knowledge about the speech data is input into the multilingual processing model. In this case, the first network 210 and the second network 220 can operate in parallel, and two texts 242 and 244 can be output from the first network 210 and the second network 220, respectively. Therefore, the challenge of using multiple pipelines is selecting the output from the appropriate network as the final output of the multilingual processing model 230. In implementations of the present disclosure, a selection strategy is provided for enabling a completely language-agnostic mode for network selection. Essentially, a first score can be obtained from the speech data 240 by the first network 210, and a second score can be obtained from the speech data 240 by the second network 220. A network can then be selected from the first network 210 and the second network 220 based on the difference between the first score and the second score. Using these implementations, the problem of selecting the appropriate network is converted into a mathematical problem, making it easy and efficient to select a network.

[0048] In an implementation of the present disclosure, a first score is determined based on a first probability score of at least one first language token identified from the speech data by the first network, and a second score is determined based on a second probability score of at least one second language token identified from the speech data by the second network. During the inference phase, each network can identify multiple language tokens (e.g., words in a corresponding language) from the speech data 240. Log probability scores can then be determined for the identified language tokens (e.g., a portion of all identified language tokens). Specifically, a first log probability score can be obtained from the first network 210, and a second log probability score can be obtained from the second network 210.

[0049] The difference can be determined by comparing the first log probability score and the second log probability score. If the difference is higher than a predetermined threshold (τ), it is shown that the difference is reliable for selecting an appropriate network, and thus the network corresponding to the larger score of the first score and the second score can be selected from the first network and the second network. Here, the predetermined threshold τ can be set to a specific value between 0 and 1 according to the working environment. For example, τ can be set to 0.5 or another value. With these implementations, the speech data 240 can be processed in a fast and efficient manner without waiting for all language tags to be identified from the speech data 240.

[0050] In an implementation of the present disclosure, if the difference is below a predetermined threshold (τ), it indicates that the first score and the second score are close to each other, and the difference associated with a portion of the identified language tokens is insufficient to select an appropriate network. In this case, an average probability score can be determined for each network. Specifically, a first average probability score is determined for a first plurality of first tokens identified from the speech data by the first network, and a second average probability score is determined for a second plurality of second tokens obtained from the speech data by the second network. In other words, it requires that the entire speech data 240 be processed by both the first network and the second network, and all language tokens are identified for determining the average probability score.

[0051] In addition, a network can be selected from the first network and the second network, and the selected network corresponds to the larger average probability score of the first average probability score and the second average probability score. With these implementations, although the decoding speed is relatively slow, it can ensure that the selection is made in a more accurate manner, and thus the speech data 110 can be converted into text with a higher level of accuracy.

[0052] In implementations of the present disclosure, adjusting the predetermined threshold allows for managing decoding speed. For example, setting a smaller threshold enables decision making without using both decoders to calculate the average log probability score of the remaining tokens. Additionally, a bias score (β) can be added to the average log probability score of the new decoder to prioritize one decoder over another.

[0053] For a given input test audio, the proposed solution will generate an output sequence from each decoder (i.e., decoders 214 and 224). The output sequence with the higher average log-likelihood score can be selected. Depending on the decoder architecture deployed, the scores may require additional adjustments to match the score range, such as scaling and normalization. Note that the output sequence is not limited to transcribed text. In the implementation of the present disclosure, the output of the multilingual processing model includes any of the text, language identifier, emotion, domain, or theme associated with the speech data.

[0054] Based on the above architecture, multi-language processing model 230 can achieve various goals. Although the above paragraphs describe multi-language processing model 230 using the mASR task as an example, multi-language processing model 230 can also be trained to identify the language identifier of speech data. In this case, the language identifier (e.g., the name of the language) can be output as English, French, etc. Alternatively and / or additionally, multi-language processing model 230 can be trained to identify the speaker's emotion, etc. Using these implementations, multi-language processing model 230 can be adapted to achieve various tasks by expanding language coverage with supplementary networks associated with new languages.

[0055] The following paragraphs will describe more details about the implementation and experimental results of the proposed solution. In some implementations of the present disclosure, the mASR model can adopt an encoder-decoder converter architecture, and the model parameters can be initialized according to any existing method. Fine-tuning can then be achieved based on a predetermined dataset for 500k steps. For example, the final mASR model may include 427M parameters and is constructed as follows: the encoder includes two convolutional layers with filters of width 3 and stride 2. After these convolutional layers, there are 24 converter blocks with 1,024 hidden states, 16 attention heads, and a feedforward dimension of 4,096 using an activation function. The decoder consists of four converter blocks with similar hidden states, attention heads, and feedforward dimensions as the encoder.

[0056] The 39 languages ​​covered by the First Network may include: Arabic, Bengali, Bulgarian, Burmese, Czech, Dutch, English, Filipino, Finnish, French, German, Greek, Hebrew, Hindi, Hungarian, Indonesian, Italian, Japanese, Kannada, Khmer, Korean, Malay, Malayalam, Marathi, Nepali, Pashto, Polish, Portuguese, Punjabi, Romanian, Russian, Spanish, Swedish, Tamil, Telugu, Thai, Turkish, Urdu and Vietnamese.

[0057] Furthermore, the second network can cover 19 languages: Asturian, Cebuano, Fula, Ganda, Igbo, Irish, Kabwidianu, Kamba, Kyrgyz, Luo, Northern Sotho, Nyanja, Oriya, Oromo, Sorani Kurdish, Umbundu, Wolof, Xhosa, and Zulu. These languages ​​represent six different language families, with approximately 10 hours of training data available for each. Importantly, none of these 19 languages ​​were previously initialized by the 39-language mASR model. The output vocabulary for these languages ​​was constructed from uniform text using a byte-level BPE algorithm with a size set to 2,000, and no text normalization was applied. During fine-tuning, the second network was fine-tuned for 10,000 steps. The fine-tuned model was tested in three scenarios: 1) group-aware, 2) language-aware, and 3) language-agnostic, and the results are shown in Table 1 below. Table 1 Word error rate results for mASR

[0058] Table 1 shows the word error rate (WER) results for both new and existing languages: WER results for 19 new languages ​​and 39 existing languages ​​across language-aware, group-aware, and language-agnostic scenarios. In the language-agnostic scenario, the language tag threshold (τ) is set to 0.5 and the bias score (β) is set to 0.1. The results in Table 1 show that although the training datasets for the 19 new languages ​​are relatively small, the WER for all languages ​​can reach close to 27%, which is much better than existing solutions.

[0059] Figure 5 An example diagram 500 is shown showing the effect of multilingual automatic speech recognition according to the implementation of the present disclosure. Figure 5 As shown, the horizontal axis represents the number of parameters, and the vertical axis represents the average character error rate (CER) of the experimental results. Curve 510 corresponds to the experimental results of the model including a decoder with an LSTM network having 128 hidden states, curve 510 corresponds to the experimental results of the model including a decoder with an LSTM network having 512 hidden states, and curve 530 corresponds to the experimental results of the model including the decoder and encoder settings. Curve 530 shows that more parameters lead to a lower average CER.

[0060] With the proposed solution, large and powerful ASR models can be extended to new languages ​​without compromising the performance of existing languages ​​and without requiring corresponding data. For example, an existing ASR model can be extended by using (multiple) supplementary networks for processing new languages ​​without negatively impacting the performance of the existing ASR model, which poses a significant challenge.

[0061] The proposed solution reduces the time, data requirements, and computing resources required to support new languages. For example, in some scenarios, an enterprise may urgently need to support a new language to address an immediate need. However, supporting a new language is typically a time-consuming process, often taking several years. Much of this time is dedicated to data collection and training, which typically requires hundreds to thousands of hours of labeled data. With the proposed solution, only a few hours of data are needed to support a new language with acceptable performance.

[0062] The above paragraphs have described the details for multilingual speech processing. According to the implementation of the present disclosure, a method for multilingual speech processing is provided. For more details about the method, please refer to Figure 6 ,in, Figure 6 An example flow chart of a method 600 for multilingual speech processing according to an implementation of the present disclosure is shown. At block 610, a determination is made as to whether speech data has been received. If speech data has been received, method 600 proceeds to block 620. At block 620, a network corresponding to the speech data is identified from a first network and a second network included in a multilingual processing model, the first network being associated with a first set of languages, and the second network being associated with a second set of languages. At block 630, an output determined by the identified network based on the speech data is provided as an output of the multilingual processing model.

[0063] In an implementation of the present disclosure, identifying a network includes: obtaining first scores of speech data respectively through a first network, and obtaining second scores of speech data respectively through a second network; and selecting a network from the first network and the second network based on a difference between the first scores and the second scores.

[0064] In an implementation of the present disclosure, the first score is determined based on a first probability score of at least one first language token recognized from the speech data by the first network, and the second score is determined based on a second probability score of at least one second language token recognized from the speech data by the second network.

[0065] In an implementation of the present disclosure, selecting a network based on a difference between the first score and the second score includes: in response to determining that the difference is above a predetermined threshold, selecting a network corresponding to a larger score from the first network and the second network.

[0066] In an implementation of the present disclosure, selecting a network based on a difference between a first score and a second score includes: in response to determining that the difference is below a predetermined threshold, determining a first average probability score for a first plurality of first labels recognized from the speech data by the first network, and determining a second average probability score for a second plurality of second labels obtained from the speech data by the second network; and selecting a network from the first network and the second network that has a larger average probability score than the first average probability score and the second average probability score.

[0067] In an implementation of the present disclosure, determining a network includes: obtaining a language classification of the speech data, the language classification indicating a language group of the first and second groups to which the speech data belongs; and selecting a network from the first and second networks based on the language classification.

[0068] In an implementation of the present disclosure, first parameters of a first network are trained using a first data set represented by a first set of languages, second parameters of a second network are trained using a second data set represented by a second set of languages, the first parameters are fixed during the training process of the second parameters of the second network, and the second set of languages ​​is a new language excluded from the first set of languages.

[0069] In an implementation of the present disclosure, the first network includes a first encoder, the second network includes a second encoder, the second encoder shares at least one layer among the multiple layers included in the first encoder, and the total number of layers included in the first encoder is equal to the total number of layers included in the second encoder.

[0070] In an implementation of the present disclosure, in response to a first data space of the first network being different from a second space of the second network, a mapping component is added between the first network and the second network.

[0071] In an implementation of the present disclosure, the output of the multilingual processing model includes any of the following: text, language identification, emotion, domain, or topic associated with the speech data.

[0072] According to an implementation of the present disclosure, a device for multilingual speech processing is provided. The device includes: a recognition module configured to, in response to receiving speech data, identify a network corresponding to the speech data from a first network and a second network included in a multilingual processing model, the first network being associated with a first group of languages, and the second network being associated with a second group of languages; and a providing module configured to provide an output determined by the identified network based on the speech data as an output of the multilingual processing model. The device also includes other modules for implementing the other steps in the above-described method.

[0073] According to an implementation of the present disclosure, an electronic device for implementing method 600 is provided. The electronic device includes a computer processor coupled to a computer-readable memory unit containing instructions that, when executed by the computer processor, implement a method for multilingual speech processing. The method includes: in response to receiving speech data, identifying a network corresponding to the speech data from a first network and a second network included in a multilingual processing model, the first network being associated with a first set of languages, and the second network being associated with a second set of languages; and providing an output determined by the identified network based on the speech data as an output of the multilingual processing model.

[0074] In an implementation of the present disclosure, identifying a network includes: obtaining a first score of speech data by a first network and a second score of speech data by a second network, respectively; and selecting a network from the first network and the second network based on a difference between the first score and the second score.

[0075] In an implementation of the present disclosure, the first score is determined based on a first probability score of at least one first language token recognized from the speech data by the first network, and the second score is determined based on a second probability score of at least one second language token recognized from the speech data by the second network.

[0076] In an implementation of the present disclosure, selecting a network based on a difference between the first score and the second score includes: in response to determining that the difference is above a predetermined threshold, selecting a network corresponding to a larger score from the first network and the second network.

[0077] In an implementation of the present disclosure, selecting a network based on a difference between a first score and a second score includes: in response to determining that the difference is below a predetermined threshold, determining a first average probability score for a first plurality of first tokens recognized from speech data by the first network, and determining a second average probability score for a second plurality of second tokens obtained from speech data by the second network, respectively; and selecting a network from the first network and the second network corresponding to a larger average probability score of the first and second average probability scores.

[0078] In an implementation of the present disclosure, determining a network includes: obtaining a language classification of the speech data, the language classification indicating a language group of the first and second groups to which the speech data belongs; and selecting a network from the first and second networks based on the language classification.

[0079] In an implementation of the present disclosure, first parameters of a first network are trained using a first data set represented by a first set of languages, second parameters of a second network are trained using a second data set represented by a second set of languages, the first parameters are fixed during the training process of the second parameters of the second network, and the second set of languages ​​is a new language excluded from the first set of languages.

[0080] In an implementation of the present disclosure, the first network includes a first encoder, the second network includes a second encoder, the second encoder shares at least one layer of the plurality of layers included in the first encoder, and the total number of layers included in the first encoder is equal to the total number of layers included in the second encoder.

[0081] In an implementation of the present disclosure, in response to a first data space of the first network being different from a second space of the second network, a mapping component is added between the first network and the second network.

[0082] In an implementation of the present disclosure, the output of the multilingual processing model includes any of the following: any one of text, language identification, emotion, domain, or topic associated with the speech data.

[0083] According to an implementation of the present disclosure, a computer program product includes a computer-readable storage medium having program instructions embodied therewith, and the program instructions can be executed by an electronic device to enable the electronic device to perform method 600.

[0084] Figure 7 1 shows a block diagram of a computing device 700 in which various implementations of the present disclosure may be implemented. It should be understood that Figure 7 The computing device 700 shown in FIG is for illustration purposes only and does not in any way imply any limitation on the functionality and scope of the present disclosure. The computing device 700 can be used to implement the above-mentioned methods in the implementation of the present application. Figure 7 As shown, computing device 700 may be a general-purpose computing device and may include at least one or more processors or processing units 710 , a memory 720 , a storage unit 730 , one or more communication units 740 , one or more input devices 750 , and one or more output devices 760 .

[0085] Processing unit 710 may be a physical or virtual processor and may implement various processes based on programs stored in memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 700. Processing unit 710 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0086] The computing device 700 typically includes various computer storage media. Such media can be any media accessible to the computing device 700, including but not limited to volatile and non-volatile media, or removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory), or any combination thereof. The storage unit 730 can be any removable or non-removable medium and can include machine-readable media such as memory, a flash drive, a disk, or another other medium that can be used to store information and / or data and can be accessed in the computing device 700.

[0087] The computing device 700 may also include additional removable / non-removable, volatile / non-volatile storage media. Figure 7 Although not shown, a magnetic disk drive for reading from and / or writing to a removable and non-volatile magnetic disk and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In this case, each drive may be connected to the bus (not shown) via one or more data medium interfaces.

[0088] The communication unit 740 communicates with another computing device via a communication medium. In addition, the functionality of the components in the computing device 700 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 700 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0089] Input device 750 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 760 may be one or more of various output devices, such as a display, speaker, printer, etc. Computing device 700 may also communicate with one or more external devices (not shown) (such as storage devices and display devices) by means of communication unit 740, one or more of which enable a user to interact with computing device 700 or any device (such as a network card, modem, etc.) that enables computing device 700 to communicate with one or more other computing devices (if desired). Such communication may be performed via an input / output (I / O) interface (not shown).

[0090] In some implementations, some or all components of computing device 1000 may also be arranged in a cloud computing architecture rather than integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some implementations, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various implementations, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides applications over a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on servers at a remote location. Computing resources in a cloud computing environment can be consolidated or distributed across locations in remote data centers. Cloud computing infrastructure can provide services through shared data centers, although they appear to be a single access point for users. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, they can be provided from a conventional server or installed directly or otherwise on a client device.

[0091] The functions described herein may be performed, at least in part, by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0092] The program code for performing the methods of the subject matter described herein can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely or partially on the machine, partially on the machine as a stand-alone software package, partially on a remote machine, or entirely on a remote machine or server.

[0093] In the context of the present disclosure, a machine-readable medium may be any tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0094] In addition, although operation is shown in a particular order, this should not be understood as requiring to perform such operation in the particular order shown or in sequence, or requiring to perform all operations shown to achieve the desired result. Under certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details have been included in the above discussion, these details should not be interpreted as limiting the scope of the subject matter described herein, but should be interpreted as describing the features that may be specific to a particular implementation. Some features described in the context of a separate implementation may also be implemented in combination in a single implementation. On the contrary, the various features described in a single implementation may also be implemented individually or in a plurality of implementations in any suitable subcombination.

[0095] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter specified in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0096] It will be appreciated from the foregoing that the presently disclosed technology has been described herein for illustrative purposes only, but various modifications may be made without departing from the scope of the present disclosure. Accordingly, the technology of the present disclosure is not to be limited except as set forth in the appended claims.

[0097] The implementation of the subject matter and functional operations described in this disclosure can be implemented in various systems, digital electronic circuits or computer software, firmware or hardware, including the structures disclosed in this specification and their structural equivalents, or a combination of one or more of them. The implementation of the subject matter described in this specification can be implemented as one or more computer program products, that is, one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by a data processing device or for controlling the operation of the data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition that implements a machine-readable propagation signal, or a combination of one or more of them. The term "data processing unit" or "data processing device" includes all devices, equipment and machines for processing data, such as a programmable processor, a computer or multiple processors or computers. In addition to hardware, the device may also include code that creates an execution environment for the computer program in question, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0098] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or code portions). A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.

[0099] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include or be operatively coupled to receive data from or transfer data to one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices. The processor and memory may be supplemented by or incorporated into dedicated logic circuitry.

[0100] The specification and drawings are to be regarded as exemplary only, in which the exemplary arrangements are examples. As used herein, the use of "or" is intended to include "and / or" unless the context clearly dictates otherwise.

[0101] Although this disclosure contains many details, these details should not be interpreted as limitations on the scope of any disclosure or claimed content, but rather should be interpreted as descriptions of features that may be specific to a particular implementation of a particular disclosure. Certain features described in this disclosure in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented in multiple implementations, either individually or in any suitable subcombination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed, in some cases, one or more features from a claimed combination may be excised from the combination, and a claimed combination may be directed to a subcombination or variations of the subcombination.

[0102] Similarly, although operations are shown in a particular order in the accompanying drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed to achieve the desired results. Furthermore, the separation of various system components in the implementations described in this disclosure should not be understood as requiring such separation in all implementations. Only some implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this disclosure.

Claims

1. A method for multilingual speech processing, comprising: In response to receiving speech data, identifying a network corresponding to the speech data from a first network and a second network included in a multilingual processing model, the first network being associated with a first set of languages ​​and the second network being associated with a second set of languages; and An output determined by the identified network based on the speech data is provided as an output of the multilingual processing model.

2. The method of claim 1 , wherein identifying the network comprises: obtaining a first score of the voice data from the first network and a second score of the voice data from the second network; as well as The network is selected from the first network and the second network based on a difference between the first score and the second score.

3. The method of claim 2 , wherein the first score is determined based on a first probability score of at least one first language token recognized by the first network from the speech data, and the second score is determined based on a second probability score of at least one second language token recognized by the second network from the speech data.

4. The method of claim 3 , wherein selecting the network based on a difference between the first score and the second score comprises: In response to determining that the difference is above a predetermined threshold, The network corresponding to the larger score of the first score and the second score is selected from the first network and the second network.

5. The method of claim 3 , wherein selecting the network based on a difference between the first score and the second score comprises: In response to determining that the difference is below a predetermined threshold, determining a first average probability score for a first plurality of first tokens identified by the first network from the speech data, and determining a second average probability score for a second plurality of second tokens obtained by the second network from the speech data; as well as The network corresponding to the larger average probability score of the first average probability score and the second average probability score is selected from the first network and the second network.

6. The method of claim 1 , wherein determining the network comprises: obtaining a language classification of the speech data, where the language classification indicates a language group in the first group and the second group to which the speech data belongs; as well as The network is selected from the first network and the second network based on the language classification.

7. The method of claim 1 , wherein first parameters of the first network are trained using a first dataset represented by the first set of languages, second parameters of the second network are trained using a second dataset represented by the second set of languages, the first parameters are fixed during the training of the second parameters of the second network, and the second set of languages ​​is a new language excluded from the first set of languages.

8. The method according to claim 1, wherein the first network includes a first encoder, the second network includes a second encoder, the second encoder shares at least one layer among the plurality of layers included in the first encoder, and the total number of layers included in the first encoder is equal to the total number of layers included in the second encoder.

9. The method of claim 1, wherein in response to a first data space of the first network being different from a second space of the second network, a mapping component is added between the first network and the second network.

10. The method of claim 1, wherein the output of the multilingual processing model comprises any of the following: text, language identification, sentiment, domain, or topic associated with the speech data.

11. An electronic device comprising a computer processor coupled to a computer readable memory unit, the memory unit comprising instructions that, when executed by the computer processor, implement a method for multilingual speech processing, the method comprising: In response to receiving speech data, identifying a network corresponding to the speech data from a first network and a second network included in a multilingual processing model, the first network being associated with a first set of languages ​​and the second network being associated with a second set of languages; and An output determined by the identified network based on the speech data is provided as an output of the multilingual processing model.

12. The apparatus of claim 11 , wherein identifying the network comprises: obtaining a first score of the voice data through the first network, and obtaining a second score of the voice data through the second network; as well as The network is selected from the first network and the second network based on a difference between the first score and the second score.

13. The apparatus of claim 12 , wherein the first score is determined based on a first probability score of at least one first language token recognized by the first network from the speech data, and the second score is determined based on a second probability score of at least one second language token recognized by the second network from the speech data.

14. The apparatus of claim 13, wherein selecting the network based on a difference between the first score and the second score comprises: In response to determining that the difference is above a predetermined threshold, The network corresponding to the larger score of the first score and the second score is selected from the first network and the second network.

15. The apparatus of claim 13, wherein selecting the network based on a difference between the first score and the second score comprises: In response to determining that the difference is below a predetermined threshold, determining a first average probability score for a first plurality of first tokens identified by the first network from the speech data, and determining a second average probability score for a second plurality of second tokens obtained by the second network from the speech data; The network corresponding to the larger average probability score of the first average probability score and the second average probability score is selected from the first network and the second network.

16. The apparatus of claim 11, wherein determining the network comprises: obtaining a language classification of the speech data, the language classification indicating a group of languages ​​in the first group and the second group to which the speech data belongs; as well as The network is selected from the first network and the second network based on the language classification.

17. The apparatus of claim 11, wherein first parameters of the first network are trained using a first dataset represented by the first set of languages, second parameters of the second network are trained using a second dataset represented by the second set of languages, the first parameters are fixed during the training process of the second parameters of the second network, and the second set of languages ​​is a new language excluded from the first set of languages.

18. The apparatus of claim 11, wherein the first network comprises a first encoder, the second network comprises a second encoder, the second encoder shares at least one layer among a plurality of layers included in the first encoder, and a total number of layers included in the first encoder is equal to a total number of layers included in the second encoder.

19. The apparatus of claim 11, wherein in response to a first data space of the first network being different from a second space of the second network, a mapping component is added between the first network and the second network, and an output of the multilingual processing model comprises any one of: text, language identifier, emotion, domain, or topic associated with speech data.

20. A non-transitory computer program product, comprising a computer-readable storage medium having program instructions embodied therewith, the program instructions being executable by an electronic device to cause the electronic device to perform a method for multilingual speech processing, the method comprising: In response to receiving speech data, identifying a network corresponding to the speech data from a first network and a second network included in a multilingual processing model, the first network being associated with a first set of languages ​​and the second network being associated with a second set of languages; and An output determined by the identified network based on the speech data is provided as an output of the multilingual processing model.