Voice processing method and device, equipment and storage medium

By updating the speech synthesis model to process tasks in the target language category and generating training sample sets, the problem of insufficient data for low-resource language labeling is solved, and the training efficiency and recognition accuracy of the speech processing model are improved.

CN120220646APending Publication Date: 2025-06-27BYTEDANCE TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510472159.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When training speech recognition models in the prior art, the labeling data of low-resource languages ​​is insufficient, resulting in unstable model processing effects and limited generalization capabilities.

Method used

By obtaining the annotated sample set of target language categories, the trained speech synthesis model is updated so that it can convert text of target language categories into speech, and using the model to generate a training sample set for training speech processing models.

Benefits of technology

High-quality audio data augmentation is achieved, the training sample set used to train speech processing models is expanded, and the training efficiency of the model and the recognition accuracy of low-resource languages ​​are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220646A_ABST
    Figure CN120220646A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice processing method and device, equipment and a storage medium. The method comprises the steps that a labeling sample set of a target language category is acquired, the labeling sample set comprises multiple pairs of samples, and each pair of samples comprises a text sample of the target language category and a voice sample corresponding to the text sample; using the labeled sample set to update a trained speech synthesis model, the speech synthesis model being trained to convert a text of at least one language category into speech of at least one language category, the at least one language category being different from the target language category; and generating a training sample set for the first speech processing model by using the updated speech synthesis model based on the first text of the target language category. Therefore, the speech synthesis model for the target language category can be updated, and the processing effect of the speech processing model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to methods, apparatuses, devices, and computer-readable storage media for speech processing. Background Art

[0002] With the continuous development of science and technology, the processing of speech also accounts for an increasingly large proportion in human-computer interaction applications. For example, text can be converted into corresponding speech signals through recognition and understanding. Correspondingly, speech signals can also be converted into corresponding text through recognition and understanding. Processing speech brings many conveniences to users. Summary of the Invention

[0003] In a first aspect of the present disclosure, there is provided a method for speech processing, including: obtaining an annotated sample set of a target language category, the annotated sample set including multiple pairs of samples, each pair of samples including a text sample of the target language category and a speech sample corresponding to the text sample; using the annotated sample set to update a trained speech synthesis model, the speech synthesis model being trained to convert text of at least one language category into speech of at least one language category, the at least one language category being different from the target language category; and based on a first text of the target language category, using the updated speech synthesis model to generate a training sample set for a first speech processing model.

[0004] In a second aspect of the present disclosure, there is provided a device for speech processing, including: a sample set obtaining module configured to obtain an annotated sample set of a target language category, the annotated sample set including multiple pairs of samples, each pair of samples including a text sample of the target language category and a speech sample corresponding to the text sample; a speech synthesis model updating module configured to use the annotated sample set to update a trained speech synthesis model, the speech synthesis model being trained to convert text of at least one language category into speech of at least one language category, the at least one language category being different from the target language category; and a training sample set generating module configured to, based on a first text of the target language category, use the updated speech synthesis model to generate a training sample set for a first speech processing model.

[0005] In a third aspect of the present disclosure, there is provided an electronic device. The device includes at least one processor; and at least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to execute the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0007] According to a fifth aspect of the present disclosure, there is provided a computer program product including computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the method of the first aspect is implemented.

[0008] It should be understood that the content described in this content part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0011] Figure 2 A flowchart showing a process for speech processing according to some embodiments of the present disclosure;

[0012] Figure 3 A schematic diagram showing an example architecture for updating a trained speech synthesis model according to some embodiments of the present disclosure;

[0013] Figure 4 A schematic diagram showing a performance display effect diagram for an updated speech synthesis model according to some embodiments of the present disclosure;

[0014] Figure 5 A schematic diagram showing an example architecture for generating training speech according to some embodiments of the present disclosure;

[0015] Figure 6 A block diagram showing a device for speech processing according to some embodiments of the present disclosure; and

[0016] Figure 7 A block diagram showing a device capable of implementing one or more embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0018] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0019] In this article, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately executed after "A", but may include one or more intermediate steps.

[0020] It can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0021] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained through appropriate means in accordance with relevant laws and regulations.

[0022] For example, when receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information, so that the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operations of the technical solution of the present disclosure according to the prompt message.

[0023] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0024] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0025] As briefly mentioned above, speech can be converted into corresponding text through recognition and understanding, which has a wide range of applications in fields such as human-computer interaction, voice assistants, meeting records, real-time translation, and so on. Conventionally, an Automatic Speech Recognition (ASR) model that can convert speech into corresponding text can be trained by means of self-supervised pre-training methods based on large-scale unlabeled speech data, cross-lingual transfer learning techniques that utilize high-resource language knowledge, and model architecture optimization for low-resource scenarios. However, when training the ASR model in the above ways, due to insufficient labeled data for low-resource languages, problems such as unstable processing effects and limited generalization ability of the ASR model occur.

[0026] Furthermore, a Text-to-Speech (TTS) model can be used to synthesize training data. That is, the TTS model is used to convert text into synthesized speech, and then the synthesized speech is used for the training of the ASR model. In this way, a large amount of training data is required to train the TTS model, and the naturalness and diversity of the synthesized speech are insufficient, resulting in the ASR model overfitting to the fixed speech patterns generated by the TTS. Correspondingly, voice conversion technology can also be used to convert the speech of a speaker into the speech style of another speaker, thereby expanding the training data. However, it is difficult to maintain the semantic content consistency of the converted speech in this way, and the diversity of the generated data is limited when the number of speakers is small. Additionally, data can also be expanded by adding various noises, reverberations, and other interferences to the original speech signal. In this way, the quality of the expanded data is poor, and no new speech content can be generated.

[0027] In view of this, the present disclosure proposes an improved solution for speech processing. According to the solution of the embodiments of the present disclosure, an annotated sample set of the target language category is obtained. The annotated sample set includes multiple pairs of samples, and each pair of samples includes a text sample of the target language category and a speech sample corresponding to the text sample. Further, using the annotated sample set, an already-trained speech synthesis model is updated. The speech synthesis model is trained to convert text of at least one language category into speech of at least one language category, and the at least one language category is different from the target language category. Based on the first text of the target language category, using the updated speech synthesis model, a training sample set for the first speech processing model is generated.

[0028] In this way, the trained speech synthesis model is updated using the labeled sample set of the target language category, so that the speech synthesis model has the ability to synthesize the speech of the target language category. Furthermore, a large number of samples of the target language category can be generated using the updated speech synthesis model for training the speech processing model. That is, by making the speech synthesis model learn the acoustic features and speech patterns of the speaker from the labeled sample set of the target language category, high-quality audio data augmentation can be achieved, thereby effectively expanding the training sample set for training the speech processing model. In this way, on the basis of the pre-trained speech synthesis model, the speech synthesis model can be enabled to handle tasks for the target language category, thereby generating training samples of the target language category for the speech processing model. In this way, the training efficiency of the speech processing model can be improved.

[0029] Example embodiments of the present disclosure are described below with reference to the accompanying drawings.

[0030] Figure 1 1 is a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 In the environment 100, it is desirable to train and use such a speech synthesis model 130, which is configured for a variety of application environments. An example of the speech synthesis model 130 is a TTS model. Figure 1 In the example of , the speech synthesis model 130 can output speech for the text sample set 112 of the target language category based on the text sample set 112 included in the annotated sample set 110 of the target language category input by the user, etc. In some embodiments, the target language category can refer to a variety of languages ​​formed according to the pronunciation, vocabulary, grammatical rules, etc. of the language. For example, the target language category can indicate languages ​​of different language families. For another example, the target language category can indicate different languages, such as English, Chinese, etc. The annotated sample set 110 of the target language category may include a text sample set 112 and a speech sample set 113. In some embodiments, the text sample set 112 can also be obtained based on the speech sample set 113 by using a speech processing model. For example, the speech sample set 113 can be input into the speech processing model to obtain the text sample set 112. The speech processing model here can indicate that it has been trained, that is, it can handle the task of converting the speech of the target language category into the text of the target language category. Furthermore, in the process of updating the speech synthesis model 130 , the speech synthesis model 130 may be updated based on the difference between the speech samples in the speech sample set 113 and the predicted speech output by the speech synthesis model 130 .

[0031] In some embodiments, the speech synthesis model 130 can be pre-trained. For example, the speech synthesis model already has the ability to convert text in at least one language category into speech in at least one language category. On this basis, the trained speech synthesis model 130 is updated to be used for the task of converting text in a target language category into speech in the target language category. The target language category is different from the at least one language category. Thus, the updated speech synthesis model 130 can be utilized ′ to generate a training sample set for training the speech processing model. Here, the speech processing model refers to a model to be trained to be capable of processing the task of converting speech in a target language category into text in the target language category.

[0032] As Figure 1 shown, the environment 100 includes a model update system 150. Figure 1 The upper part shows the process of the model training phase, and the lower part shows the process of the model application phase. Before training, the parameter values of the speech synthesis model 130 can have the values before update, that is, the pre-trained parameter values obtained through the pre-training process. The speech synthesis model 130 can be trained via a text sample set 112 and a speech sample set 113. During the training process, the parameter values of the speech synthesis model 130 can be updated and adjusted. After the speech synthesis model 130 is trained, the speech synthesis model 130 is obtained ′ . At this time, the parameter values of the speech processing model 130 ′ have been updated, and based on the updated parameter values, the speech synthesis model 130 ′ can be used to implement various types of text processing tasks in the model application phase. For example, the speech synthesis model 130 ′ can be used to implement tasks such as text-to-speech tasks, translation tasks, and so on.

[0033] In the model training phase, based on an annotated sample set 110 of the target language category including the text sample set 112 and the speech sample set 113 for updating the speech synthesis model 130, the model update system 150 can be used to update the speech synthesis model 130. Here, the speech synthesis model 130 can generate a text sample set 112 corresponding to the speech sample set 113 through the already trained speech processing model based on the annotated sample set 110 of the target language category. Subsequently, the model update system 150 can also update the speech synthesis model 130 based on the multiple text samples included in the text sample set 112 and the speech samples of the target language category included in the speech sample set 113. Specifically, the training process can be iteratively performed using a large number of training samples. After training is completed, the speech synthesis model 130 can include knowledge about the task to be processed. In the model application phase, the speech synthesis model 130 ′ (the speech processing model 130 at this time′ (with updated parameter values) can be used to perform corresponding tasks. For example, in a speech recognition task, the speech synthesis model 130 ′ can receive the text 142 and output the training speech 144 in the corresponding target language category for training the speech processing model.

[0034] In Figure 1 , the model update system 150 can include any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. The terminal device can refer to any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server includes, but is not limited to, mainframes, edge computing nodes, computing devices in a cloud environment, and so on.

[0035] It should be understood that Figure 1 the components and arrangements in the illustrated environment 100 are merely examples, and the computing systems suitable for implementing the exemplary implementations described in this disclosure may include one or more different components, other components, and / or different arrangements. The implementations of this disclosure are not limited in this regard.

[0036] Some example embodiments of this disclosure will be described below with continued reference to the accompanying drawings. In the following, the example embodiments will be mainly described with respect to the model update system 150.

[0037] The following will refer to Figure 2 to describe the solution of this disclosure for speech processing. Figure 2 The flowchart of an example process 200 for speech processing according to some embodiments of this disclosure is shown. In the embodiments of this disclosure, it is desired to update such a speech synthesis model 130, which is configured to be updated by an annotated set of annotation samples and then output training speech. That is, after being trained by the annotated sample set, based on the input text 142, it accurately outputs high-quality training speech 144. Furthermore, the training speech 144 can be used to train the speech processing model. The speech processing model is used to process the input speech to achieve a predetermined task or function. An example of the speech processing model can be an ASR model, which converts the input speech into text. It can be understood that the speech processing model can be a model that implements other types of speech processing functions, such as a model for transforming speech styles.

[0038] Therefore, the following will refer to Figures 2 to 5This disclosure describes updating a trained speech synthesis model 130 based on an annotated sample set to accurately output high-quality training speech for training a speech processing model. Figure 2 FIG. 200 is a flowchart showing a process for speech processing according to some embodiments of the present disclosure. Figure 3 FIG. 300 is a schematic diagram showing an example architecture for updating a trained speech synthesis model 130 according to some embodiments of the present disclosure.

[0039] Referring to Figure 2 , at block 210, the model update system 150 obtains an annotated sample set of a target language category. In some embodiments, the annotated sample set includes multiple pairs of samples, each pair including a text sample of the target language category and a speech sample corresponding to the text sample. In some embodiments, the target language category may indicate a language such as English, German, French, Italian, etc. In some embodiments, the target language category may include low-resource languages. A low-resource language may refer to a language lacking large-scale text data, annotation data, or computational resource support, e.g., a language with a small number of speakers.

[0040] In some embodiments, the model update system 150 may utilize a trained speech processing model to obtain an annotated sample set of a target language category. That is, the model update system 150 obtains a target text corresponding to the target speech according to the target speech of the target language category by using a second speech processing model. Further, the target speech and the target text are used as a pair of samples in the annotated sample set. In some embodiments, the second speech processing model is configured to convert speeches of multiple language categories including the target language category into texts of multiple language categories respectively. It can be understood that the second speech processing model has been trained and can be used to process the task of converting a speech of the target language category into a text of the target language category.

[0041] Using the second speech processing model, a target text corresponding to the target speech of the target language category can be obtained, and the target speech and the target text can be labeled as an update sample set for updating the trained speech synthesis model. Thus, the data requirements for updating the trained speech synthesis model and improving the effect of the first speech processing model can be reduced. For example, language can be quickly adapted in this way, and the cost of obtaining training data and the input of human resources can also be reduced.

[0042] In some embodiments, the model update system 150 may generate a text sample set for a target language category based on source text of the target language category. For example, source text of the target language category may be obtained from various types of public resources such as encyclopedias, websites, e-books, and so on. In some examples, the model update system 150 uses a trained speech synthesis model to obtain a speech sample set corresponding to the text sample set respectively, and then labels the text sample set and the speech sample set as a sample set for updating the trained speech synthesis model. However, these are merely exemplary, and the present disclosure does not limit this.

[0043] In some examples, using the labeled sample set of the target language category, the pre-trained speech synthesis model 130 may be supervised based on the labeled data to learn the correspondence between the input and the output. Further, the speech synthesis model 130 is used to generate a training sample set for training the speech processing model. In this way, the recognition accuracy of the speech processing model in low-resource languages can be improved.

[0044] In block 220, the model update system 150 updates the trained speech synthesis model using the labeled sample set. In some embodiments, the speech synthesis model 130 has been trained to convert text of at least one language category into speech of at least one language category, and at least one language category is different from the target language category. In some embodiments, the model update system 150 may use the text sample set 112 included in the labeled sample set 110 and the speech sample set 113 corresponding to the text sample set 112 to update the trained speech synthesis model 130. The speech synthesis model 130 has been trained to handle the task of converting text of at least one language category into speech of at least one language category. On this basis, the model update system 150 uses the labeled sample set 110 belonging to the target language category to update the speech synthesis model 130, so that the updated speech synthesis model 130 ′ can handle the task of converting text of the target language category into speech of the target language category.

[0045] The following will refer to Figure 3 to describe how the model update system 150 updates the speech synthesis model 130 using the labeled sample set 110.

[0046] In some embodiments, the model update system 150 may update the vocabulary for the trained speech synthesis model according to the text set of the target language category. The updated vocabulary includes at least one token for the target speech category. In some embodiments, the text set of the target language category may indicate the text samples included in each pair of samples in the labeled sample set 110 of the target language category obtained by the model update system 150.

[0047] For example, in Figure 3In the example architecture 300 shown, if the model update system 150 obtains a text sample set 112 of a target language category (i.e., a text set), it can expand the vocabulary of the trained speech synthesis model 130. The trained speech synthesis model 130 has the ability to convert text in certain language categories into speech in these language categories, but does not yet have the ability to process text in the target language category. Therefore, based on the collected text sample set 112 of the target language category, the model update system 150 can expand the vocabulary of the trained speech synthesis model 130. In some embodiments, the model update system 150 can use a data compression algorithm to update the vocabulary of the speech synthesis model 130. An example of a data compression algorithm is the byte pair encoding (BPE) algorithm, which is configured to determine a token pair with a frequency greater than a frequency threshold based on the frequencies of adjacent tokens among multiple tokens corresponding to the text set, and replace the token pair with a single token.

[0048] In some examples, the model update system 150 can first determine the frequencies of adjacent tokens among the multiple tokens corresponding to the text sample set 112, and then replace adjacent tokens with a frequency greater than the frequency threshold with a single token not included in the multiple tokens. For example, for a certain sentence in the text sample set 112, the sentence is segmented into individual words or characters, and each word / character is a token. Then, find two adjacent tokens that are the same among these tokens and replace them with a character. For example, for aaabdaaabac, the adjacent byte pair aa appears with a relatively high frequency, so aa can be replaced with the new character Z. Iterate in this way to update the vocabulary of the trained speech synthesis model 130.

[0049] In some embodiments, the model update system 150 can update the dimension of the embedding layer included in the trained speech synthesis model based on the updated vocabulary. In the Figure 3 example architecture 300 shown, the model update system 150 can use the updated vocabulary to determine the text sample tokens corresponding to the text sample set 112. Correspondingly, the model update system 150 can also use the updated embedding layer 313 to obtain the speech features corresponding to the speech set 311, such as timbre features, style features, and other features. In some embodiments, the embedding layer can be used to embed the speech features corresponding to the speaker's speech set into the tokens corresponding to the text sample set respectively. In some examples, the speech set 311 can include speech belonging to the target language category (which can also be referred to as "speaker speech"). The model update system 150 can use the updated attention layer to embed the speech features corresponding to the speech set 311 into the text sample tokens corresponding to the text sample set 112 respectively.

[0050] Further, based on the updated vocabulary, the dimension of the embedding layer 313 of the speech synthesis model 130 can also be updated accordingly. For example, assume that the original vocabulary includes AA word tokens, and the dimension of each word is BB dimensions. Therefore, the current dimension of the embedding layer 313 is AA×BB. Then, as AAA word tokens are added to the vocabulary, the dimension of the embedding layer 313 is updated to (AA + AAA)×BB. In some embodiments, the model update system 150 can initialize the parameters of at least one word token. That is, the model update system 150 initializes the parameters of at least one word token added to the vocabulary.

[0051] By building on the original vocabulary included in the trained speech synthesis model 130, the vocabulary of the trained speech synthesis model 130 can be quickly updated. This enables rapid expansion of support for new voices based on the existing speech synthesis model 130, reduces the update time of the trained speech synthesis model 130, and simplifies the development process of the speech synthesis model 130 for processing target language categories.

[0052] For the extended vocabulary of the trained speech synthesis model 130, the model update system 150 can update the trained speech synthesis model 130 in the following manner. For ease of understanding below, an example of any pair of samples (e.g., the first pair of samples) in the labeled sample set will be used to describe how the model update system 150 updates the trained speech synthesis model 130. The first pair of samples includes a first text sample and a first speech sample corresponding to the first text sample.

[0053] In some embodiments, the model update system 150 can use the trained speech synthesis model to generate a first predicted speech corresponding to the first text sample. Further, the model update system 150 can update the trained speech synthesis model 130 based on the difference between the first predicted speech and the first speech sample. Referring to Figure 3 Figure, the model update system 150 can input the first text sample in the text sample set 112 into the trained speech synthesis model 130 to obtain the predicted speech 317 corresponding to the first text sample.

[0054] In some examples, the model update system 150 may input a first text sample and a certain voice in the voice set 311 into the trained voice synthesis model 130. Further, through the attention layer included in the trained voice synthesis model 130, the acoustic sample tokens 315 corresponding to the first text sample and the voice are obtained. That is, the model update system 150 embeds the voice features corresponding to the voice into the text sample tokens 314 corresponding to the first text sample by using the attention layer, and then generates the acoustic sample tokens 315. The model update system 150 uses the vocoder 316 included in the trained voice synthesis model 130 to generate a predicted voice 317 for the acoustic sample tokens 315. Subsequently, the model update system 150 updates the trained voice synthesis model 130 according to the difference between the predicted voice 317 and the first voice sample corresponding to the first text sample.

[0055] In some embodiments, the model update system 150 may update the trained voice synthesis model 130 by updating the parameters of the attention layer in the voice synthesis model 130. For example, during the process of updating the trained voice synthesis model 130 by using the labeled sample set, the parameters of the self-attention layer included in the trained voice synthesis model 130 are also updated accordingly. In this embodiment, the parameters of other model layers in the voice synthesis model 130 except the attention layer may not be updated. That is, only the parameters of the attention layer can be fine-tuned. For the voice synthesis model 130 based on the attention mechanism, the attention layer is the most important model layer. By only fine-tuning or updating the parameters of the attention layer, on the one hand, it can ensure that the voice synthesis model 130 has the ability to process the target language category, and on the other hand, it can minimize the training cost.

[0056] In some embodiments, the model update system 150 may update at least one weight included in the voice synthesis model 130 by using a predetermined learning rate, and then update the trained voice synthesis model 130. In some embodiments, the predetermined learning rate is used to control the numerical range for updating at least one weight. In some examples, the model update system 150 may use a predetermined learning rate (for example, a relatively small learning rate) to update at least one weight included in the voice synthesis model 130. In some embodiments, the learning rate can indicate a key hyperparameter in deep learning, which can determine the speed of weight update of the trained voice synthesis model 130 during the training process. By using the predetermined learning rate to update at least one weight included in the trained voice synthesis model 130, it can ensure the ′ stability and performance of the updated voice synthesis model 130.

[0057] In some embodiments, the model update system 150 may utilize regularization to update the trained speech synthesis model 130. In some examples, the model update system 150 may limit the complexity of the updated speech synthesis model 130 by adding a predetermined rule ′ so as to improve the generalization ability of the updated speech synthesis model 130 ′ . For example, during the process of updating the trained speech synthesis model 130, a rule (e.g., "data within XX hours is valid") is added to limit the complexity of the updated speech synthesis model 130 ′ . Alternatively, the model update system 150 may limit the complexity of the updated speech synthesis model 130 by adding a penalty term to the loss function ′ so as to prevent overfitting and improve the generalization ability of the updated speech synthesis model 130 ′ .

[0058] In summary, based on the trained speech synthesis model 130, the speech synthesis model 130 is further trained. In this way, a relatively high synthetic speech quality can be maintained, so that the updated speech synthesis model 130 ′ is versatile and also improves the computational efficiency of the updated speech synthesis model 130 ′ .

[0059] The above describes how the model update system 150 updates the trained speech synthesis model 130 based on the labeled sample set. Next, it continues to describe how the model update system 150 evaluates the quality of the updated speech synthesis model 130 ′ .

[0060] In some embodiments, the model update system 150 may utilize the updated speech synthesis model 130 ′ to generate a second predicted speech corresponding to a second text of a target language category. Accordingly, the model update system 150 may utilize a third speech processing model to obtain a predicted text for the second predicted speech. Subsequently, the model update system 150 may determine the performance of the updated speech synthesis model 130 based on the difference between the predicted text and the second text and a standard threshold ′ . It should be understood that the model update system 150 may utilize the updated speech synthesis model 130 ′ to generate a predicted speech corresponding to any text (e.g., the second text) among multiple texts of a target language category. Furthermore, based on the difference between the predicted text corresponding to the predicted speech and any text and a standard threshold, the performance of the updated speech synthesis model 130 ′ is determined. For ease of understanding in the following description, the second text and its corresponding second predicted speech will be used as an example for description

[0061] In some embodiments, the third speech processing model is configured to convert speech in one or more language categories into text in one or more language categories, where the one or more language categories include the target language category. In some examples, the third speech processing model and the second speech processing model may be the same trained speech processing model, or may be different and separately trained speech processing models. The present disclosure does not limit this. The model update system 150 may utilize the already trained third speech processing model to obtain the predicted text of the second predicted speech.

[0062] That is, the model update system 150 may utilize the updated speech synthesis model 130 ′ to generate the second predicted speech of the second text. Correspondingly, the model update system 150 utilizes another already trained speech processing model to obtain the predicted text corresponding to the second predicted speech. Subsequently, the model update system 150 determines the difference between the predicted text and the second text, and determines the performance of the updated speech synthesis model 130 ′ based on the difference between the two and a standard threshold. In some embodiments, the model update system 150 may set an evaluation metric for the second predicted speech based on pronunciation accuracy to evaluate the performance of the updated speech synthesis model 130 ′ .

[0063] In some embodiments, during the process of updating the trained speech synthesis model 130 based on the labeled sample set within different update durations, the performance of the updated speech synthesis model 130 ′ also changes accordingly. In some embodiments, an updated speech synthesis model 130 ′ with performance exceeding the standard threshold is required to output the training speech for training the first speech processing model. Then, during the process of updating the trained speech synthesis model 130 based on the labeled sample set, if it is known that after the performance of the speech synthesis model 130 exceeds a certain threshold (e.g., 0.2 or any other threshold), the training speech output can improve the performance of the first speech processing model, then this threshold can be set as the standard threshold for evaluating the performance of the updated speech synthesis model 130 ′ . In this way, through the method of the present disclosure, the trained speech synthesis model 130 can be updated within a shorter update duration, so that the updated speech synthesis model 130 ′ generates high-quality synthetic data.

[0064] The following is a reference Figure 4 to describe the performance of the speech synthesis model 130 updated by using the labeled sample set different numbers of times during the process of updating the trained speech synthesis model 130 at different update durations ′ .Figure 4 Shows a demonstration effect diagram of the performance of an updated speech synthesis model according to some embodiments of the present disclosure.

[0065] As Figure 4 shown in the example effect diagram, the abscissa indicates the duration of updating the trained speech synthesis model 130 based on the labeled sample set. The ordinate indicates the performance of the updated speech synthesis model 130 ′ , that is, the smaller the error rate of the updated speech synthesis model 130 ′ outputting the training speech, the better the performance of the updated speech synthesis model 130 ′ . As Figure 4 shown, the broken line 411 in Figure 4 corresponds to the labeled sample set being used 2 times (this is just exemplary, for example, it can be any other appropriate number) during the process of updating the trained speech synthesis model 130. As Figure 4 shown, the broken line 412 in Figure 4 corresponds to the labeled sample set being used 5 times (this is just exemplary, for example, it can be any other appropriate number) during the process of updating the trained speech synthesis model 130. As

[0066] From Figure 4 it can be seen that through the solution of the present disclosure, by using the labeled sample set a predetermined number of times (for example, 10 times or any other appropriate number) and using the speaker speech for a predetermined duration (for example, 100 hours or any other appropriate duration) during the process of updating the trained speech synthesis model 130, the updated speech synthesis model 130 ′ can generate training speech that is closer to the speaker speech.

[0067] In this way, based on multi-dimensional performance evaluation, it can be ensured that the updated speech synthesis model 130 ′ can meet the application requirements, thereby also improving the level of the training speech being closer to the speaker speech in terms of naturalness and clarity.

[0068] The above describes how to evaluate the performance of the updated speech synthesis model 130 ′ . The following continues to refer to Figure 2 and Figure 5 to describe how to use the updated speech synthesis model 130 ′ to generate a training sample set for the speech processing model. Figure 5A schematic diagram of an example architecture 500 for generating training speech according to some embodiments of the present disclosure is shown.

[0069] Continuing with process 200, at block 230, the model update system 150 generates a training sample set for the first speech processing model based on the first text of the target language category, using the updated speech synthesis model. In some embodiments, the model update system 150 can determine the training samples for the first speech processing model according to the training speech 144 ′ generated by the updated speech synthesis model 130. It should be understood that the model update system 150 can use the updated speech synthesis model 130 ′ to generate the training speech corresponding to any text (e.g., the first text) among multiple texts of the target language category. For ease of understanding in the following description, the first text will be used as an example to describe how to generate the training sample set for the speech processing model.

[0070] In some embodiments, the model update system 150 can use the updated vocabulary in the updated speech synthesis model 130 ′ to determine the text tokens corresponding to the first text. Correspondingly, the model update system 150, based on the text tokens and speech features, can use the attention layer included in the updated speech synthesis model 130 ′ to generate the acoustic tokens of the text tokens. Then, the model update system 150 can generate the training speech based on the acoustic tokens using the vocoder. Furthermore, the training speech and the first text are used as a pair of training samples in the training sample set. Similarly, for other texts among the multiple texts belonging to the target language category, the above method can also be used to obtain the training speech corresponding to the text, and then the text and its corresponding training speech are used as a pair of training samples in the training sample set.

[0071] In the example architecture 500 as Figure 5 shown, the model update system 150 can use the vocabulary of the updated speech synthesis model 130 ′ to obtain the text tokens 514 corresponding to the text 142 (i.e., the first text). Correspondingly, the model update system 150 can also obtain the speech features corresponding to the speech 511 according to the speech 511 belonging to the target language category, using the embedding layer (e.g., the style embedding layer) included in the updated speech synthesis model 130 ′ In some embodiments, the embedding layer can indicate embedding the speech features corresponding to the speaker's speech into the tokens corresponding to the text. In some embodiments, the speech 511 can include the timbre of the speaker, such as the A timbre of speaker A, the B timbre of speaker B, etc. In some embodiments, the speech 511 can include the style of the speaker, such as a rigorous style, a humorous style, etc.

[0072] Further, after the model update system 150 determines the speech features of the text token 514 and the speech 511 (such as timbre features, style features, and other features), it can utilize the updated speech synthesis model 130 ′ The included attention layer to generate the acoustic token 515 of the text token 514. That is, the model update system 150 can utilize the attention layer to embed the speech features corresponding to the speech 511 into the text token 514 to generate the acoustic token 515. Thus, training speech with rich speaker styles and speech variations can be generated. Subsequently, the model update system 150 utilizes the vocoder 316 to generate the training speech 144 corresponding to the acoustic token 515. Furthermore, the training speech 144 and the text 142 can be used as a pair of training samples for training the first speech processing model. In some embodiments, a pair of samples for training the first speech processing model may include the text 142 and the speech corresponding to the text 142.

[0073] In this way, by using the updated speech synthesis model to collect a training sample set for training the speech processing model, the effect of the speech processing model in processing tasks can be improved. Further, based on the training speech with rich speaker styles generated by the updated speech synthesis model, the robustness of the speech processing model is improved.

[0074] In summary, using the labeled sample set of the target language category to update the trained speech synthesis model enables the speech synthesis model to have the ability to synthesize speech of the target language category. Further, a large number of samples of the target language category can be generated using the updated speech synthesis model for training the speech processing model. That is, by enabling the speech synthesis model to learn the acoustic features and speech patterns of the speaker from the labeled sample set of the target language category, high-quality audio data augmentation can be achieved, thereby effectively expanding the training sample set for training the speech processing model. In this way, based on the pre-trained speech synthesis model, the speech synthesis model can be enabled to process tasks for the target language category, thereby generating training samples of the target language category for the speech processing model. In this way, the training efficiency of the speech processing model can be improved.

[0075] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above methods or processes.

[0076] Figure 6 A schematic structural block diagram of an apparatus 600 for speech processing according to certain embodiments of the present disclosure is shown. The apparatus 600 can be implemented as or included in the model update system 150. Each module / component in the apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.

[0077] AsFigure 6 As shown in Figure 6 , the apparatus 600 includes a sample set acquisition module 610 configured to acquire an annotated sample set of a target language category. The annotated sample set includes multiple pairs of samples, and each pair of samples includes a text sample of the target language category and a speech sample corresponding to the text sample. The apparatus 600 further includes a speech synthesis model update module 620 configured to update a trained speech synthesis model using the annotated sample set. The speech synthesis model is trained to convert text of at least one language category into speech of at least one language category, and the at least one language category is different from the target language category. The apparatus 600 further includes a training sample set generation module 630 configured to generate a training sample set for a first speech processing model based on a first text of the target language category using the updated speech synthesis model.

[0078] In some embodiments, the speech synthesis model update module 620 is further configured to update a vocabulary for the trained speech synthesis model based on a text set of the target language category. The updated vocabulary includes at least one token for the target speech category; update the dimension of the embedding layer included in the trained speech synthesis model based on the updated vocabulary; and initialize the parameters of at least one token.

[0079] In some embodiments, the speech synthesis model update module 620 is further configured to update the vocabulary using a data compression algorithm. The data compression algorithm is configured to determine a token pair with a frequency greater than a frequency threshold based on the frequencies of adjacent tokens among multiple tokens corresponding to the text set, and replace the token pair with a single token.

[0080] In some embodiments, the training sample set generation module 630 is further configured to determine text tokens corresponding to the first text using the updated vocabulary in the updated speech synthesis model; determine speech features corresponding to the speech based on the speech of the target language category using the embedding layer included in the updated speech synthesis model; and generate acoustic tokens of the text tokens based on the text tokens and the speech features using the attention layer included in the updated speech synthesis model; generate training speech using a vocoder based on the acoustic tokens, and use the training speech and the first text as a pair of training samples in the training sample set.

[0081] In some embodiments, the speech includes at least one of the following: the timbre of the speaker, or the style of the speaker.

[0082] In some embodiments, the sample set acquisition module 610 is further configured to acquire a target text corresponding to the target speech using a second speech processing model based on the target speech of the target language category. The second speech processing model is configured to convert speech of multiple language categories into text of multiple language categories respectively, and the multiple language categories include the target language category; and use the target speech and the target text as a pair of samples in the annotated sample set.

[0083] In some embodiments, the speech synthesis model update module 620 is further configured to, for a first pair of samples in the labeled sample set, where the first pair of samples includes a first text sample of a target language category and a first speech sample corresponding to the first text sample, use the trained speech synthesis model to generate a first predicted speech corresponding to the first text sample; and update the trained speech synthesis model based on the difference between the first predicted speech and the first speech sample.

[0084] In some embodiments, updating the speech synthesis model includes at least one of the following: updating the parameters of the attention layer in the speech synthesis model, updating at least one weight included in the speech synthesis model using a predetermined learning rate, where the predetermined learning rate is used to control the numerical range for updating the at least one weight, or updating the speech synthesis model using regularization.

[0085] In some embodiments, the apparatus 600 further includes a performance determination module configured to use the updated speech synthesis model to generate a second predicted speech corresponding to a second text of the target language category; use a third speech processing model to obtain a predicted text for the second predicted speech, where the third speech processing model is configured to convert speech of one or more language categories into text of one or more language categories, and the one or more language categories include the target language category; and determine the performance of the updated speech synthesis model based on the difference between the predicted text and the second text and a standard threshold.

[0086] The units and / or modules included in the apparatus 600 may be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units and / or modules in the apparatus 600 may be at least partially implemented by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0087] It should be understood that one or more steps in the above methods may be performed by a suitable electronic device or a combination of electronic devices. Such an electronic device or combination of electronic devices may include, for example, Figure 1 the model update system 150 in

[0088] Figure 7 A block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood thatFigure 7 The illustrated electronic device 700 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Figure 7 The illustrated electronic device 700 can be used to implement Figure 1 the model update system 150.

[0089] As Figure 7 illustrated, the electronic device 700 is in the form of a general-purpose electronic device. The components of the electronic device 700 can include, but are not limited to, one or more processors 710 or processing units, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processor 710 can be an actual or virtual processor and can execute various processes according to the programs stored in the memory 720. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 700.

[0090] The electronic device 700 generally includes multiple computer storage media. Such media can be any accessible media available to the electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 700.

[0091] The electronic device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 7 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 720 can include a computer program product 725 having one or more program modules that are configured to execute the various methods or actions of the various embodiments of the present disclosure.

[0092] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 700 can be implemented with a single computing cluster or multiple computer machines that are capable of communicating via a communication connection. Accordingly, the electronic device 700 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.

[0093] The input device 750 can be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 760 can be one or more output devices such as a display, speaker, printer, etc. The electronic device 700 can also communicate, as needed, with one or more external devices (not shown) via the communication unit 740, the external devices such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 700, or communicate with any device that enables the electronic device 700 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0094] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.

[0095] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0096] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0097] Computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0098] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified function or act, or by a combination of dedicated hardware and computer instructions.

[0099] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled artisans in the art to understand the various implementations disclosed herein.

Claims

1. A method for speech processing, comprising: Acquire a labeled sample set of a target language category, the labeled sample set comprising a plurality of pairs of samples, each pair of samples comprising a text sample of the target language category and a speech sample corresponding to the text sample; Using the annotated sample set, updating a trained speech synthesis model, the speech synthesis model being trained to convert text in at least one language category into speech in the at least one language category, the at least one language category being different from the target language category; as well as Based on the first text of the target language category, a training sample set for a first speech processing model is generated using the updated speech synthesis model.

2. The method according to claim 1, wherein updating the trained speech synthesis model comprises: Based on the text set of the target language category, updating a vocabulary for the trained speech synthesis model, the updated vocabulary including at least one word element for the target speech category; Based on the updated vocabulary, updating the dimension of the embedding layer included in the trained speech synthesis model; as well as Initialize parameters of the at least one word-gram.

3. The method of claim 2, wherein updating the vocabulary for the trained speech synthesis model comprises: The word list is updated using a data compression algorithm, wherein the data compression algorithm is configured to determine word-gram pairs greater than a frequency threshold based on frequencies of adjacent word-grams in a plurality of word-grams corresponding to the text set, and replace the word-gram pairs with a single word-gram.

4. The method according to claim 1, wherein generating a training sample set for a speech processing model comprises: Determining a text word corresponding to the first text using the updated word list in the updated speech synthesis model; Based on the speech of the target language category, determining speech features corresponding to the speech using the embedding layer included in the updated speech synthesis model; as well as Based on the text word unit and the speech feature, using the attention layer included in the updated speech synthesis model, generate an acoustic word unit of the text word unit; Based on the acoustic word unit, a training speech is generated by using a vocoder, so that the training speech and the first text are used as a pair of training samples of the training sample set.

5. The method according to claim 4, wherein the speech comprises at least one of the following: The speaker's timbre, or The speaker's style.

6. The method according to claim 1, wherein obtaining the annotated sample set of the target language category comprises: Based on the target speech of the target language category, using a second speech processing model to obtain a target text corresponding to the target speech, the second speech processing model being configured to convert speech of multiple language categories into text of the multiple language categories, respectively, the multiple language categories including the target language category; as well as The target speech and the target text are taken as a pair of samples in the annotated sample set.

7. The method of claim 1, wherein updating the trained speech synthesis model comprises: For a first pair of samples in the labeled sample set, the first pair of samples includes a first text sample of the target language category and a first speech sample corresponding to the first text sample, Generate a first predicted speech corresponding to the first text sample using the trained speech synthesis model; as well as The trained speech synthesis model is updated based on the difference between the first predicted speech and the first speech sample.

8. The method according to claim 1, wherein updating the speech synthesis model comprises at least one of the following: updating the parameters of the attention layer in the speech synthesis model, updating at least one weight included in the speech synthesis model using a predetermined learning rate, wherein the predetermined learning rate is used to control a numerical range for updating the at least one weight, or The speech synthesis model is updated using regularization.

9. The method according to claim 1, further comprising: generating, using the updated speech synthesis model, a second predicted speech corresponding to a second text of the target language category; Obtaining predicted text for the second predicted speech using a third speech processing model, wherein the third speech processing model is configured to convert speech of one or more language categories into text of the one or more language categories, respectively, wherein the one or more language categories include the target language category; as well as Based on the difference between the predicted text and the second text and a standard threshold, the performance of the updated speech synthesis model is determined.

10. A device for speech processing, comprising: A sample set acquisition module is configured to acquire a labeled sample set of a target language category, wherein the labeled sample set includes a plurality of pairs of samples, each pair of samples including a text sample of the target language category and a speech sample corresponding to the text sample; a speech synthesis model updating module, configured to update a trained speech synthesis model using the annotated sample set, wherein the speech synthesis model is trained to convert text of at least one language category into speech of the at least one language category, the at least one language category being different from the target language category; as well as The training sample set generation module is configured to generate a training sample set for a first speech processing model based on the first text of the target language category and using the updated speech synthesis model.

11. An electronic device, comprising: at least one processor; as well as At least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the electronic device to perform the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program, wherein the computer program implements the method according to any one of claims 1 to 9 when executed by a processor.

Citation Information

Cited By

  • Cross-modal conversion method and system suitable for multiple languages and multiple voices, and medium

    CN120766659A