NPU-based Chinese-English bilingual text-to-speech conversion method and system

By optimizing the text processing and acoustic feature encoding of the VITS model on the NPU, efficient and natural speech generation of Chinese-English bilingual text to speech is achieved, solving the problems of high resource consumption and low speech quality, and adapting to NPU hardware deployment.

CN120932627APending Publication Date: 2025-11-11GUANGZHOU BAOLUN ELECTRONICS CO LTD

Patent Information

Application Number
CN202511117623.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies consume a lot of resources and produce low-quality speech during the Chinese-English bilingual text-to-speech process, and the VITS model is difficult to deploy directly on fixed-shape NPU hardware.

Method used

A bilingual Chinese-English text-to-speech method based on NPU is adopted. By segmenting and converting the mixed Chinese and English text into words, phoneme ID and language ID sequences are generated. The phoneme duration and acoustic feature distribution are predicted by combining text latent variables. The model is optimized using random noise tensors and independent submodules to achieve efficient conversion.

Benefits of technology

It efficiently processes mixed Chinese and English text within a single model, generating high-quality speech data with natural pronunciation and low resource consumption, and is adaptable to NPU hardware deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932627A_ABST
    Figure CN120932627A_ABST
Patent Text Reader

Abstract

The invention discloses a Chinese-English bilingual text-to-speech conversion method and system based on NPU, and belongs to the technical field of speech processing, and the method comprises the steps: carrying out the word segmentation processing and phoneme conversion of a Chinese-English mixed text based on the language type of each segment in the Chinese-English mixed text, and obtaining a text input sequence; performing vector combination on the phoneme ID sequence and the language ID sequence, and extracting text semantic features from a vector combination result to obtain text hidden variables; according to the text hidden variable, predicting the duration of each phoneme and a prior distribution parameter in an acoustic potential feature space; performing multi-level transformation on the phoneme alignment result and the prior distribution parameter to obtain a bilingual acoustic feature sequence; and converting the bilingual acoustic feature sequence into a corresponding voice waveform. Therefore, by implementing the method, the device and the system, the problems that the Chinese-English bilingual text occupies more resources in the voice conversion process and the output voice quality is lower in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of speech processing technology, specifically relating to a method and system for converting Chinese and English bilingual text to speech based on an NPU. Background Technology

[0002] Text-to-speech (TTS) is a technology that converts written text into human-readable speech. The most common model for implementing TTS is VITS (High Naturalness End-to-End Model), which can generate a corresponding speech waveform from input text in one language in a single step through autoencoding, while maintaining the naturalness of the speech.

[0003] While a single VITS model can convert monolingual text to speech, in situations like simultaneous interpretation where multiple languages ​​need to be converted to speech at the same time, deploying a separate model for each language leads to high system complexity, high resource consumption, and model redundancy. However, concatenating independent monolingual models or using a general multilingual model often results in unnatural pronunciation, disrupted rhythm, and inconsistent timbre during Chinese-English code conversion. Furthermore, the original VITS model architecture includes dynamic shape operations and stochastic operations, making it difficult to directly and efficiently deploy on acceleration hardware such as NPUs (Neural Processing Units) that require fixed input shapes. Summary of the Invention

[0004] This application proposes an NPU-based method and system for converting Chinese and English bilingual text to speech, which can solve the problems of high resource consumption and low output speech quality in the existing technology during the process of converting Chinese and English bilingual text to speech.

[0005] The first aspect of this application provides a Chinese-English bilingual text-to-speech method based on an NPU, the method being applied to a VITS model deployed on an NPU, comprising:

[0006] Based on the language type of each segment in the mixed Chinese-English text, word segmentation and phoneme conversion are performed on the mixed Chinese-English text to obtain a text input sequence; wherein, the text input sequence includes a phoneme ID sequence and a language ID sequence;

[0007] The phoneme ID sequence and language ID sequence are combined into vectors, and textual semantic features are extracted from the vector combination results to obtain textual latent variables;

[0008] Based on textual latent variables, predict the duration of each phoneme and its prior distribution parameters in the acoustic latent feature space;

[0009] The phoneme alignment results based on the duration and the prior distribution parameters are transformed at several levels to obtain the bilingual acoustic feature sequence.

[0010] The bilingual acoustic feature sequence is converted into the corresponding speech waveform.

[0011] The above scheme first converts the mixed Chinese-English text into a phoneme sequence, performing different word segmentation processes on Chinese and English separately during the conversion. This yields both phoneme ID sequences and language ID sequences that can guide pronunciation style, providing support for subsequent speech waveform encoding based on the acoustic features of each language. Then, text latent variables representing deep semantic information, such as prosodic features, are extracted from the mixed Chinese-English phoneme input sequence. The duration of each phoneme is predicted, and the text features are mapped to the speech latent space, yielding corresponding prior distribution parameters. This provides data for subsequent conversion into a complex acoustic feature distribution. Phoneme alignment associates each acoustic feature frame with its corresponding text feature, aligning the duration of the text and speech. The alignment results are then adjusted based on the prior distribution parameters to obtain a bilingual acoustic feature sequence containing precise acoustic details. This allows for the accurate conversion of mixed bilingual text into high-quality speech data using only a single model.

[0012] In one possible implementation of the first aspect, based on the language type of each segment in the mixed Chinese-English text, word segmentation and phoneme conversion are performed on the mixed Chinese-English text to obtain the text input sequence, specifically:

[0013] The language type of the mixed Chinese and English text is identified in segments, and the mixed Chinese and English text is segmented into several text units based on the identification results.

[0014] Based on the language type of the text unit, the text unit is converted into a phoneme sequence;

[0015] The Chinese and English phonemes in the phoneme sequence are associated with the corresponding speech information, and the associated phoneme sequence is converted into a data ID sequence to obtain a phoneme ID sequence and a language ID sequence.

[0016] The above scheme first identifies Chinese and English characters in mixed Chinese-English text, then performs word segmentation and phoneme conversion based on the different language types, obtaining Chinese and English phoneme sequences. Next, it associates each phoneme with corresponding language information, resulting in a phoneme ID sequence and a language ID sequence. These ID sequences determine the language type and specific content of each phoneme. The language ID sequence guides the corresponding pronunciation style in subsequent speech conversion, while the phoneme ID sequence specifies the specific pronunciation content.

[0017] In one possible implementation of the first aspect, the text unit is converted into a phoneme sequence according to the language type of the text unit, specifically as follows:

[0018] The text units belonging to Chinese are converted into Pinyin sequences, and the phonemes in the Pinyin sequences are mapped to the corresponding initials and finals by taking into account tone information;

[0019] The text units belonging to English are converted into English phoneme sequences.

[0020] In one possible implementation of the first aspect, the phoneme ID sequence and the language ID sequence are combined into vectors, and textual semantic features are extracted from the vector combination result to obtain textual latent variables, specifically:

[0021] Map the phoneme ID sequence and the language ID sequence to phoneme embedding vectors and language embedding vectors, respectively;

[0022] During the mapping process, the phoneme ID sequence is positionally encoded according to the position of each phoneme in the phoneme ID sequence.

[0023] Based on the positional encoding results, the phoneme embedding vector and the language embedding vector are combined to generate a hybrid embedding vector;

[0024] Textual semantic features are extracted by learning the contextual relationships in the hybrid embedding vectors, thus obtaining textual latent variables.

[0025] The above scheme uses positional encoding to mark the location of each phoneme in the sequence, assisting in subsequent embedding vector combination to form a hybrid embedding vector that includes content, sequence, and linguistic style. Then, by performing context-dependent feature extraction on the hybrid embedding vector, phoneme-level textual latent variables are obtained.

[0026] In one possible implementation of the first aspect, the duration of each phoneme and its prior distribution parameters in the acoustic feature latent space are predicted based on textual latent variables, specifically as follows:

[0027] Based on the textual latent variables, the pronunciation duration of each phoneme is predicted in both Chinese and English environments to obtain the duration; wherein, the logarithm of the pronunciation duration is taken.

[0028] A prior distribution model is constructed to map textual latent variables to the acoustic feature latent space, yielding the mean and log-variance of the prior distribution parameters. These prior distribution parameters provide acoustic features for subsequent speech waveform generation.

[0029] The above scheme determines the duration by predicting how many frames of speech each phoneme corresponds to using latent variables in the text. Then...

[0030] In one possible implementation of the first aspect, the phoneme alignment result based on the duration and the prior distribution parameters are transformed at several levels to obtain a bilingual acoustic feature sequence, specifically:

[0031] Align the acoustic feature frames corresponding to the duration with the text latent variables to obtain the phoneme alignment result;

[0032] Based on a preset random noise tensor, latent variables are obtained by sampling on the prior distribution parameters; wherein, the random noise tensor is a random number of fixed length;

[0033] The latent variables are subjected to several levels of affine transformation, and the transformation results are adjusted according to the phoneme alignment results during each level of affine transformation to obtain a bilingual acoustic feature sequence; wherein the affine transformation is reversible.

[0034] The above scheme uses a fixed-length random noise tensor for random sampling, which ensures that the length of the sampled content is fixed and can be implemented on an NPU.

[0035] In one possible implementation of the first aspect, the acoustic feature frames corresponding to the duration are aligned with text latent variables to obtain the phoneme alignment result, specifically as follows:

[0036] The acoustic feature distribution of each phoneme is determined based on the number of acoustic feature frames corresponding to the duration.

[0037] Based on the acoustic feature distribution, by extending the text latent variables in the time dimension, each acoustic feature frame is aligned with the text semantic features to obtain the phoneme alignment result.

[0038] The above scheme aligns acoustic feature frames with text latent variables to align the nonlinear relationship between text and speech, and the resulting phoneme alignment can provide the rhythm of human speech.

[0039] In one possible implementation of the first aspect, the VITS model is specifically as follows:

[0040] The duration of each phoneme is predicted based on the independent sub-modules provided by the VITS model; wherein, the independent sub-modules can be executed or deployed independently and run on a CPU or DSP.

[0041] When the VITS model performs random sampling, it loads and directly uses the random noise tensor.

[0042] In the process of converting mixed Chinese and English text into a bilingual acoustic feature sequence, the first operator that cannot be converted by the NPU is identified, and data replacement is performed on the first operator; wherein, the data replacement includes replacing comparison operators, replacing cumulative summation operations, and replacing index assignment operations.

[0043] The above scheme enables the VITS model for text-to-speech conversion to be deployed on NPUs that require fixed-shape inputs. It achieves this by using independently deployable submodules to predict duration, thus ensuring the normal processing of outputs with variable durations. Introducing random noise tensors to handle random sampling ensures the consistency of the model's output. Furthermore, substitution operations transform difficult-to-convert operators into those more easily understood and executed by the NPU.

[0044] The second aspect of this application provides an NPU-based Chinese-English bilingual text-to-speech system, the system comprising: a text unification module, a text feature extraction module, a data encoding module, an acoustic feature sequence acquisition module, and a speech conversion module;

[0045] The text unification module is used to perform word segmentation and phoneme conversion on the mixed Chinese and English text based on the language type of each segment, to obtain a text input sequence; wherein, the text input sequence includes a phoneme ID sequence and a language ID sequence;

[0046] The text feature extraction module is used to combine phoneme ID sequences and language ID sequences into vectors, and extract text semantic features from the vector combination results to obtain text latent variables.

[0047] The data encoding module is used to predict the duration of each phoneme and its prior distribution parameters in the acoustic latent feature space based on text latent variables.

[0048] The acoustic feature sequence acquisition module is used to perform several levels of transformation on the phoneme alignment result based on the duration and the prior distribution parameters to obtain a bilingual acoustic feature sequence.

[0049] The speech conversion module is used to convert the bilingual acoustic feature sequence into the corresponding speech waveform.

[0050] A third aspect of this application provides a terminal device, the device comprising: a terminal device including a processor and a memory, the memory storing a computer program, wherein the processor executes the computer program to implement the steps of the NPU-based Chinese-English bilingual text-to-speech method as described in any one of the embodiments of this application. Attached Figure Description

[0051] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0052] Figure 1This is a schematic diagram illustrating the specific process of an NPU-based Chinese-English bilingual text-to-speech method according to an embodiment of this application;

[0053] Figure 2 This is a structural diagram of an NPU-based Chinese-English bilingual text-to-speech system provided in one embodiment of this application;

[0054] Figure 3 This application provides a structural diagram of a terminal device according to one embodiment. Detailed Implementation

[0055] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0056] It should be understood that the step numbers used in the text are for ease of description only and are not intended to limit the order in which the steps are performed.

[0057] First Embodiment

[0058] For speech conversion in different languages, a separate speech conversion model is usually deployed for each language. However, in the case of mixed languages, deploying multiple models simultaneously leads to problems such as high system complexity and large resource consumption. Therefore, the embodiments of this application improve the original VITS architecture, realizing efficient processing of speech conversion of mixed Chinese and English text within a single model, and obtaining high-quality Chinese and English speech with natural pronunciation.

[0059] like Figure 1 As shown, to address the problem in existing technologies that bilingual text in Chinese and English consumes a lot of resources and produces low-quality output speech during the speech conversion process, the first embodiment of this application provides a detailed flowchart of a bilingual text-to-speech method based on an NPU. This embodiment's NPU-based bilingual text-to-speech method includes steps S1 to S5, detailed below:

[0060] Step S1: Based on the language type of each segment in the mixed Chinese and English text, perform word segmentation and phoneme conversion on the mixed Chinese and English text to obtain the text input sequence.

[0061] In this embodiment, a VITS model deployed on an NPU is used for text-to-speech conversion. An NPU, or Neural Processing Unit, is a new type of hardware chip specifically designed to accelerate artificial intelligence and deep learning tasks. The VITS model is an end-to-end speech synthesis model that combines variational inference, normalized flow, and adversarial training, and can improve speech diversity and naturalness through latent variable modeling and stochastic duration prediction.

[0062] In order to achieve natural Chinese-English bilingual and transcoded speech synthesis, the embodiments of this application have made targeted designs and optimizations to the key components of the text processing front-end, text encoder and acoustic feature encoder on the original VITS architecture.

[0063] For the text processing front-end, the input mixed Chinese and English text undergoes text regularization, language identification, word segmentation, phoneme conversion, unified phoneme representation, serialization, and padding sequentially. Text regularization cleanses and standardizes the input mixed Chinese and English text, including converting full-width characters to half-width characters and standardizing the representation format of numbers, dates, and symbols. Language identification identifies the language attributes of each segment of the text, recognizing both Chinese and English.

[0064] Optionally, embodiments of this application employ fine-grained language recognition methods based on dictionary analysis or machine learning models to achieve language identification within sentences.

[0065] Word segmentation involves implementing different segmentation strategies for different languages. Chinese text uses segmenters suitable for Chinese processing, such as segmentation methods based on dictionaries or pre-trained models; English text is usually segmented based on spaces and punctuation marks, and fixed phrases are treated as a whole for processing.

[0066] After word segmentation, text units composed of multiple Chinese characters or English words are obtained. These text units are then converted into corresponding phoneme sequences according to language type. Specifically, Chinese text units are converted into Pinyin sequences, and the phonemes in the Pinyin sequences are mapped to their corresponding initials and finals by taking tone information into account. English text units are converted into English phoneme sequences based on existing pronunciation dictionaries.

[0067] Unified phoneme representation associates each phoneme in the converted phoneme sequence with corresponding speech information, thus labeling the phonemes so that the model can learn the pronunciation patterns of Chinese or English during subsequent speech conversion. In other embodiments, unified phoneme representation is achieved by mapping Chinese and English phonemes to a shared phoneme set.

[0068] The phoneme sequence associated with the speech information is converted into a data ID sequence, resulting in a phoneme ID sequence and a language ID sequence. The ID sequences are then padded to a fixed maximum length according to the NPU's input length requirements. The language ID sequence identifies the speech type of each phoneme, while the phoneme ID sequence provides the specific content of the speech. The language ID sequence guides the corresponding pronunciation style in subsequent speech conversion, while the phoneme ID sequence specifies the specific pronunciation content.

[0069] Generate a text input sequence based on the phoneme ID sequence and the language ID sequence.

[0070] Step S2 involves combining the phoneme ID sequence and the language ID sequence into vectors, and extracting textual semantic features from the vector combination results to obtain textual latent variables.

[0071] The phoneme ID sequence and language ID sequence are mapped to phoneme embedding vectors and language embedding vectors, respectively, and the position information is encoded in the embedding vectors during the mapping process.

[0072] Specifically, the positional information of each phoneme in the phoneme ID sequence is obtained. This positional information is itself a mathematical vector. This positional information is encoded in the phoneme embedding vector and the language embedding vector. Then, based on the positional encoding result, the phoneme embedding vector and the language embedding vector are combined to generate a hybrid embedding vector that includes content, order, and language style. The combination methods include concatenation and addition.

[0073] As an improvement to the above scheme, the VITS model can also introduce multiple speaker IDs in the context of mixed Chinese and English text and map them into speaker embedding vectors, which can be used as one of the data for subsequent vector combination.

[0074] The hybrid embedding vectors are input into the multi-layer Transformer Encoder structure of the VITS model. This structure, based on a pre-defined self-attention mechanism, can effectively capture long-distance dependencies within the input phoneme sequence, thereby learning context-dependent text features and obtaining phoneme-level textual latent variables that represent the semantic information of the text. These textual latent variables are abstract representations that are not directly observed during text processing but are implicitly contained within the model. They are typically used to capture the latent structural, semantic, or stylistic features of text and serve as a bridge for text-to-speech conversion.

[0075] Step S3: Based on the textual latent variables, predict the duration of each phoneme and its prior distribution parameters in the acoustic latent feature space.

[0076] In this embodiment, a lightweight network is used after a multi-layer Transformer Encoder structure. This lightweight network processes the text's latent variables to predict the pronunciation duration of each phoneme in both Chinese and English environments. The logarithm of each pronunciation duration is then taken to obtain the overall duration of each phoneme. This lightweight network is trained to accurately predict the typical pronunciation duration of phonemes in different languages ​​and can smoothly handle the duration transitions in the conversion process.

[0077] Optionally, embodiments of this application may employ a small feedforward network or a convolutional network as a lightweight network.

[0078] Following the multi-layer Transformer Encoder structure is a prior encoder, which models the prior distribution of the textual latent variables mapped to the acoustic feature latent space, and outputs the mean and log-variance of this Gaussian prior distribution to obtain the prior distribution parameters. These prior distribution parameters can be used for subsequent decoding to obtain the corresponding speech sequence.

[0079] Step S4 involves performing several levels of transformation on the phoneme alignment results based on the duration and the prior distribution parameters to obtain a bilingual acoustic feature sequence.

[0080] In the VITS model, a bilingual acoustic feature generator generates high-quality and rhythmically natural Chinese and English acoustic features, providing data support for subsequent conversion into speech waveforms.

[0081] In the bilingual acoustic feature generator, an attention alignment mechanism (or equivalent alignment information in other embodiments) is first generated based on the duration of the phonemes obtained above. Based on this attention alignment mechanism, the text latent variables are aligned temporally with the acoustic feature frames corresponding to the durations, resulting in phoneme alignment. This alignment process essentially extends the text latent variables, ensuring that each acoustic feature frame can be associated with corresponding text feature information. This achieves the construction of a non-linear alignment relationship between text latent variables and speech latent variables.

[0082] In addition, the dynamic calculation part of the attention alignment mechanism can be processed in a separate submodule provided by the NPU.

[0083] Then, latent variables are obtained by sampling based on prior distribution parameters. Because the length of the sampling results may be inconsistent, in order to adapt to the NPU inference process and ensure the accuracy of the results, a preset random noise tensor with a fixed length of random numbers is used for random sampling.

[0084] The core decoding module (Flow-based Decoder) is one of the core components of the bilingual acoustic feature generator, consisting of a series of reversible normalized flow modules stacked together. The core decoding module takes the obtained phoneme alignment results and latent variables as input, and through multi-level flow transformations, uses affine transformations to progressively and reversibly transform the latent variables into a complex distribution of target acoustic features. During each affine transformation, the transformation results are adjusted based on the phoneme alignment results, enabling the model to generate fine acoustic details according to the specific content of the input text (including language type, phonemes, prosody, etc.), resulting in a high-precision bilingual acoustic feature sequence.

[0085] Furthermore, the core decoding module can learn and reproduce the acoustic characteristics of Chinese and English respectively, and achieve a smooth transition of acoustic parameters during language switching (code conversion) to obtain a smooth acoustic feature.

[0086] Finally, the bilingual acoustic feature sequence is output as a Mel spectrum, the length of which corresponds to the total number of predicted frames, and is processed to a fixed maximum frame length during NPU deployment.

[0087] The obtained bilingual acoustic feature sequence demonstrates how abstract textual language information is converted into concrete physical sound descriptions. This conversion process is completely unified; it does not involve generating Chinese and English sounds separately and then splicing them together. Instead, it utilizes a single model to directly generate a complete acoustic sequence containing both languages ​​with a smooth transition, significantly reducing the resources required for text-to-speech conversion.

[0088] Step S5: Convert the bilingual acoustic feature sequence into the corresponding speech waveform.

[0089] The received Mel spectrum containing bilingual acoustic feature sequences is converted into a speech waveform that sounds natural to the ear using a vocoder, generating high-quality and rhythmically natural Chinese and English speech.

[0090] Furthermore, because NPUs require inputs of fixed shapes, while the VITS model used in this embodiment involves many dynamic and variable intermediate parameters during text conversion, the following optimization strategies are adopted in order to enable the VITS model to be efficiently deployed on the NPU, aiming to solve problems such as dynamic shape operations and NPU-incompatible operators in the original model:

[0091] (1) Establishing independent submodules. Because the original VITS model includes operations in its core forward propagation process that cause the shape of the intermediate result tensor to depend on the actual length of the input text and the predicted speech duration. For example, calculating the attention alignment mechanism and the output padding mask based on the predicted phoneme duration generates tensors with dynamic shapes, which is inconsistent with the requirement of most NPUs to have fixed input / output shapes. Therefore, this embodiment logically splits the acoustic model part of the original VITS model (mainly referring to the process from text encoding to acoustic feature generation) into multiple independent submodules that can be executed or deployed independently. This split allows for the isolation of parts containing dynamic shape calculations, or operations sensitive to dynamic shapes, for special processing.

[0092] The independent submodules can run on a CPU or DSP to handle dynamic calculations that are difficult to implement efficiently on a fixed-shape NPU, including predicting the duration of phoneme utterances, generating attention mechanisms, and masking the decoder output. The independent submodules ultimately trim or pad the output to the fixed shape required by the NPU.

[0093] (2) Introducing a random noise tensor. In the VITS model, randomness is often introduced at certain stages to improve the naturalness and diversity of the generated speech. This randomness means that even if the input is the same, the output may be different each time the model runs. This is detrimental to model derivation that requires deterministic behavior and to reproducible and verifiable inference on the NPU. Therefore, this application introduces a pre-generated random noise tensor of fixed length. By loading and directly using the random noise tensor, the determinism of the model throughout the inference process can be ensured.

[0094] (3) Replacing specific operators. Because some operators cannot be fully supported or optimized by the NPU conversion toolchain, direct conversion may fail, or the converted model may have poor performance, severely reduced accuracy, or even generate dynamic shapes or control flows that the NPU cannot handle. Therefore, this application identifies these NPU-incompatible operators and replaces them with functionally equivalent combinations of operations that are easier for the NPU toolchain to understand and execute efficiently. Among these replacement operations are replacing comparison operators, replacing cumulative summation operations, and replacing index assignment operations.

[0095] Implementing the embodiments of this application has the following beneficial effects:

[0096] This application first converts mixed Chinese and English text into a phoneme sequence, performing different word segmentation processes on Chinese and English during the conversion. This yields both phoneme ID sequences and language ID sequences that guide pronunciation style, supporting subsequent speech waveform encoding based on the acoustic features of each language. Then, latent text variables representing deep semantic information (prosodic features) are extracted from the mixed Chinese and English phoneme input sequence. The duration of each phoneme is predicted, and the text features are mapped to the speech latent space, yielding corresponding prior distribution parameters. This provides data for subsequent conversion into complex acoustic feature distributions. Phoneme alignment associates each acoustic feature frame with its corresponding text feature, aligning the duration of text and speech. The alignment results are then adjusted based on the prior distribution parameters to obtain a bilingual acoustic feature sequence containing precise acoustic details. This allows for accurate conversion of mixed bilingual text into high-quality speech data using only a single model.

[0097] Second Embodiment

[0098] Furthermore, in order to implement the NPU-based Chinese-English bilingual text-to-speech system corresponding to the above method embodiments, and to achieve the corresponding functions and technical effects, Figure 2 A structural diagram of an NPU-based Chinese-English bilingual text-to-speech system is provided. For ease of explanation, only the parts relevant to this embodiment are shown. The NPU-based Chinese-English bilingual text-to-speech system provided in this application embodiment includes:

[0099] The text unification module 201 is used to perform word segmentation and phoneme conversion on the mixed Chinese and English text based on the language type of each segment in the mixed Chinese and English text to obtain a text input sequence; wherein, the text input sequence includes a phoneme ID sequence and a language ID sequence.

[0100] In this embodiment, a VITS model deployed on an NPU is used for text-to-speech conversion. An NPU, or Neural Processing Unit, is a new type of hardware chip specifically designed to accelerate artificial intelligence and deep learning tasks. The VITS model is an end-to-end speech synthesis model that combines variational inference, normalized flow, and adversarial training, and can improve speech diversity and naturalness through latent variable modeling and stochastic duration prediction.

[0101] In order to achieve natural Chinese-English bilingual and transcoded speech synthesis, the embodiments of this application have made targeted designs and optimizations to the key components of the text processing front-end, text encoder and acoustic feature encoder on the original VITS architecture.

[0102] For the text processing front-end, the input mixed Chinese and English text undergoes text regularization, language identification, word segmentation, phoneme conversion, unified phoneme representation, serialization, and padding sequentially. Text regularization cleanses and standardizes the input mixed Chinese and English text, including converting full-width characters to half-width characters and standardizing the representation format of numbers, dates, and symbols. Language identification identifies the language attributes of each segment of the text, recognizing both Chinese and English.

[0103] Optionally, embodiments of this application employ fine-grained language recognition methods based on dictionary analysis or machine learning models to achieve language identification within sentences.

[0104] Word segmentation involves implementing different segmentation strategies for different languages. Chinese text uses segmenters suitable for Chinese processing, such as segmentation methods based on dictionaries or pre-trained models; English text is usually segmented based on spaces and punctuation marks, and fixed phrases are treated as a whole for processing.

[0105] After word segmentation, text units composed of multiple Chinese characters or English words are obtained. These text units are then converted into corresponding phoneme sequences according to language type. Specifically, Chinese text units are converted into Pinyin sequences, and the phonemes in the Pinyin sequences are mapped to their corresponding initials and finals by taking tone information into account. English text units are converted into English phoneme sequences based on existing pronunciation dictionaries.

[0106] Unified phoneme representation associates each phoneme in the converted phoneme sequence with corresponding speech information, thus labeling the phonemes so that the model can learn the pronunciation patterns of Chinese or English during subsequent speech conversion. In other embodiments, unified phoneme representation is achieved by mapping Chinese and English phonemes to a shared phoneme set.

[0107] The phoneme sequence associated with the speech information is converted into a data ID sequence, resulting in a phoneme ID sequence and a language ID sequence. The ID sequences are then padded to a fixed maximum length according to the NPU's input length requirements. The language ID sequence identifies the speech type of each phoneme, while the phoneme ID sequence provides the specific content of the speech. The language ID sequence guides the corresponding pronunciation style in subsequent speech conversion, while the phoneme ID sequence specifies the specific pronunciation content.

[0108] Generate a text input sequence based on the phoneme ID sequence and the language ID sequence.

[0109] The text feature extraction module 202 is used to perform vector combination on the phoneme ID sequence and the language ID sequence, and to extract text semantic features from the vector combination result to obtain text latent variables.

[0110] In this embodiment of the application, the phoneme ID sequence and the language ID sequence are mapped to phoneme embedding vector and language embedding vector, respectively, and the position information is encoded in the embedding vector during the mapping process.

[0111] Specifically, the positional information of each phoneme in the phoneme ID sequence is obtained. This positional information is itself a mathematical vector. This positional information is encoded in the phoneme embedding vector and the language embedding vector. Then, based on the positional encoding result, the phoneme embedding vector and the language embedding vector are combined to generate a hybrid embedding vector that includes content, order, and language style. The combination methods include concatenation and addition.

[0112] As an improvement to the above scheme, the VITS model can also introduce multiple speaker IDs in the context of mixed Chinese and English text and map them into speaker embedding vectors, which can be used as one of the data for subsequent vector combination.

[0113] The hybrid embedding vectors are input into the multi-layer Transformer Encoder structure of the VITS model. This structure, based on a pre-defined self-attention mechanism, can effectively capture long-distance dependencies within the input phoneme sequence, thereby learning context-dependent text features and obtaining phoneme-level textual latent variables that represent the semantic information of the text. These textual latent variables are abstract representations that are not directly observed during text processing but are implicitly contained within the model. They are typically used to capture the latent structural, semantic, or stylistic features of text and serve as a bridge for text-to-speech conversion.

[0114] The data encoding module 203 is used to predict the duration of each phoneme and its prior distribution parameters in the acoustic latent feature space based on the text latent variables.

[0115] In this embodiment, a lightweight network is selected and connected after a multi-layer Transformer Encoder structure. This lightweight network processes the latent variables of the text, predicts the pronunciation duration of each phoneme in both Chinese and English environments, and takes the logarithm of the pronunciation duration to obtain the overall duration of each phoneme. The lightweight network is trained to accurately predict the typical pronunciation duration of phonemes in different languages ​​and can smoothly handle the duration transitions in the conversion process.

[0116] Optionally, embodiments of this application may employ a small feedforward network or a convolutional network as a lightweight network.

[0117] Following the multi-layer Transformer Encoder structure is a prior encoder, which models the prior distribution of the textual latent variables mapped to the acoustic feature latent space, and outputs the mean and log-variance of this Gaussian prior distribution to obtain the prior distribution parameters. These prior distribution parameters can be used for subsequent decoding to obtain the corresponding speech sequence.

[0118] The acoustic feature sequence acquisition module 204 is used to perform several levels of transformation on the phoneme alignment result based on the duration and the prior distribution parameters to obtain a bilingual acoustic feature sequence.

[0119] In the VITS model, a bilingual acoustic feature generator generates high-quality and rhythmically natural Chinese and English acoustic features, providing data support for subsequent conversion into speech waveforms.

[0120] In the bilingual acoustic feature generator, an attention alignment mechanism (or equivalent alignment information in other embodiments) is first generated based on the duration of the phonemes obtained above. Based on this attention alignment mechanism, the text latent variables are aligned temporally with the acoustic feature frames corresponding to the durations, resulting in phoneme alignment. This alignment process essentially extends the text latent variables, ensuring that each acoustic feature frame can be associated with corresponding text feature information. This achieves the construction of a non-linear alignment relationship between text latent variables and speech latent variables.

[0121] In addition, the dynamic calculation part of the attention alignment mechanism can be processed in a separate submodule provided by the NPU.

[0122] Then, latent variables are obtained by sampling based on prior distribution parameters. Because the length of the sampling results may be inconsistent, in order to adapt to the NPU inference process and ensure the accuracy of the results, a preset random noise tensor with a fixed length of random numbers is used for random sampling.

[0123] The core decoding module (Flow-based Decoder) is one of the core components of the bilingual acoustic feature generator, consisting of a series of reversible normalized flow modules stacked together. The core decoding module takes the obtained phoneme alignment results and latent variables as input, and through multi-level flow transformations, uses affine transformations to progressively and reversibly transform the latent variables into a complex distribution of target acoustic features. During each affine transformation, the transformation results are adjusted based on the phoneme alignment results, enabling the model to generate fine acoustic details according to the specific content of the input text (including language type, phonemes, prosody, etc.), resulting in a high-precision bilingual acoustic feature sequence.

[0124] Furthermore, the core decoding module can learn and reproduce the acoustic characteristics of Chinese and English respectively, and achieve a smooth transition of acoustic parameters during language switching (code conversion) to obtain a smooth acoustic feature.

[0125] Finally, the bilingual acoustic feature sequence is output as a Mel spectrum, the length of which corresponds to the total number of predicted frames, and is processed to a fixed maximum frame length during NPU deployment.

[0126] The obtained bilingual acoustic feature sequence demonstrates how abstract textual language information is converted into concrete physical sound descriptions. This conversion process is completely unified; it does not involve generating Chinese and English sounds separately and then splicing them together. Instead, it utilizes a single model to directly generate a complete acoustic sequence containing both languages ​​with a smooth transition, significantly reducing the resources required for text-to-speech conversion.

[0127] The speech conversion module 205 is used to convert the bilingual acoustic feature sequence into the corresponding speech waveform.

[0128] In this embodiment of the application, the received Mel spectrum containing bilingual acoustic feature sequences is converted into a speech waveform that sounds natural to the ear using a vocoder, thereby generating high-quality and rhythmically natural Chinese and English speech.

[0129] Implementing the embodiments of this application has the following beneficial effects:

[0130] This application first converts mixed Chinese and English text into a phoneme sequence, performing different word segmentation processes on Chinese and English during the conversion. This yields both phoneme ID sequences and language ID sequences that guide pronunciation style, supporting subsequent speech waveform encoding based on the acoustic features of each language. Then, latent text variables representing deep semantic information (prosodic features) are extracted from the mixed Chinese and English phoneme input sequence. The duration of each phoneme is predicted, and the text features are mapped to the speech latent space, yielding corresponding prior distribution parameters. This provides data for subsequent conversion into complex acoustic feature distributions. Phoneme alignment associates each acoustic feature frame with its corresponding text feature, aligning the duration of text and speech. The alignment results are then adjusted based on the prior distribution parameters to obtain a bilingual acoustic feature sequence containing precise acoustic details. This allows for accurate conversion of mixed bilingual text into high-quality speech data using only a single model.

[0131] Furthermore, Figure 3 This is a structural diagram of a terminal device provided in one embodiment of this application. Figure 3 As shown, the terminal device 3 of this embodiment includes: at least one processor 30 (in... Figure 3The present invention includes only one of the following: a memory 31 and a computer program 32 stored in the memory 31 and executable on the at least one processor. When the processor 30 executes the computer program 32, it can implement the steps of the NPU-based Chinese-English bilingual text-to-speech method according to any one of the embodiments of this application.

[0132] The terminal device 3 may be a computing device such as a desktop computer, a cloud server, or a laptop computer, and the computing device may include, but is not limited to, a processor 30 and a memory 31. Figure 3 This is merely an example of terminal device 3 and does not constitute a limitation on terminal device 3. It may include more or fewer components than those shown in the figure.

[0133] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, or improvements made by those skilled in the art within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A Chinese-English bilingual text-to-speech method based on an NPU, applied to a VITS model deployed on an NPU, characterized in that... include: Based on the language type of each segment in the mixed Chinese-English text, word segmentation and phoneme conversion are performed on the mixed Chinese-English text to obtain a text input sequence; wherein, the text input sequence includes a phoneme ID sequence and a language ID sequence; The phoneme ID sequence and language ID sequence are combined into vectors, and textual semantic features are extracted from the vector combination results to obtain textual latent variables; Based on textual latent variables, predict the duration of each phoneme and its prior distribution parameters in the acoustic latent feature space; The phoneme alignment results based on the duration and the prior distribution parameters are transformed at several levels to obtain the bilingual acoustic feature sequence. The bilingual acoustic feature sequence is converted into the corresponding speech waveform.

2. The NPU-based Chinese-English bilingual text-to-speech method according to claim 1, characterized in that, Based on the language types of each segment in the mixed Chinese-English text, word segmentation and phoneme conversion are performed on the mixed Chinese-English text to obtain the text input sequence, specifically as follows: The language type of the mixed Chinese and English text is identified in segments, and the mixed Chinese and English text is segmented into several text units based on the identification results. Based on the language type of the text unit, the text unit is converted into a phoneme sequence; The Chinese and English phonemes in the phoneme sequence are associated with the corresponding speech information, and the associated phoneme sequence is converted into a data ID sequence to obtain a phoneme ID sequence and a language ID sequence.

3. The NPU-based Chinese-English bilingual text-to-speech method according to claim 2, characterized in that, The step of converting the text unit into a phoneme sequence according to the language type of the text unit specifically involves: The text units belonging to Chinese are converted into Pinyin sequences, and the phonemes in the Pinyin sequences are mapped to the corresponding initials and finals by taking into account tone information; The text units belonging to English are converted into English phoneme sequences.

4. The NPU-based Chinese-English bilingual text-to-speech method according to claim 1, characterized in that, The process of combining phoneme ID sequences and language ID sequences into vectors, and extracting textual semantic features from the vector combination results to obtain textual latent variables, specifically involves: Map the phoneme ID sequence and the language ID sequence to phoneme embedding vectors and language embedding vectors, respectively; During the mapping process, the phoneme ID sequence is positionally encoded according to the position of each phoneme in the phoneme ID sequence. Based on the positional encoding results, the phoneme embedding vector and the language embedding vector are combined to generate a hybrid embedding vector; Textual semantic features are extracted by learning the contextual relationships in the hybrid embedding vectors, thus obtaining textual latent variables.

5. The NPU-based Chinese-English bilingual text-to-speech method according to claim 1, characterized in that, The process of predicting the duration of each phoneme and its prior distribution parameters in the acoustic feature latent space based on textual latent variables specifically involves: Based on the textual latent variables, the pronunciation duration of each phoneme is predicted in both Chinese and English environments to obtain the duration; wherein, the logarithm of the pronunciation duration is taken. The prior distribution of textual latent variables mapped to the acoustic feature latent space is modeled to obtain the mean and logarithmic variance of the prior distribution parameters.

6. The NPU-based Chinese-English bilingual text-to-speech method according to claim 1, characterized in that, The process of performing several levels of transformation based on the phoneme alignment results of the duration and the prior distribution parameters to obtain the bilingual acoustic feature sequence is as follows: Align the acoustic feature frames corresponding to the duration with the text latent variables to obtain the phoneme alignment result; Based on a preset random noise tensor, latent variables are obtained by sampling on the prior distribution parameters; wherein, the random noise tensor is a random number of fixed length; The latent variables are subjected to several levels of affine transformation, and the transformation results are adjusted according to the phoneme alignment results during each level of affine transformation to obtain a bilingual acoustic feature sequence; wherein the affine transformation is reversible.

7. The NPU-based Chinese-English bilingual text-to-speech method according to claim 6, characterized in that, The step of aligning the acoustic feature frames corresponding to the duration with the text latent variables to obtain the phoneme alignment result is specifically as follows: The acoustic feature distribution of each phoneme is determined based on the number of acoustic feature frames corresponding to the duration. Based on the acoustic feature distribution, by extending the text latent variables in the time dimension, each acoustic feature frame is aligned with the text semantic features to obtain the phoneme alignment result.

8. The NPU-based Chinese-English bilingual text-to-speech method according to any one of claims 1 to 7, characterized in that, The VITS model is specifically as follows: The duration of each phoneme is predicted based on the independent sub-modules provided by the VITS model; wherein, the independent sub-modules can be executed or deployed independently and run on a CPU or DSP; When the VITS model performs random sampling, it loads and directly uses the random noise tensor. In the process of converting mixed Chinese and English text into a bilingual acoustic feature sequence, the first operator that cannot be converted by the NPU is identified, and data replacement is performed on the first operator; wherein, the data replacement includes replacing comparison operators, replacing cumulative summation operations, and replacing index assignment operations.

9. A Chinese-English bilingual text-to-speech system based on an NPU, characterized in that, It includes: a text unification module, a text feature extraction module, a data encoding module, an acoustic feature sequence acquisition module, and a speech conversion module; The text unification module is used to perform word segmentation and phoneme conversion on the mixed Chinese and English text based on the language type of each segment, to obtain a text input sequence; wherein, the text input sequence includes a phoneme ID sequence and a language ID sequence; The text feature extraction module is used to combine phoneme ID sequences and language ID sequences into vectors, and extract text semantic features from the vector combination results to obtain text latent variables; The data encoding module is used to predict the duration of each phoneme and its prior distribution parameters in the acoustic latent feature space based on text latent variables. The acoustic feature sequence acquisition module is used to perform several levels of transformation on the phoneme alignment result based on the duration and the prior distribution parameters to obtain a bilingual acoustic feature sequence. The speech conversion module is used to convert the bilingual acoustic feature sequence into the corresponding speech waveform.

10. A terminal device, characterized in that, It includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the steps of the NPU-based Chinese-English bilingual text-to-speech method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Chinese and English mixed speech synthesis method and device

    CN112151005A

  • Speech synthesis method, device and equipment and computer readable storage medium

    CN112365878A

  • Multi-speaker and multi-language speech synthesis method and system thereof

    CN112435650A

  • Chinese and English mixed speech synthesis method and device, electronic equipment and storage medium

    CN113380221A

  • Parallel speech synthesis method and device based on variational auto-encoder

    CN113450761A

Cited By

  • Text feature extraction method and system

    CN121366567A