Speech synthesis method, speech synthesis device, electronic equipment and storage medium
By utilizing sample speech sets and adjusting the parameters of the semantic modeler in text-to-speech technology, the problem of low accuracy in cross-language speech synthesis is solved, achieving more efficient cross-language speech synthesis results.
Patent Information
- Application Number
- CN202610012779.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-06
- Publication Date
- 2026-02-10
AI Technical Summary
Existing text-to-speech technologies suffer from low accuracy in cross-language speech synthesis, especially when the amount of text in a particular language in the text dataset is small, which affects the accuracy of speech synthesis.
By acquiring sample speech sets and sample texts, semantic modeling is performed using an initial semantic modeler. Combined with speech quantization and parameter adjustment, a target semantic modeler is obtained for cross-language speech synthesis, improving the cross-language capability of semantic modeling. The target semantic modeler is then used to reconstruct speech from the target text.
It improves the accuracy of cross-language speech synthesis, overcomes the adverse effects caused by the uneven distribution of text volume in different languages, and enhances the effect of cross-language speech synthesis.
Smart Images

Figure CN121506097A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and is applicable to the financial and medical fields. In particular, it relates to a speech synthesis method, speech synthesis device, electronic device, and storage medium. Background Technology
[0002] Currently, Text-to-Speech (TTS) technology aims to synthesize speech from text. For example, in intelligent customer service scenarios in the financial sector, bank / securities company customer service systems convert text information (such as account notifications, transaction confirmations, and interest rate change announcements) into speech in real time, thereby providing customers with voice self-service. Another example is in medical self-service navigation scenarios in the healthcare sector, where hospital self-service guidance / information kiosks convert department locations, treatment procedures, and pharmacy information into voice announcements, improving the accessibility of medical care.
[0003] In related technologies, TTS models are typically trained using text datasets. These datasets can include text in multiple languages (such as English and Chinese), enabling the TTS model to perform cross-language speech synthesis. However, if the amount of text in a particular language in the dataset is limited, the TTS model may not adequately support that language, affecting the accuracy of speech synthesis.
[0004] Therefore, the relevant technologies suffer from low accuracy in cross-language speech synthesis. Summary of the Invention
[0005] The main objective of this application is to propose a speech synthesis method, speech synthesis device, electronic device, and storage medium that can reduce the adverse effects of uneven distribution of text volume in different languages on cross-language speech synthesis and improve the accuracy of cross-language speech synthesis.
[0006] To achieve the above objectives, a first aspect of this application proposes a speech synthesis method, the method comprising: Obtain the target text and determine the target language to which the target text belongs; Obtain a sample speech set and sample text belonging to the target language; wherein, the sample speech set includes a first sample speech belonging to the target language and a second sample speech belonging to a non-target language, and the speech content of each first sample speech and each second sample speech is the same and is the sample text; The sample text and the target language are semantically modeled using a preset initial semantic modeler to obtain a first sample semantic unit sequence; Speech quantization is performed on the sample speech in the sample speech set to obtain a second sample semantic unit sequence; The initial semantic modeler is adjusted according to the first sample semantic unit sequence and the second sample semantic unit sequence to obtain the target semantic modeler; The target semantic modeler is used to perform semantic modeling on the target text and the target language to obtain a sequence of text semantic units. Speech reconstruction is performed based on the sequence of semantic units in the text to obtain the target synthesized speech.
[0007] Optionally, adjusting the parameters of the initial semantic modeler based on the first sample semantic unit sequence and the second sample semantic unit sequence to obtain the target semantic modeler includes: The semantic loss data is obtained by calculating the loss based on the first sample semantic unit sequence and the second sample semantic unit sequence. The first sample semantic unit sequence is used to perform text conversion to obtain the first converted text, and the second sample semantic unit sequence is used to perform text conversion to obtain the second converted text. Loss data is obtained by calculating the loss based on the first converted text and the second converted text; The initial semantic modeler is adjusted based on the semantic loss data and the text loss data to obtain the target semantic modeler.
[0008] Optionally, the step of performing speech quantization on the sample speech to obtain a second sample semantic unit sequence includes: Speech recognition is performed on each of the sample speech to obtain an initial semantic unit; Clustering the initial semantic units of each of the sample speech samples yields the second sample semantic unit sequence.
[0009] Optionally, the step of reconstructing speech based on the text semantic unit sequence to obtain the target synthesized speech includes: The target Mel spectrum is obtained by performing flow matching on the text semantic unit sequence using a preset target condition flow matching model. The target Mel spectrum is converted into speech to obtain the target synthesized speech.
[0010] Optionally, before performing flow matching on the text semantic unit sequence using a preset target condition flow matching model to obtain the target Mel spectrum, the method further includes: The target semantic modeler performs semantic modeling on the sample text and the target language to obtain a sequence of target sample semantic units. The sample noise is determined, and the sample noise and the target sample semantic unit sequence are flow matched using an initial conditional flow matching model to obtain the sample prediction vector field; The auxiliary vector field is obtained by calculating the difference between the sample noise and the reference Mel spectrum corresponding to the sample speech; Loss calculation is performed based on the sample prediction vector field and the auxiliary vector field to obtain flow matching loss data; The parameters of the initial conditional flow matching model are adjusted based on the flow matching loss data to obtain the target conditional flow matching model.
[0011] Optionally, the step of performing flow matching on the text semantic unit sequence using a preset target condition flow matching model to obtain the target Mel spectrum includes: Obtain reference speech data that does not belong to the target language; Speaker coding is performed on the reference speech data to obtain speaker features; The target noise is determined, and the target noise, the text semantic unit sequence, and the speaker features are matched using the target conditional flow matching model to obtain the target Mel spectrum.
[0012] Optionally, the step of calculating the loss based on the sample prediction vector field and the auxiliary vector field to obtain the flow matching loss data includes: The difference between the sample prediction vector field and the auxiliary vector field is calculated to obtain the vector field difference. The L2 norm is calculated based on the vector field difference to obtain the flow matching loss data.
[0013] To achieve the above objectives, a second aspect of this application provides a speech synthesis apparatus, the apparatus comprising: The text acquisition module is used to acquire target text and determine the target language to which the target text belongs; The sample acquisition module is used to acquire a sample speech set and sample text belonging to the target language; wherein, the sample speech set includes a first sample speech belonging to the target language and a second sample speech belonging to a non-target language, and the speech content of each first sample speech and each second sample speech is the same and is the sample text; The first semantic modeling module is used to perform semantic modeling on the sample text and the target language through a preset initial semantic modeler to obtain a first sample semantic unit sequence; The speech quantization module is used to perform speech quantization on the sample speech to obtain a second sample semantic unit sequence; The parameter adjustment module is used to adjust the parameters of the initial semantic modeler according to the first sample semantic unit sequence and the second sample semantic unit sequence to obtain the target semantic modeler. The second semantic modeling module is used to perform semantic modeling on the target text and the target language through the target semantic modeler to obtain a sequence of text semantic units. The speech reconstruction module is used to reconstruct speech based on the text semantic unit sequence to obtain the target synthesized speech.
[0014] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the speech synthesis method described in the first aspect.
[0015] To achieve the above objectives, a fourth aspect of the present application provides a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the speech synthesis method described in the first aspect.
[0016] The speech synthesis method, speech synthesis device, electronic device, and storage medium proposed in this application train a target semantic modeler based on a sample speech set and sample text belonging to the target language. Since the sample speech set contains sample speech from multiple languages, and the speech content of each sample speech is sample text, the target semantic modeler can ignore the influence of language on semantic modeling, thereby improving the cross-lingual capability of semantic modeling. Then, the target semantic modeler is used to perform semantic modeling on the target text and target language, obtaining a sequence of text semantic units; finally, speech reconstruction is performed based on the text semantic unit sequence to obtain the target synthesized speech. This overcomes the adverse effects of the limited number of target language sample speech in the sample speech set on the cross-lingual capability of the target semantic modeler, i.e., it reduces the adverse effects of uneven distribution of text volume in different languages on cross-lingual speech synthesis, thereby improving the accuracy of cross-lingual speech synthesis.
[0017] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0018] Figure 1 This is a flowchart of the speech synthesis method provided in the embodiments of this application; Figure 2 yes Figure 1 The flowchart for step 104 in the document; Figure 3 yes Figure 1 The flowchart for step 105 in the document; Figure 4 yes Figure 1 The flowchart for step 107 in the document; Figure 5 yes Figure 4 The flowchart for step 401 in the document; Figure 6 This is a flowchart of a speech synthesis method provided in another embodiment of this application; Figure 7 This is a block diagram of the module structure of the speech synthesis device provided in the embodiments of this application; Figure 8 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0020] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0022] First, let's analyze some of the terms used in this application: Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0023] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). It is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information and image processing, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.
[0024] Zero-shot speech synthesis is a technique that can generate the speech of any speaker without specific speech data. This technique utilizes deep learning and neural networks, trained on datasets of multiple speakers, to generate outputs with a specific speaker's speech pattern after input text. Zero-shot speech synthesis does not require collecting independent speech data for each speaker, thus greatly improving synthesis efficiency and flexibility.
[0025] Current text-to-speech (TTS) systems have made significant progress in monolingual and partially multilingual environments, especially in English and Chinese, where models trained on large-scale data can synthesize highly natural and consistent speech. However, these technologies have significant shortcomings in supporting multiple languages, particularly low-resource languages. When performing speech synthesis, these technologies typically suffer from at least one of the following drawbacks: 1. Limited language coverage: Most existing TTS systems focus on English, Chinese, and a few European languages, lacking unified support for multilingual environments such as Hindi languages, thus limiting the technology's widespread applicability. 2. Weak cross-lingual generalization ability: Although some models possess zero-shot or few-shot cross-lingual capabilities, they often exhibit decreased sound quality, unnatural prosody, or incorrect speech content in low-resource languages. 3. Limitations in modeling methods: Current mainstream methods are mostly based on diffusion models or VQVAE for semantic and acoustic modeling. These methods are complex to train, slow inference, and struggle to maintain stability under multilingual conditions. 4. Severe data imbalance: Major languages (such as English) constitute the vast majority of training data, while there is insufficient speech data for low-resource languages (minor languages). The model tends to favor major languages, resulting in poor performance in minor languages. 5. Insufficient expressiveness: Traditional methods struggle to capture both semantic information (content, emotion, intonation) and acoustic information (timbre, prosody), easily leading to insufficient expressiveness or style loss in cross-linguistic or multi-speaker environments.
[0026] Therefore, existing technologies cannot simultaneously achieve multilingual coverage, cross-language generalization ability, synthesis efficiency, and naturalness within a unified framework, and a new method is urgently needed to address these shortcomings.
[0027] Based on this, embodiments of this application propose a speech synthesis method, a speech synthesis device, an electronic device, and a computer-readable storage medium, which can improve the accuracy of cross-language speech synthesis.
[0028] The speech synthesis method provided in this application can be applied to terminals and servers, or it can be software running on the server. The server can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or it can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the speech synthesis method, but is not limited to the above forms.
[0029] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include server computers, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0030] This application provides a speech synthesis method, a speech synthesis device, an electronic device, and a computer-readable storage medium, which are specifically described through the following embodiments. First, the speech synthesis method in the embodiments of this application is described.
[0031] It should be noted that in each specific embodiment of this application, when it is necessary to process data related to the user's identity or characteristics, such as the user's voice data, the user's permission or consent will be obtained first. Moreover, the collection, use and processing of this data will comply with relevant laws, regulations and standards.
[0032] Reference Figure 1 , Figure 1This is an optional flowchart of the speech synthesis method provided in the embodiments of this application, which may include, but is not limited to, steps 101 to 107.
[0033] Step 101: Obtain the target text and determine the target language to which the target text belongs; Step 102: Obtain a sample speech set and sample text belonging to the target language; wherein, the sample speech set includes a first sample speech belonging to the target language and a second sample speech belonging to a non-target language, and the speech content of each first sample speech and each second sample speech is the same and is both sample text. Step 103: Semantic modeling of the sample text and target language is performed using a preset initial semantic modeler to obtain the first sample semantic unit sequence; Step 104: Perform speech quantization on the sample speech in the sample speech set to obtain the second sample semantic unit sequence; Step 105: Adjust the parameters of the initial semantic modeler according to the first sample semantic unit sequence and the second sample semantic unit sequence to obtain the target semantic modeler; Step 106: Semantic modeling of the target text and target language is performed using a target semantic modeler to obtain a sequence of text semantic units; Step 107: Reconstruct speech based on the text semantic unit sequence to obtain the target synthesized speech.
[0034] Steps 101 to 107, as illustrated in this embodiment, train a target semantic modeler based on a sample speech set and sample text belonging to the target language. Since the sample speech set contains sample speech from multiple languages, and the speech content of each sample speech is sample text, the target semantic modeler can ignore the influence of language on semantic modeling, thereby improving the cross-lingual capability of semantic modeling. Then, the target semantic modeler is used to perform semantic modeling on the target text and target language, obtaining a text semantic unit sequence; subsequently, speech reconstruction is performed based on the text semantic unit sequence to obtain the target synthesized speech. This overcomes the adverse effect of the limited sample speech of the target language in the sample speech set on the cross-lingual capability of the target semantic modeler, i.e., it reduces the adverse effect of the uneven distribution of text volume in different languages on cross-lingual speech synthesis, thereby improving the accuracy of cross-lingual speech synthesis.
[0035] In step 101 of some embodiments, the target text is acquired, and the target language of the target text is determined. The target text is the text to be converted into speech. The target language is the language to which the target text belongs. For example, if the target text is generated based on English, then the target language is English. Or, for example, if the target text is generated based on Chinese, then the target language is Chinese.
[0036] In one example, within the context of intelligent customer service in the financial sector, the target text could be text information generated by a bank's / securities firm's customer service system, such as account notifications, transaction confirmations, or interest rate change announcements. For instance, the target text could be "Effective from [effective date], the benchmark interest rate will be adjusted from [original interest rate]% to [new interest rate]%".
[0037] In another example, in the medical self-service navigation scenario, the target text can be text information generated by hospital self-service guidance / information kiosks, such as department locations, consultation procedures, and medication dispensing point information. For example, the target text could be "Please go to consultation room 1 for your consultation, patient XX." Or, the target text could be "I am the intelligent medical assistant XX. If you want to go to consultation room 1, please walk forward 10 meters. The consultation room is on your left."
[0038] In another example, in a short video dubbing scenario, given the target text "This is a short video about how to brush your teeth correctly," a target synthesized speech can be generated based on the target text, English, and the voice of the voice actor. The speaker of the target synthesized speech is the voice actor, the target language is English, and the speech content is "This is a short video about how to brush your teeth correctly." The target audio and video are then merged based on the corresponding video frames of the target synthesized speech and the target text.
[0039] In step 102 of some embodiments, a sample speech set and sample text belonging to the target language are obtained. The sample speech set includes multiple sample speech samples. Specifically, the sample speech set includes first sample speech belonging to the target language and second sample speech belonging to a non-target language. The speech content of each first sample speech and each second sample speech is the same and both are sample text. There can be multiple second sample speech samples belonging to non-target languages, and any two second sample speech samples may belong to the same or different languages.
[0040] For example, if the target language is Hindi, then the non-target language is not Hindi, and may specifically include English and Chinese. The sample speech set may include at least one first sample speech belonging to Hindi, at least one second sample speech belonging to English, and at least one second sample speech belonging to Chinese.
[0041] Both the first and second sample audio recordings contain sample text. For example, if the sample text is "interest rate change announcement," then the sample audio set includes: sample audio in Hindi with the content "interest rate change announcement," sample audio in English with the content "interest rate change announcement," sample audio in Chinese with the content "interest rate change announcement," and so on.
[0042] In one embodiment, if two sample speech items are in the same language, then the speakers of the two sample speech items are different. For example, if there are multiple first sample speech items, then any two first sample speech items will have different speakers. Similarly, if there are multiple second sample speech items belonging to the same non-target language (e.g., English or Chinese), then the speakers of the two second sample speech items will be different. The speakers of two second sample speech items belonging to different non-target languages may be the same or different.
[0043] In step 103 of some embodiments, semantic modeling is performed on the sample text and the target language using a preset initial semantic modeler to obtain a first sample semantic unit sequence. The initial semantic modeler is used to convert text into semantic units so that speech can be subsequently constructed based on the semantic units. The initial semantic modeler can be a Gemma-based decoder Transformer. The first sample semantic unit sequence includes multiple first sample semantic units. Each first sample semantic unit can represent the text semantics at a corresponding position in the sample text.
[0044] In step 104 of some embodiments, the sample speech is quantized to obtain a second sample semantic unit sequence. The second sample semantic unit sequence includes multiple second sample semantic units. Each second sample semantic unit can represent the semantic content of the speech content at a corresponding position in the sample speech.
[0045] In one example, the VQVAE (Vector-Quantized Variational Autoencoder) method can be used to quantize speech into semantic units. The VQVAE method works by mapping the raw speech waveform or intermediate representation (such as Mel-spectral coefficients, acoustic features, latent vectors) to discrete vector representations (codebook vectors), which are then interpreted as semantic units.
[0046] In one embodiment, reference is made to Figure 2 Step 104 may include: Step 201: Perform speech recognition on each sample speech to obtain the initial semantic unit; Step 202: Cluster the initial semantic units of each sample speech to obtain the second sample semantic unit sequence.
[0047] In step 201, wav2vec2.0 can be used to perform speech recognition on each sample speech to obtain initial semantic units. wav2vec2.0 is an end-to-end self-supervised learning framework proposed by Facebook AI Research (FAIR) for learning representations from raw speech signals to improve the performance of downstream speech recognition (ASR) tasks, especially when labeled data is limited. wav2vec2.0 is pre-trained on massive cross-lingual corpora and possesses powerful cross-lingual semantic modeling capabilities. The working principle of wav2vec2.0 includes: a convolutional feature extractor (usually a multi-layer 1D / 2D convolution) encodes the raw waveform of the speech or its high-level features to obtain a latent representation sequence. A quantization module (Vector Quantization, VQ) discretizes the continuous latent representation sequence into discrete codebook vectors, forming the initial semantic units.
[0048] In step 202, the k-means clustering algorithm can be used to cluster the initial semantic units of each sample speech to obtain the second sample semantic unit sequence. For example, a stable semantic token dictionary (approximately 10k in size) can be obtained through k-means discretization, ensuring cross-language applicability.
[0049] The advantage of the embodiments of steps 201 to 202 above is that by using wav2vec2.0 + k-means clustering to extract semantic discrete representations, generalizable speech semantic units are obtained.
[0050] In step 105 of some embodiments, the parameters of the initial semantic modeler are adjusted based on the first sample semantic unit sequence and the second sample semantic unit sequence to obtain the target semantic modeler. The target semantic modeler is used to convert text into semantic units so that speech can be subsequently constructed based on these semantic units. The model structure of the target semantic modeler is the same as that of the initial semantic modeler, but their model parameters differ. The target semantic modeler can be a Gemma-based decoder (Transformer). The initial semantic modeler can be trained using language modeling (LM) to obtain the target semantic modeler.
[0051] In one embodiment, reference is made to Figure 3 Step 105 may include: Step 301: Calculate the loss based on the first sample semantic unit sequence and the second sample semantic unit sequence to obtain semantic loss data; Step 302: Perform text conversion based on the first sample semantic unit sequence to obtain the first converted text, and perform text conversion based on the second sample semantic unit sequence to obtain the second converted text; Step 303: Calculate the loss based on the first and second converted texts to obtain text loss data; Step 304: Adjust the parameters of the initial semantic modeler based on the semantic loss data and text loss data to obtain the target semantic modeler.
[0052] The advantage of the embodiments of steps 301 to 304 above is that the target semantic modeler is obtained by jointly training semantic loss data and text loss data, which improves the accuracy of semantic modeling.
[0053] In step 301, the cosine similarity distance calculation formula can be used to calculate the loss between the first sample semantic unit sequence and the second sample semantic unit sequence to obtain semantic loss data. The semantic loss data is used to characterize the direct difference between the two semantic unit sequences.
[0054] In step 302, a relevant text decoder can be set for text conversion to obtain a first converted text based on a first sample semantic unit sequence, and to obtain a second converted text based on a second sample semantic unit sequence.
[0055] In step 303, loss functions such as cross-entropy loss function or contrast loss function can be used to calculate the loss of the first transformed text and the second transformed text to obtain text loss data.
[0056] In step 304, the parameters of the initial semantic modeler can be adjusted based on the average of the semantic loss data and the text loss data to obtain the target semantic modeler.
[0057] In one embodiment, step 304 may include: obtaining semantic loss weights and text loss weights, wherein the semantic loss weights are greater than the text loss weights; multiplying the semantic loss weights and semantic loss data to obtain a first loss value; multiplying the text loss weights and text loss data to obtain a second loss value; adding the first loss value and the second loss value to obtain a target loss value; adjusting the parameters of the initial semantic modeler according to the target loss value until the target loss value is less than a preset loss threshold, and using the adjusted initial semantic modeler as the target semantic modeler.
[0058] The advantage of the above embodiments is that they not only enable the target semantic modeler to have cross-language prediction capabilities, but also introduce loss weights of different sizes (semantic loss weights are greater than text loss weights) during training, thus ensuring the core position of semantics.
[0059] In step 106 of some embodiments, a target semantic modeler performs semantic modeling on the target text and target language to obtain a sequence of text semantic units. The target semantic modeler has the ability to integrate information such as content, emotion, and intonation of speech. Specifically, step 106 may include: encoding the target text to obtain a target word feature sequence; encoding the target language to obtain language features; and performing semantic modeling on the target word feature sequence and language features using the target semantic modeler to obtain a sequence of text semantic units.
[0060] In one embodiment, step 106 may include: acquiring target speech data of the target speaker; performing speaker encoding on the target speech data to obtain target speaker features; and performing semantic modeling on the target speaker features, target word feature sequence, and language features using a target semantic modeler to obtain a text semantic unit sequence. In this way, speaker encoding is introduced during semantic modeling, further improving the generalization of semantic modeling.
[0061] In step 107 of some embodiments, speech reconstruction is performed based on the sequence of text semantic units to obtain the target synthesized speech. This step realizes the reconstruction from text semantics to speech.
[0062] In one embodiment, reference is made to Figure 4 Step 107 may include: Step 401: Perform flow matching on the text semantic unit sequence using a preset target condition flow matching model to obtain the target Mel spectrum; Step 402: Perform speech conversion on the target Mel spectrum to obtain the target synthesized speech.
[0063] In step 401, the idea behind flow matching is to find a gradual transformation process from a simple distribution (such as Gaussian noise) to a complex data distribution based on the approximate modeling and sampling of high-dimensional data distributions. This process allows samples that approximate real data to be obtained after several transformations, starting from initial noise sampling. The Conditional Flow Matching (CFM) model learns a parameterized vector field (or transformation rule) to gradually move the data distribution over time or iteration steps until a target distribution is reached. Compared to the diffusion model used in related technologies, the conditional flow matching model used in this application achieves efficient semantic-to-acoustic mapping, balancing speech generation quality and inference efficiency. The CFM model has the ability to integrate information such as speaker timbre, speech rate, and environmental features.
[0064] The meaning of a vector field: For each point x in the data space, define a vector v. θ (x,t) indicates the direction and speed of movement of the point at time t, and θ is a model parameter. Essentially, v θIt encodes the local "density gradient information" or "potential structural information" of the data distribution, so that the overall flow conforms to the geometry of the target distribution.
[0065] It should be noted how the "flow" from noise to data is achieved: (1) Continuous time perspective: Treating time t as a continuous variable, a gradual transformation process from noise distribution to data distribution is defined. At each time point t, an instantaneous "vector field" or transformation is defined to guide the evolution direction and rate of the sample. Through integration / numerical solution, the initial distribution p0(x) (noise) is gradually mapped to the approximate data distribution p. T (x). Discrete-time perspective: The process is divided into K discrete steps t=1,...,K. At each step, a parameterized transformation / update rule is applied to gradually move the samples from the current distribution to the next distribution. This type of discrete step often corresponds to a progressively training sequence of network modules (such as one or more neural network layers / modules).
[0066] In one embodiment, reference is made to Figure 5 Step 401 may include: Step 501: Obtain reference speech data that does not belong to the target language; Step 502: Speaker coding is performed on the reference speech data to obtain speaker features; Step 503: Determine the target noise, and perform flow matching on the target noise, text semantic unit sequence and speaker features through the target conditional flow matching model to obtain the target Mel spectrum.
[0067] Specifically, a speaker encoder can be introduced to encode the speaker in the reference speech data to obtain speaker features. A fixed-dimensional speaker embedding can be extracted from multiple reference audio segments and used as a conditional input to CFM. This embodiment achieves: (1) multi-speaker modeling: the model can generate speech with multiple timbres; cross-language transfer: (2) in low-resource languages, the naturalness and consistency of speech can be improved through existing speaker timbres transfer. Theoretically, this conditional modeling is based on the principles of transfer learning and multimodal alignment, enabling the model to have better cross-language generalization capabilities.
[0068] In step 402, the target Mel spectrum can be converted into high-fidelity audio using a BigVGAN vocoder, thus obtaining the target synthesized speech.
[0069] The advantage of the embodiments of steps 401 to 402 described above is that they enable efficient semantic-to-acoustic mapping while taking into account both speech generation quality and inference efficiency.
[0070] In one embodiment, reference is made to Figure 6 Before step 401, the speech synthesis method may further include: Step 601: Semantic modeling of the sample text and target language is performed using a target semantic modeler to obtain a sequence of target sample semantic units; Step 602: Determine the sample noise, and perform flow matching between the sample noise and the target sample semantic unit sequence through the initial conditional flow matching model to obtain the sample prediction vector field; Step 603: Calculate the difference between the sample noise and the reference Mel spectrum corresponding to the sample speech to obtain the auxiliary vector field; Step 604: Calculate the loss based on the sample prediction vector field and the auxiliary vector field to obtain the flow matching loss data; Step 605: Adjust the parameters of the initial conditional flow matching model based on the flow matching loss data to obtain the target conditional flow matching model.
[0071] In step 601, the sample text is first encoded to obtain a sample word feature sequence; the target language is encoded to obtain language features; and the target semantic modeler performs semantic modeling on the sample word feature sequence and language features to obtain a target sample semantic unit sequence.
[0072] In step 602, the sample prediction vector field can be represented as: , where v t (x) is the sample prediction vector field at time step t, x is the sample noise, y is the sequence of semantic units of the target sample, and ref is the speaker feature (which can be omitted).
[0073] In step 603, the auxiliary vector field can be represented as: x1-x0, where x0 is noise and x1 is the reference Mel spectrum (also known as the true Mel spectrum) corresponding to the sample speech. t (x) is used to guide the distribution from x0 to x1.
[0074] In step 604, the flow matching loss data is used to characterize the model performance of the initial conditional flow matching model. A larger flow matching loss indicates worse model performance, while a smaller flow matching loss indicates better model performance.
[0075] In one embodiment, step 604 may include: calculating the difference between the sample prediction vector field and the auxiliary vector field to obtain the vector field difference; and calculating the L2 norm based on the vector field difference to obtain the flow matching loss data.
[0076] Specifically, the flow matching loss data can be shown below: , Among them, L CFM For stream matching loss data, P t(x) represents the data distribution at time step t (also known as the Mel spectrum at time step t). This ensures that the Mel spectrum generated by the target conditional flow matching model conforms to the semantic conditions while maintaining consistency in speaker features.
[0077] The advantages of the embodiments of steps 601 to 605 above are: 1. High computational efficiency: Compared with the diffusion model which requires hundreds of iterations, CFM only requires fewer iterations to generate high-quality results; 2. Strong conditional modeling capability: CFM can explicitly introduce conditions (such as semantic unit sequences and speaker features) to ensure that the generated results are consistent with the conditions; 3. Strong theoretical interpretability: Stream matching establishes a continuous mapping between noise distribution and target distribution through the optimal transport (OT) theory, which is mathematically simpler.
[0078] Based on the above embodiments, this application can achieve at least the following beneficial effects: (1) It can be trained on a dataset of approximately 20,000 hours covering 22 languages, ensuring broad coverage of multiple languages. (2) It adopts data augmentation and cleaning mechanisms to ensure the diversity and quality of training data. (3) Large-scale multilingual training enables the model (which includes a target semantic modeler and a target conditional flow matching model) to learn shared representations between languages in the latent space, thereby achieving cross-lingual zero-shot synthesis.
[0079] Please see Figure 7 This application also provides a speech synthesis apparatus that can implement the above-described speech synthesis method. Figure 7The present application provides a block diagram of the module structure of a speech synthesis device, which includes: a text acquisition module 701, a sample acquisition module 702, a first semantic modeling module 703, a speech quantization module 704, a parameter adjustment module 705, a second semantic modeling module 706, and a speech reconstruction module 707. The system comprises the following modules: a text acquisition module 701, used to acquire target text and determine the target language to which the target text belongs; a sample acquisition module 702, used to acquire a sample speech set and sample text belonging to the target language; wherein the sample speech set includes a first sample speech belonging to the target language and a second sample speech belonging to a non-target language, and the speech content of each first sample speech and each second sample speech is the same and is both sample text; a first semantic modeling module 703, used to perform semantic modeling on the sample text and the target language using a preset initial semantic modeler to obtain a first sample semantic unit sequence; a speech quantization module 704, used to perform speech quantization on the sample speech to obtain a second sample semantic unit sequence; a parameter adjustment module 705, used to adjust the parameters of the initial semantic modeler according to the first sample semantic unit sequence and the second sample semantic unit sequence to obtain a target semantic modeler; a second semantic modeling module 706, used to perform semantic modeling on the target text and the target language using the target semantic modeler to obtain a text semantic unit sequence; and a speech reconstruction module 707, used to perform speech reconstruction based on the text semantic unit sequence to obtain target synthesized speech.
[0080] In one embodiment, the speech synthesis device further includes: a model training module, configured to: perform semantic modeling on sample text and target language using a target semantic modeler to obtain a target sample semantic unit sequence; determine sample noise, and perform flow matching on the sample noise and the target sample semantic unit sequence using an initial conditional flow matching model to obtain a sample prediction vector field; calculate the difference between the sample noise and the reference Mel spectrum corresponding to the sample speech to obtain an auxiliary vector field; calculate the loss based on the sample prediction vector field and the auxiliary vector field to obtain flow matching loss data; and adjust the parameters of the initial conditional flow matching model based on the flow matching loss data to obtain a target conditional flow matching model.
[0081] It should be noted that the specific implementation of this speech synthesis device is basically the same as the specific implementation of the speech synthesis method described above, and will not be repeated here.
[0082] This application also provides an electronic device, which includes: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for communication between the processor and the memory. When the program is executed by the processor, it implements the aforementioned speech synthesis method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0083] Please see Figure 8 , Figure 8 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 801 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 802 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 using the speech synthesis method of the embodiments of this application. The 803 input / output interface is used to implement information input and output. The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804); The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.
[0084] This application embodiment also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, which can be executed by one or more processors to implement the above-described speech synthesis method.
[0085] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0086] The speech synthesis method, speech synthesis device, electronic device, and storage medium proposed in this application first train a target semantic modeler based on a sample speech set and sample text belonging to the target language. Since the sample speech set contains sample speech from multiple languages, and the speech content of each sample speech is sample text, the target semantic modeler can ignore the influence of language on semantic modeling, thereby improving the cross-lingual capability of semantic modeling. Then, the target semantic modeler is used to perform semantic modeling on the target text and target language, obtaining a sequence of text semantic units; finally, speech reconstruction is performed based on the text semantic unit sequence to obtain the target synthesized speech. This overcomes the adverse effect of the limited number of sample speech in the target language in the sample speech set on the cross-lingual capability of the target semantic modeler, thereby improving the accuracy of cross-lingual speech synthesis.
[0087] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0088] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0089] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0090] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0091] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0092] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0093] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0094] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0095] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0096] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0097] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A speech synthesis method, characterized in that, The method includes: Obtain the target text and determine the target language to which the target text belongs; Obtain a sample speech set and sample text belonging to the target language; wherein, the sample speech set includes a first sample speech belonging to the target language and a second sample speech belonging to a non-target language, and the speech content of each first sample speech and each second sample speech is the same and is the sample text; The sample text and the target language are semantically modeled using a preset initial semantic modeler to obtain a first sample semantic unit sequence; Speech quantization is performed on the sample speech in the sample speech set to obtain a second sample semantic unit sequence; The initial semantic modeler is adjusted according to the first sample semantic unit sequence and the second sample semantic unit sequence to obtain the target semantic modeler; The target semantic modeler is used to perform semantic modeling on the target text and the target language to obtain a sequence of text semantic units. Speech reconstruction is performed based on the sequence of semantic units in the text to obtain the target synthesized speech.
2. The method according to claim 1, characterized in that, The step of adjusting the parameters of the initial semantic modeler based on the first sample semantic unit sequence and the second sample semantic unit sequence to obtain the target semantic modeler includes: The semantic loss data is obtained by calculating the loss based on the first sample semantic unit sequence and the second sample semantic unit sequence. The first sample semantic unit sequence is used to perform text conversion to obtain the first converted text, and the second sample semantic unit sequence is used to perform text conversion to obtain the second converted text. Loss data is obtained by calculating the loss based on the first converted text and the second converted text; The initial semantic modeler is adjusted based on the semantic loss data and the text loss data to obtain the target semantic modeler.
3. The method according to claim 1, characterized in that, The step of performing speech quantization on the sample speech in the sample speech set to obtain a second sample semantic unit sequence includes: Speech recognition is performed on each of the sample speech to obtain an initial semantic unit; Clustering the initial semantic units of each of the sample speech samples yields the second sample semantic unit sequence.
4. The method according to any one of claims 1 to 3, characterized in that, The step of reconstructing speech based on the text semantic unit sequence to obtain the target synthesized speech includes: The target Mel spectrum is obtained by performing flow matching on the text semantic unit sequence using a preset target condition flow matching model. The target Mel spectrum is converted into speech to obtain the target synthesized speech.
5. The method according to claim 4, characterized in that, Before performing flow matching on the text semantic unit sequence using a preset target condition flow matching model to obtain the target Mel spectrum, the method further includes: The target semantic modeler performs semantic modeling on the sample text and the target language to obtain a sequence of target sample semantic units. The sample noise is determined, and the sample noise and the target sample semantic unit sequence are flow matched using an initial conditional flow matching model to obtain the sample prediction vector field; The auxiliary vector field is obtained by calculating the difference between the sample noise and the reference Mel spectrum corresponding to the sample speech; Loss calculation is performed based on the sample prediction vector field and the auxiliary vector field to obtain flow matching loss data; The parameters of the initial conditional flow matching model are adjusted based on the flow matching loss data to obtain the target conditional flow matching model.
6. The method according to claim 4, characterized in that, The step of performing flow matching on the text semantic unit sequence using a preset target conditional flow matching model to obtain the target Mel spectrum includes: Obtain reference speech data that does not belong to the target language; Speaker coding is performed on the reference speech data to obtain speaker features; The target noise is determined, and the target noise, the text semantic unit sequence, and the speaker features are matched using the target conditional flow matching model to obtain the target Mel spectrum.
7. The method according to claim 5, characterized in that, The step of calculating the loss based on the sample prediction vector field and the auxiliary vector field to obtain the flow matching loss data includes: The difference between the sample prediction vector field and the auxiliary vector field is calculated to obtain the vector field difference. The L2 norm is calculated based on the vector field difference to obtain the flow matching loss data.
8. A speech synthesis device, characterized in that, The device includes: The text acquisition module is used to acquire target text and determine the target language to which the target text belongs; The sample acquisition module is used to acquire a sample speech set and sample text belonging to the target language; wherein, the sample speech set includes a first sample speech belonging to the target language and a second sample speech belonging to a non-target language, and the speech content of each first sample speech and each second sample speech is the same and is the sample text; The first semantic modeling module is used to perform semantic modeling on the sample text and the target language through a preset initial semantic modeler to obtain a first sample semantic unit sequence; The speech quantization module is used to perform speech quantization on the sample speech in the sample speech set to obtain a second sample semantic unit sequence; The parameter adjustment module is used to adjust the parameters of the initial semantic modeler according to the first sample semantic unit sequence and the second sample semantic unit sequence to obtain the target semantic modeler. The second semantic modeling module is used to perform semantic modeling on the target text and the target language through the target semantic modeler to obtain a sequence of text semantic units. The speech reconstruction module is used to reconstruct speech based on the text semantic unit sequence to obtain the target synthesized speech.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the speech synthesis method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech synthesis method according to any one of claims 1 to 7.