Speech synthesis method and device, computer equipment and storage medium
By constructing a joint optimization loss function and parameter optimization method, the problem that existing speech synthesis technology cannot reflect the unique voice characteristics of users is solved, and personalized speech synthesis is generated to meet the needs of diverse application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-03
AI Technical Summary
Existing speech synthesis technology is unable to meet the diverse application scenarios and cannot reflect the unique voice characteristics of different users, resulting in synthesized speech lacking personalization and realism.
By acquiring training data of user identifiers, training text, and training speech, the original multi-user speech synthesis model is invoked, a joint optimization loss function is constructed, and the model parameters are optimized to generate an optimized multi-user speech synthesis model. Finally, personalized synthesized speech is generated based on the target user identifier.
It enables the generation of voices with different user-specific characteristics, meeting diverse voice synthesis needs and improving the personalization and realism of voice synthesis.
Smart Images

Figure CN121789633A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to a speech synthesis method, apparatus, computer device, and storage medium. Background Technology
[0002] Speech synthesis technology, which converts text information into speech signals, has been widely used in many fields, such as intelligent voice assistants, audiobook production, and voice navigation, especially in the fields of fintech, healthcare, and elderly care.
[0003] Traditional speech synthesis methods are typically based on a single speech model that learns from a large amount of speech data to generate generic speech. However, this generic model fails to capture the unique voice characteristics of different users, such as timbre, intonation, and speech rate, resulting in synthesized speech that lacks personalization and realism.
[0004] This shows that existing speech synthesis technology is insufficient to meet the diverse application scenarios. Summary of the Invention
[0005] The purpose of this application is to provide a speech synthesis method, apparatus, computer device, and storage medium to solve the problem that existing speech synthesis technologies are unable to meet diverse application scenarios.
[0006] To address the aforementioned technical problems, this application provides a speech synthesis method, employing the following technical solution: Obtain model training data, wherein the model training data includes a training user identifier, training text corresponding to the training user identifier, and training speech corresponding to the training user identifier; The original multi-user speech synthesis model is invoked, and the user identifier and the training text are input into the original multi-user speech synthesis model to obtain synthesized speech; Construct a joint optimized loss function based on the synthesized speech and the training speech; The parameters of the original multi-user speech synthesis model are optimized according to the joint optimization loss function to obtain an optimized multi-user speech synthesis model. Obtain the text to be synthesized and the target user identifier; The text to be synthesized and the target user identifier are input into the optimized multi-user speech synthesis model to obtain the target synthesized speech.
[0007] To address the aforementioned technical problems, this application also provides a speech synthesis device, which employs the following technical solution: The training data acquisition module is used to acquire model training data, wherein the model training data includes a training user identifier, training text corresponding to the training user identifier, and training speech corresponding to the training user identifier. The synthesized speech acquisition module is used to call the original multi-user speech synthesis model and input the user identifier and the training text into the original multi-user speech synthesis model to obtain synthesized speech. A joint optimization loss function construction module is used to construct a joint optimization loss function based on the synthesized speech and the training speech; The parameter optimization module is used to optimize the parameters of the original multi-user speech synthesis model according to the joint optimization loss function to obtain an optimized multi-user speech synthesis model. The data acquisition module is used to acquire the text to be synthesized and the target user identifier; The speech synthesis output module is used to input the text to be synthesized and the target user identifier into the optimized multi-user speech synthesis model to obtain the target synthesized speech.
[0008] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution: It includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the speech synthesis method as described above.
[0009] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below: The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the speech synthesis method described above.
[0010] This application provides a speech synthesis method, comprising: acquiring model training data, wherein the model training data includes training user identifiers, training text corresponding to the training user identifiers, and training speech corresponding to the training user identifiers; invoking an original multi-user speech synthesis model and inputting the user identifiers and the training text into the original multi-user speech synthesis model to obtain synthesized speech; constructing a joint optimization loss function based on the synthesized speech and the training speech; optimizing the parameters of the original multi-user speech synthesis model based on the joint optimization loss function to obtain an optimized multi-user speech synthesis model; acquiring text to be synthesized and a target user identifier; and inputting the text to be synthesized and the target user identifier into the optimized multi-user speech synthesis model to obtain target synthesized speech. Compared with the prior art, this application can utilize multi-user training data to optimize the original speech synthesis model, thereby generating speech with different user-specific characteristics, meeting diverse speech synthesis needs. Attached Figure Description
[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart illustrating the implementation of the speech synthesis method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the speech synthesis device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0016] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0017] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0019] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0020] It should be noted that the speech synthesis method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the speech synthesis device is generally set in the server / terminal device.
[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0022] Continue to refer to Figure 2 The diagram shows a flowchart of an embodiment of the speech synthesis method according to this application. The speech synthesis method described above includes steps S201, S202, S203, S204, S205, and S206.
[0023] In step S201, model training data is obtained, wherein the model training data includes training user identifiers, training text corresponding to the training user identifiers, and training speech corresponding to the training user identifiers.
[0024] In this application's embodiments, model training data can be collected in various ways. For example, volunteers can be organized to participate in voice recording activities, and each volunteer can be assigned a unique training user identifier. Volunteers can then read pre-prepared training text aloud, and their voices can be recorded using professional recording equipment, thereby obtaining training text and training voice corresponding to the training user identifiers. The collected data should cover users of different ages, genders, and regions to ensure the model's generalization ability.
[0025] In step S202, the original multi-user speech synthesis model is invoked, and the user identifier and training text are input into the original multi-user speech synthesis model to obtain synthesized speech.
[0026] In the embodiments of this application, the multi-user speech synthesis model applicable to this application is mainly used to generate natural and fluent speech of multiple user dialogues under the conditions of input text and speaker. The multi-user speech synthesis model can be a model built based on a large language model. Specifically, the multi-user speech synthesis model can be, but is not limited to, Transformers model, LLaMA model (Large Language Model Meta AI, a large-scale language model developed by Meta), etc. It should be understood that the examples of multi-user speech synthesis models here are only for ease of understanding and are not intended to limit this application.
[0027] In this embodiment, the core formula of the multi-user speech synthesis model is:
[0028] in, Indicates the first A voice token, Indicates the target text to be synthesized. Indicates the first A voice token, This represents the model parameters of a large language model. This means generating the first speech token given the target text, previously generated speech tokens, and parameters of a large language model. The probability of a voice token.
[0029] In step S203, a joint optimization loss function is constructed based on the synthesized speech and the training speech.
[0030] In this embodiment, the joint optimization loss function can be composed of multiple sub-loss functions to comprehensively measure the differences between synthesized speech and training speech. For example, it can include a spectral loss function to measure the differences in spectral features between synthesized speech and training speech; and a prosodic loss function to measure the differences in prosody such as intonation and speech rate. These sub-loss functions are combined according to certain weights to obtain the joint optimization loss function.
[0031] In step S204, the parameters of the original multi-user speech synthesis model are optimized according to the joint optimization loss function to obtain the optimized multi-user speech synthesis model.
[0032] In this embodiment, the parameters of the original multi-user speech synthesis model are optimized using an optimization algorithm based on the constructed joint optimization loss function. Commonly used optimization algorithms include stochastic gradient descent (SGD) and Adam. During the optimization process, the gradient of the joint optimization loss function with respect to the model parameters is calculated, and then the model parameters are updated in the opposite direction of the gradient, causing the value of the loss function to gradually decrease. After multiple iterations of training, until the loss function converges or reaches the preset number of iterations, the optimized multi-user speech synthesis model is obtained.
[0033] In step S205, the text to be synthesized and the target user identifier are obtained.
[0034] In the embodiments of this application, a user inputs a request through their terminal device, which is received by the system. The terminal device may be a mobile terminal such as a mobile phone, smartphone, laptop, digital broadcast receiver, PDA (personal digital assistant), PAD (tablet computer), PMP (portable multimedia player), navigation device, etc., or a fixed terminal such as a digital TV, desktop computer, etc. It should be understood that the examples of terminal devices given herein are for convenience of understanding only and are not intended to limit this application.
[0035] In step S206, the text to be synthesized and the target user identifier are input into the optimized multi-user speech synthesis model to obtain the target synthesized speech.
[0036] In practical applications, particularly in the financial sector, intelligent customer service systems need to provide highly personalized, compliant, and secure voice interaction services. For example, a bank needs to provide VIP clients with a dedicated, synthesized voice for their financial advisors. This synthesized voice must retain the advisor's unique intonation while adhering to financial terminology standards. This method achieves this goal through a multi-user voice synthesis model, specifically: (1) Model training data acquisition: First, voice samples were collected from 50 senior financial advisors, with each advisor assigned a unique user identifier (e.g., ID_001 to ID_050). Then, a standard financial script library was selected, including financial product introduction texts (e.g., "This structured deposit has an annualized yield of 3.5%-4.2%, with a minimum investment of 100,000 yuan"), risk disclosure texts, etc. Finally, the real voices of each advisor reading the training texts were recorded at a sampling rate of 44.1kHz and a depth of 16 bits to ensure no background noise. (2) Original multi-user speech synthesis model call: First, a Transformer-based TTS architecture is adopted, embedding a user identifier encoding module. Inputting a user identifier (e.g., ID_001) and training text, the text encoder generates a phoneme sequence, which, combined with the user identifier embedding vector, drives the acoustic model to generate a Mel spectrum. Then, inputting ID_001 and the text "The subscription fee rate for this fund is 1.5%", the model first maps the user identifier to a 128-dimensional embedding vector, fuses it with the text encoding output, and then generates synthesized speech through a vocoder. This synthesized speech should reflect the unique characteristics of ID_001, such as speech rate (e.g., 0.6x speech rate) and pitch (fundamental frequency 180Hz). (3) Construction of joint optimization loss function: Based on sampling synthesized speech and real speech at the same time point, the mean square error between corresponding sampling points is calculated to construct a joint optimization loss function; (4) Model parameter optimization: The Adam optimizer was used with an initial learning rate of 0.001 and a batch size of 32. Training was performed for 200 epochs on a financial corpus, with evaluation on the validation set every 10 epochs. An early stopping mechanism was triggered if the validation set loss did not decrease for five consecutive iterations. (5) Target speech synthesis generation: When a customer inquires, "What is the risk level of this financial product?", the system automatically matches the target user ID_003 (the customer's dedicated advisor). Then, the text and ID_003 are input into the optimized model to generate synthesized speech. This speech must pass a compliance check module to ensure it does not contain illegal expressions such as "guaranteed returns," and that the speech rate and pauses conform to financial language standards. Finally, a 16kHz audio stream compliant with the RTP protocol is generated and transmitted to the customer's terminal in real time.
[0037] This application provides a speech synthesis method, comprising: acquiring model training data, wherein the model training data includes training user identifiers, training text corresponding to the training user identifiers, and training speech corresponding to the training user identifiers; calling an original multi-user speech synthesis model and inputting the user identifiers and training text into the original multi-user speech synthesis model to obtain synthesized speech; constructing a joint optimization loss function based on the synthesized speech and training speech; optimizing the parameters of the original multi-user speech synthesis model based on the joint optimization loss function to obtain an optimized multi-user speech synthesis model; acquiring the text to be synthesized and the target user identifier; and inputting the text to be synthesized and the target user identifier into the optimized multi-user speech synthesis model to obtain the target synthesized speech. Compared with the prior art, this application can utilize multi-user training data to optimize the original speech synthesis model, thereby generating speech with different user-specific characteristics, meeting diverse speech synthesis needs.
[0038] In some optional implementations of the embodiments of this application, the step of constructing a joint optimization loss function based on synthesized speech and training speech specifically includes the following steps: Construct the main reconstruction loss function based on the synthesized speech and the training speech; The synthesized speech is input into the frame-level label classifier to calculate the frame-level label classification loss function; The synthesized speech and training speech are input into a pre-trained identifier encoder to calculate the embedding similarity loss function; A joint optimization loss function is constructed based on the main reconstruction loss function, the frame-level identifier classification loss function, and the embedding similarity loss function.
[0039] In this embodiment, the principal reconstruction loss function is a core function used to quantify the difference between the model-reconstructed data and the original data. Its core function is to optimize model parameters and improve reconstruction quality by minimizing this difference. This principal reconstruction loss function can be the mean squared error (MSE) loss function. Synthetic speech and training speech are sampled at the same time points, and the mean squared error between corresponding sampling points is calculated to measure the difference in waveform between the two. Specifically, the principal reconstruction loss function can be expressed as:
[0040] in, Represents the principal reconstruction loss function. Indicates the number of sampling points. Indicates the training speech at the 1st The value of each sampling point, Indicates the synthesized speech in the first... The value of each sampling point.
[0041] In this embodiment, the frame-level label classification loss function is mainly used to solve the frame-level classification task. Specifically, by quantifying the difference between the model's classification prediction for each frame and the true label, it guides the model to optimize parameters to improve classification accuracy. This frame-level label classification loss function... It can be represented as:
[0042]
[0043] in, This represents the probability distribution of user identifiers during training. Indicates the number of speakers.
[0044] In this embodiment, the embedding similarity loss function is mainly used to measure the similarity between different data samples in the embedding space. Specifically, by optimizing the model parameters, the embedding vectors of similar samples are made closer, and the embedding vectors of dissimilar samples are made further apart.
[0045] In this embodiment, the main reconstruction loss function, frame-level identifier classification loss function, and embedding similarity loss function are combined according to certain weights to obtain a joint optimization loss function, as shown in the following formula:
[0046] in, This represents the principal reconstruction loss function mentioned above. This represents the frame-level identifier classification loss function described above. Let α and β represent the embedding similarity loss function described above, and let α and β represent the weight parameters.
[0047] Compared with existing technologies, this application constructs a joint optimization loss function that includes a main reconstruction loss function, a frame-level identifier classification loss function, and an embedding similarity loss function. This function measures the difference between synthesized speech and real speech from multiple dimensions, such as overall features, frame-level features, and user identifier-related features. This enables the model to comprehensively learn speech features, thereby generating higher-quality synthesized speech that is closer to real speech.
[0048] In some optional implementations of the embodiments of this application, the step of inputting synthesized speech to a frame-level identifier classifier to calculate the frame-level identifier classification loss function specifically includes the following steps: The synthesized acoustic features of each frame of the synthesized speech are input into a frame-level label classifier to obtain the training user label probability distribution of the synthesized acoustic features of each frame. The frame-level identifier classification loss function is obtained by calculating the probability distribution of training user identifiers and the cross-entropy loss of training acoustic features and training user labels for each frame of training speech.
[0049] In this embodiment, the synthesized acoustic features of each frame of synthesized speech are sequentially input into a pre-trained frame-level labeling classifier. The frame-level labeling classifier then processes each frame's synthesized acoustic features and outputs the probability values of that frame belonging to each possible user identifier. These probability values constitute a probability distribution vector, which is the training user identifier probability distribution of each frame's synthesized acoustic features. For example, if there are N possible user identifiers, then the probability distribution vector for each frame has a dimension of N, and each element in the vector represents the probability that the frame belongs to the corresponding user identifier. Specifically, this training user identifier probability distribution can be expressed as:
[0050] In this embodiment, cross-entropy loss is primarily used to measure the difference between two probability distributions. This application trains the user identifier probability distribution and compares it with the actual frame-level speaker labels. Align and calculate cross-entropy loss:
[0051] in, Indicates the number of speakers.
[0052] Compared with existing technologies, this application inputs the synthetic acoustic features of each frame of synthesized speech into a frame-level label classifier to obtain the training user label probability distribution of the synthetic acoustic features of each frame. Then, it calculates the cross-entropy loss between the probability distribution and the training user label corresponding to each frame of training acoustic features of the training speech, and finally obtains the frame-level label classification loss function. This provides a more accurate optimization direction for model training and improves the performance of synthesized speech in user labeling-related tasks.
[0053] In some optional implementations of the embodiments of this application, the step of inputting the synthesized speech and training speech into a pre-trained identifier encoder to calculate the embedding similarity loss function specifically includes the following steps: The synthesized speech and the training speech are respectively input into the pre-trained label encoder to obtain the synthesized speech label and the training speech label; Calculate the cosine similarity between the synthesized speech identifier and the trained speech identifier; Construct an embedding similarity loss function based on cosine similarity.
[0054] In this embodiment, a pre-trained identifier encoder (such as ECAPA-TDNN or HuBERT-SPK) is used to extract the speaker embeddings of the training speech and the synthesized speech.
[0055] Calculate the cosine similarity and define the embedding similarity loss function as follows:
[0056] Compared with existing technologies, this application first inputs the synthesized speech and training speech into a pre-trained label encoder to obtain the synthesized speech label and the training speech label, respectively; then it calculates the cosine similarity between the two labels; finally, it constructs an embedding similarity loss function based on the cosine similarity, providing a more accurate optimization direction for the training of the speech model and improving the performance of the model in speech similarity-related tasks.
[0057] In some optional implementations of the embodiments of this application, after the step of obtaining model training data, the following steps are further included: The training speech is processed by frame segmentation and windowing, and then a short-time Fourier transform is performed to obtain the spectrum of the training speech. Obtain the noise spectrum in the speechless segments of the training speech; The enhanced training speech spectrum is obtained by subtracting the noise spectrum from the spectrum of each frame of the training speech. The enhanced training speech spectrum is subjected to inverse short-time Fourier transform to obtain the denoised training speech.
[0058] In this embodiment, since speech signals have short-term stationarity, meaning their characteristics are relatively stable over a short period (typically 10-30 ms), the continuous training speech signal is divided into multiple short frames to facilitate feature extraction and analysis for each frame. Specifically, an appropriate frame length and frame shift are selected. The frame length is typically the number of sampling points corresponding to 10-30 ms, and the frame shift is generally half or one-third of the frame length. Simultaneously, a fixed-length window function (such as a rectangular window, Hamming window, or Hanning window) is used to window each frame of the speech signal. The window function reduces spectral leakage, making the signal in each frame more concentrated in the frequency domain.
[0059] In the embodiments of this application, a short-time Fourier transform (STFT) is performed on each frame of the windowed speech signal to convert the time-domain signal into a frequency-domain signal, thereby obtaining the training speech spectrum.
[0060] In this embodiment, the noise spectrum is obtained from the speechless segments of the training speech. Speechless segments typically refer to the silent segments before or after the start of speech. During these time periods, the speech signal has low energy and mainly contains information such as environmental noise. The noise spectrum is obtained by performing a short-time Fourier transform on the speech signal from the speechless segments.
[0061] In this embodiment of the application, the noise spectrum is subtracted from the spectrum of each frame of the training speech to obtain the enhanced training speech spectrum.
[0062] In this embodiment, the enhanced training speech spectrum is subjected to inverse short-time Fourier transform (ISTFT) to convert the frequency domain signal back to the time domain signal, thereby obtaining the denoised training speech.
[0063] Compared with existing technologies, this application can simply and effectively remove noise from training speech, improve the quality of training speech, and has low computational complexity, which can meet the requirements of real-time processing.
[0064] In some optional implementations of the embodiments of this application, after the step of obtaining model training data, the following steps are further included: The training speech corresponding to each training user identifier is encoded using a variational autoencoder to obtain the intonation variation pattern of each training user identifier. Identify stress adjustment information in the training text based on intonation variation patterns; Personalized voice feature vectors are obtained by jointly training the accent adjustment information and the training user identifiers using a generative adversarial network. Based on the application scenario, the personalized voice feature vectors are divided into batches to obtain the adjusted dynamic feature set; The adjusted dynamic feature set is reconstructed and optimized using a variational autoencoder to obtain the optimized training speech.
[0065] In this embodiment, the Variational Autoencoder (VAE) is a generative model consisting of an encoder and a decoder. The encoder maps the input training speech to a low-dimensional latent space, obtaining latent variables that contain key feature information of the training speech. In this invention, the encoder part of the Variational Autoencoder is used to encode the training speech corresponding to each training user identifier. Through training, the encoder can learn the implicit intonation variation patterns in the speech of different users, encoding the training speech of each training user identifier into a corresponding intonation variation pattern. This pattern is represented in the form of latent variables, reflecting the unique intonation features of the user.
[0066] In this embodiment, a variational autoencoder model is constructed using TensorFlow or PyTorch, including an encoder and a decoder. The encoder consists of a multi-layer convolutional neural network (CNN) or a recurrent neural network (RNN) to extract speech features and map them to a latent space. The decoder consists of a deconvolutional neural network or a reverse recurrent neural network to reconstruct the latent variables into speech signals. Then, training speech data is input into the variational autoencoder for training. Loss functions such as mean squared error (MSE) are used to measure the difference between the reconstructed speech and the original speech. The model parameters are updated using a backpropagation algorithm until the model converges. After training, the encoder is used to encode the training speech for each training user identifier to obtain the corresponding intonation variation pattern (latest variable).
[0067] In this embodiment, based on linguistic knowledge and the characteristics of intonation changes, a mapping rule is established between intonation change patterns and stress adjustment. For example, it is stipulated that when the intonation rise exceeds a certain threshold, the stress intensity of a specific part of speech (such as verbs or adjectives) in the corresponding text needs to be increased. Then, the intonation change pattern of each training user is matched with the training text, and the stress adjustment information of each word or syllable in the training text is calculated according to the mapping rule, including the stress intensity adjustment value and whether stress needs to be added.
[0068] In this embodiment, the input data consists of accent adjustment information and training user identifiers, which are then fed into a generative adversarial network (GAN) for training along with the corresponding real speech feature vectors. An adversarial loss function (such as the least squares GAN loss function) is used to optimize the generator and discriminator. After multiple rounds of training, the generator is able to generate high-quality personalized voice feature vectors.
[0069] In this application embodiment, application scenarios are divided into different categories based on their characteristics, such as voice navigation, voice stories, and voice customer service. Based on the application scenario corresponding to the generation of personalized voice feature vectors, this application assigns the personalized voice feature vectors to corresponding batches, forming an adjusted dynamic feature set. For example, personalized voice feature vectors used for voice navigation are assigned to one batch, while those used for voice stories are assigned to another batch.
[0070] In this embodiment, the decoder part of the trained variational autoencoder takes the latent variables in the adjusted dynamic feature set as input and performs element reconstruction optimization. The decoder reconstructs the latent variables into speech signals to obtain the optimized training speech.
[0071] Compared with existing technologies, this application can accurately capture the user's intonation change pattern, accurately determine the stress adjustment information, generate a highly personalized voice feature vector, and flexibly adjust and optimize it according to different application scenarios, ultimately obtaining an optimized training voice, improving the personalization and naturalness of the voice.
[0072] In some optional implementations of the embodiments of this application, the original multi-user speech synthesis model includes a speaker embedding layer. After the step of inputting user identifiers and training text into the original multi-user speech synthesis model to obtain synthesized speech, the following steps are further included: A low-rank decomposition matrix is introduced into the speaker embedding layer based on the LORA algorithm; The low-rank decomposition matrix is trained based on the synthesized speech to obtain the optimized speaker embedding layer.
[0073] In this embodiment, firstly, the parameter matrix structure of the speaker embedding layer in the original speech model is analyzed to determine its dimension and number of parameters. It is assumed that the original speaker embedding matrix is... ,in It is the dimension of the embedded vector. The number of speakers; then, according to the LoRA algorithm, this application decomposes the original matrix W into two low-rank matrices. and The product of, i.e. ,in It is the rank of a low-rank matrix, and min(d,n), by choosing an appropriate rank This approach can significantly reduce the number of parameters while ensuring model performance. Finally, this application initializes the low-rank matrices A and B using random initialization methods, such as Xavier initialization or Kaiming initialization, or meaningful initialization based on prior knowledge to accelerate model convergence.
[0074] In this embodiment, the speaker embedding layer after introducing a low-rank decomposition matrix is integrated into the original speech model to construct a new model architecture, ensuring that other parts of the model remain unchanged and only the speaker embedding layer is optimized.
[0075] In the embodiments of this application, the training objective is set according to the specific speech task. For example, in the speech synthesis task, the training objective may be to make the speaker features of the synthesized speech as similar as possible to the target speaker; in the speaker verification task, the training objective may be to improve the model's ability to distinguish between different speakers.
[0076] In this application embodiment, an appropriate loss function is selected to measure the gap between the model's output and the training objective. Common loss functions include mean squared error (MSE) and cross-entropy loss. For optimization of the speaker embedding layer, contrastive loss or triplet loss can be used to enhance the model's ability to distinguish speaker features.
[0077] In this embodiment, the model is trained using a prepared synthetic speech dataset. Synthetic speech samples are input into the model, the loss function value is calculated, and the parameters of low-rank matrices A and B are updated using the backpropagation algorithm. Optimization algorithms such as stochastic gradient descent (SGD) or its variants (such as the Adam optimizer) are employed to accelerate model convergence. During training, a validation set can be used to monitor model performance, prevent overfitting, and adjust hyperparameters such as learning rate and batch size based on the validation set performance.
[0078] In this embodiment, training is considered converged when the model's performance on the validation set no longer significantly improves or reaches a preset number of training rounds. At this point, the parameters of the low-rank matrices A and B have been optimized to a relatively good state.
[0079] In this embodiment, the optimized low-rank matrices A and B are multiplied to obtain the optimized speaker embedding matrix W′=A×B. W′ is then used to replace the speaker embedding matrix W in the original speech model to obtain the optimized speech model, which includes the optimized speaker embedding layer.
[0080] Compared with existing technologies, this application first introduces a low-rank decomposition matrix into the speaker embedding layer based on the LoRA algorithm, decomposing the original high-dimensional speaker embedding matrix into the product of two low-rank matrices; then, it uses synthesized speech to train the low-rank decomposition matrix, and by continuously adjusting the parameters of the low-rank matrix, the speaker embedding layer can better capture and represent the features of different speakers; finally, an optimized speaker embedding layer is obtained, which improves the performance of the speech model in speaker-related tasks.
[0081] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0082] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0083] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0084] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0085] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a speech synthesis device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0086] like Figure 3 As shown, the speech synthesis device 200 of this application embodiment includes: The training data acquisition module 210 is used to acquire model training data, wherein the model training data includes training user identifiers, training text corresponding to the training user identifiers, and training speech corresponding to the training user identifiers. The synthesized speech acquisition module 220 is used to call the original multi-user speech synthesis model and input the user identifier and training text into the original multi-user speech synthesis model to obtain synthesized speech. The joint optimization loss function construction module 230 is used to construct a joint optimization loss function based on the synthesized speech and the training speech; The parameter optimization module 240 is used to optimize the parameters of the original multi-user speech synthesis model according to the joint optimization loss function to obtain an optimized multi-user speech synthesis model. The data acquisition module 250 is used to acquire the text to be synthesized and the target user identifier; The speech synthesis output module 260 is used to input the text to be synthesized and the target user identifier into the optimized multi-user speech synthesis model to obtain the target synthesized speech.
[0087] In this embodiment, a speech synthesis device 200 is provided, comprising: a training data acquisition module 210, used to acquire model training data, wherein the model training data includes training user identifiers, training text corresponding to the training user identifiers, and training speech corresponding to the training user identifiers; a synthesized speech acquisition module 220, used to call an original multi-user speech synthesis model and input the user identifiers and training text into the original multi-user speech synthesis model to obtain synthesized speech; a joint optimization loss function construction module 230, used to construct a joint optimization loss function based on the synthesized speech and training speech; a parameter optimization module 240, used to optimize the parameters of the original multi-user speech synthesis model based on the joint optimization loss function to obtain an optimized multi-user speech synthesis model; a data to be synthesized acquisition module 250, used to acquire the text to be synthesized and the target user identifier; and a speech synthesis output module 260, used to input the text to be synthesized and the target user identifier into the optimized multi-user speech synthesis model to obtain the target synthesized speech. Compared with the prior art, this application can utilize multi-user training data to optimize the original speech synthesis model, thereby generating speech with different user-specific characteristics, meeting diverse speech synthesis needs.
[0088] In some optional implementations of the embodiments of this application, the above-mentioned joint optimization loss function construction module includes: The main reconstruction loss function construction submodule is used to construct the main reconstruction loss function based on the synthesized speech and the training speech; The frame-level identifier classification loss function calculation submodule is used to input the synthesized speech into the frame-level identifier classifier to calculate the frame-level identifier classification loss function; The embedding similarity loss function calculation submodule is used to input the synthesized speech and training speech into the pre-trained identifier encoder to calculate the embedding similarity loss function; The joint optimization loss function construction submodule is used to construct a joint optimization loss function based on the main reconstruction loss function, the frame-level identifier classification loss function, and the embedding similarity loss function.
[0089] In some optional implementations of the embodiments of this application, the above-mentioned frame-level identifier classification loss function calculation submodule includes: The probability distribution calculation unit is used to input the synthesized acoustic features of each frame of synthesized speech into the frame-level label classifier to obtain the training user label probability distribution of the synthesized acoustic features of each frame. The cross-entropy loss calculation unit is used to calculate the cross-entropy loss of the training user identifier probability distribution and the training acoustic features of each frame of the training speech, resulting in a frame-level identifier classification loss function.
[0090] In some optional implementations of the embodiments of this application, the above-mentioned embedding similarity loss function calculation submodule includes: The speech identifier acquisition unit is used to input the synthesized speech and the training speech into the pre-trained identifier encoder to obtain the synthesized speech identifier and the training speech identifier, respectively. The cosine similarity calculation unit is used to calculate the cosine similarity between the synthesized speech identifier and the trained speech identifier; The embedding similarity loss function calculation unit is used to construct the embedding similarity loss function based on cosine similarity.
[0091] In some optional implementations of the embodiments of this application, the above-mentioned speech synthesis device 200 further includes: The low-rank decomposition matrix introduction module is used to introduce low-rank decomposition matrices into the speaker embedding layer according to the LORA algorithm; The matrix training module is used to train the low-rank decomposition matrix based on the synthesized speech to obtain the optimized speaker embedding layer.
[0092] In some optional implementations of the embodiments of this application, the above-mentioned speech synthesis device 200 further includes: In some optional implementations of the embodiments of this application, the above-mentioned ... module includes: To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of a computer device according to an embodiment of this application.
[0093] Computer device 300 includes a memory 310, a processor 320, and a network interface 330 that are interconnected via a system bus. It should be noted that only computer device 300 with components 310-330 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0094] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through methods such as keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0095] The memory 310 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 310 may be an internal storage unit of the computer device 300, such as the hard disk or memory of the computer device 300. In other embodiments, the memory 310 may also be an external storage device of the computer device 300, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Of course, the memory 310 may include both internal storage units and external storage devices of the computer device 300. In the embodiments of this application, the memory 310 is typically used to store the operating system and various application software installed on the computer device 300, such as computer-readable instructions for speech synthesis methods. In addition, the memory 310 can also be used to temporarily store various types of data that have been output or will be output.
[0096] In some embodiments, processor 320 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. Processor 320 is typically used to control the overall operation of computer device 300. In embodiments of this application, processor 320 is used to execute computer-readable instructions stored in memory 310 or process data, such as executing computer-readable instructions for a speech synthesis method.
[0097] The network interface 330 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 300 and other electronic devices.
[0098] The computer device provided in this application can optimize the original speech synthesis model using training data from multiple users, thereby generating speech with different user-specific characteristics to meet diverse speech synthesis needs.
[0099] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the speech synthesis method described above.
[0100] The computer-readable storage medium provided in this application can optimize the original speech synthesis model using training data from multiple users, thereby generating speech with different user-specific characteristics to meet diverse speech synthesis needs.
[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0102] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A speech synthesis method, characterized in that, Includes the following steps: Obtain model training data, wherein the model training data includes a training user identifier, training text corresponding to the training user identifier, and training speech corresponding to the training user identifier; The original multi-user speech synthesis model is invoked, and the user identifier and the training text are input into the original multi-user speech synthesis model to obtain synthesized speech; Construct a joint optimized loss function based on the synthesized speech and the training speech; The parameters of the original multi-user speech synthesis model are optimized according to the joint optimization loss function to obtain an optimized multi-user speech synthesis model. Obtain the text to be synthesized and the target user identifier; The text to be synthesized and the target user identifier are input into the optimized multi-user speech synthesis model to obtain the target synthesized speech.
2. The speech synthesis method according to claim 1, characterized in that, The step of constructing a joint optimization loss function based on the synthesized speech and the training speech specifically includes the following steps: Construct the main reconstruction loss function based on the synthesized speech and the training speech; The synthesized speech is input into a frame-level identifier classifier to calculate the frame-level identifier classification loss function; The synthesized speech and the training speech are input into a pre-trained identifier encoder to calculate the embedding similarity loss function; The joint optimization loss function is constructed based on the main reconstruction loss function, the frame-level identifier classification loss function, and the embedding similarity loss function.
3. The speech synthesis method according to claim 2, characterized in that, The step of inputting the synthesized speech into the frame-level identifier classifier to calculate the frame-level identifier classification loss function specifically includes the following steps: The synthesized acoustic features of each frame of the synthesized speech are input into the frame-level identifier classifier to obtain the training user identifier probability distribution of the synthesized acoustic features of each frame; The frame-level identifier classification loss function is obtained by calculating the probability distribution of the training user identifier and the cross-entropy loss of the training acoustic features and training user labels for each frame of the training speech.
4. The speech synthesis method according to claim 2, characterized in that, The step of inputting the synthesized speech and the training speech into a pre-trained identifier encoder to calculate the embedding similarity loss function specifically includes the following steps: The synthesized speech and the training speech are respectively input into a pre-trained label encoder to obtain the synthesized speech label and the training speech label; Calculate the cosine similarity between the synthesized speech identifier and the trained speech identifier; The embedding similarity loss function is constructed based on the cosine similarity.
5. The speech synthesis method according to claim 1, characterized in that, Following the step of obtaining model training data, the following steps are also included: The training speech is processed by frame segmentation and windowing, and then a short-time Fourier transform is performed to obtain the spectrum of the training speech. Obtain the noise spectrum in the speechless segment of the training speech; The enhanced training speech spectrum is obtained by subtracting the noise spectrum from the spectrum of each frame of the training speech; The enhanced training speech spectrum is subjected to inverse short-time Fourier transform to obtain the denoised training speech.
6. The speech synthesis method according to claim 1, characterized in that, Following the step of obtaining model training data, the following steps are also included: The training speech corresponding to each training user identifier is encoded using a variational autoencoder to obtain the intonation variation pattern of each training user identifier. Based on the intonation variation pattern, stress adjustment information is determined in the training text; A personalized voice feature vector is obtained by jointly training the accent adjustment information and the training user identifier using a generative adversarial network. The personalized voice feature vectors are divided into batches according to the application scenario to obtain an adjusted dynamic feature set. The adjusted dynamic feature set is reconstructed and optimized using a variational autoencoder to obtain the optimized training speech.
7. The speech synthesis method according to claim 1, characterized in that, The original multi-user speech synthesis model includes a speaker embedding layer. After the step of inputting the user identifier and the training text into the original multi-user speech synthesis model to obtain synthesized speech, the model further includes the following steps: A low-rank decomposition matrix is introduced into the speaker embedding layer according to the LORA algorithm; The low-rank decomposition matrix is trained based on the synthesized speech to obtain the optimized speaker embedding layer.
8. A speech synthesis device, characterized in that, include: The training data acquisition module is used to acquire model training data, wherein the model training data includes a training user identifier, training text corresponding to the training user identifier, and training speech corresponding to the training user identifier. The synthesized speech acquisition module is used to call the original multi-user speech synthesis model and input the user identifier and the training text into the original multi-user speech synthesis model to obtain synthesized speech. A joint optimization loss function construction module is used to construct a joint optimization loss function based on the synthesized speech and the training speech; The parameter optimization module is used to optimize the parameters of the original multi-user speech synthesis model according to the joint optimization loss function to obtain an optimized multi-user speech synthesis model. The data acquisition module is used to acquire the text to be synthesized and the target user identifier; The speech synthesis output module is used to input the text to be synthesized and the target user identifier into the optimized multi-user speech synthesis model to obtain the target synthesized speech.
9. A computer device, comprising a memory and a processor, characterized in that, The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements the steps of the speech synthesis method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the speech synthesis method as described in any one of claims 1 to 7.