Speech synthesis method, system, electronic device, and storage medium
By separating and stretching the speech features of the target object, a speech synthesis model is constructed to achieve high-precision synthesis of speech at a specific age. This solves the problem of insufficient datasets in existing technologies and improves the quality and naturalness of speech synthesis.
Patent Information
- Application Number
- CN202411881960.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-12-19
AI Technical Summary
Existing technologies lack the ability to synthesize speech of specific ages with high precision, mainly due to the difficulty in acquiring datasets and insufficient sample representativeness, making it difficult to accurately capture the voice features of specific ages.
By acquiring the source speech features of the target object, age decoupling is performed using an encoder to separate age-related and irrelevant information, and feature extraction and stretching are performed through an age-aware module to construct a speech synthesis model to synthesize speech of the target age.
It achieves high-precision synthesis of speech for specific ages, reduces data collection costs and complexity, and provides highly customized, natural, fluent, and realistic speech synthesis services, thereby improving the user experience.
Smart Images

Figure CN119785764B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech synthesis technology, and in particular to a speech synthesis method, system, electronic device, and storage medium. Background Technology
[0002] In recent years, with the rapid advancement of artificial intelligence technology, speech synthesis technology has gradually become one of the indispensable core components of human-computer interaction. Therefore, how to perform high-precision speech synthesis to improve the quality and naturalness of speech synthesis is an important topic that urgently needs to be studied.
[0003] To achieve high-precision synthesis of speech from a specific speaker at a specific age in existing technologies, it is often necessary to use a large-scale speech dataset covering all stages of the speaker's life. However, such datasets are difficult to obtain and lack representativeness. Therefore, they lack the ability to accurately capture the voice features of a specific age, making it difficult to achieve high-precision synthesis of speech from a specific age. Summary of the Invention
[0004] This invention provides a speech synthesis method, system, electronic device, and storage medium to address the shortcomings of existing technologies, such as the lack of datasets and the difficulty in achieving high-precision synthesis of speech at a specific age, thereby improving the high-precision synthesis of speech at a specific age.
[0005] This invention provides a speech synthesis method, comprising:
[0006] Obtain the speech features of the source speech of the target object;
[0007] The speech features of the source speech are input into the encoder in the speech synthesis model to obtain the first coding feature and the second coding feature of the source speech; the first coding feature is an age-related feature, and the second coding feature is an age-independent feature.
[0008] The first encoded feature of the source speech is input into the age perception module in the speech synthesis model to obtain the target age feature of the first age data, and the target age feature of the second age data is obtained based on the target age feature of the first age data; the first age data is the age data corresponding to the source speech, and the second age data is the age data corresponding to the speech synthesis requirements of the target object.
[0009] The target age features of the second age data, the second coding features of the source speech, and the target text to be synthesized are input into the speech synthesis module in the speech synthesis model to obtain the target speech of the target object under the second age data.
[0010] The speech synthesis model is trained based on sample speech of the sample object, as well as the age label and sample text corresponding to the sample speech.
[0011] According to the present invention, a speech synthesis method is provided, wherein obtaining the target age features of second age data based on the target age features of the first age data includes:
[0012] In the age dictionary, find the standard age features corresponding to the first age data and the standard age features corresponding to the second age data;
[0013] The deviation between the standard age feature corresponding to the first age data and the standard age feature corresponding to the second age data is calculated to obtain the stretching parameter;
[0014] Based on the stretching parameters and the target age features of the first age data, feature transformation is performed to obtain the age transformation features corresponding to the source speech;
[0015] Based on the age conversion features corresponding to the source speech, the target age features of the second age data are obtained.
[0016] According to the present invention, a speech synthesis method is provided, wherein obtaining the target age feature of the second age data based on the age conversion feature corresponding to the source speech includes:
[0017] When the number of source voices is one, the age conversion feature corresponding to the source voice is used as the target age feature of the second age data;
[0018] When there are multiple source voices, the average value among the age conversion features corresponding to the multiple source voices is calculated to obtain the target age feature of the second age data.
[0019] According to the present invention, a speech synthesis method is provided, wherein the construction steps of the age dictionary include:
[0020] In the database, obtain each of the sample voices and the age label corresponding to each of the sample voices;
[0021] The speech features of each of the sample speech samples are input into the encoder to obtain the first encoded features of each of the sample speech samples;
[0022] The first encoded feature of each of the sample speech is input into the age perception module to obtain the target age feature of the age label corresponding to each of the sample speech;
[0023] Calculate the average value of the target age features corresponding to all sample speech under each age label to obtain the standard age features corresponding to each age label;
[0024] The age dictionary is obtained by mapping and storing each age label and the corresponding standard age feature.
[0025] According to the present invention, a speech synthesis method is provided, wherein the target age feature of the second age data, the second coding feature of the source speech, and the target text to be synthesized are input into the speech synthesis module of the speech synthesis model to obtain the target speech of the target object under the second age data, comprising:
[0026] The target age feature of the second age data and the second encoded feature of the source speech are input into the decoder in the speech synthesis module to obtain the decoded features;
[0027] The decoded features and the target text are input into the acoustic model in the speech synthesis module to obtain the reconstructed acoustic features;
[0028] The reconstructed acoustic features are input into the vocoder in the speech synthesis module to obtain the target speech.
[0029] According to the present invention, a speech synthesis method is provided, wherein the speech synthesis model is constructed based on the following steps:
[0030] The speech features of the sample speech are input into the initial encoder to obtain the first and second coding features of the sample speech;
[0031] The first and second coding features of the sample speech are input into the initial decoder to obtain the sample decoding features;
[0032] The sample decoding features and the sample text are input into the initial acoustic model to obtain the first sample reconstructed acoustic features;
[0033] The initial encoder, the initial decoder, and the initial acoustic model are trained based on the age category corresponding to the first coding feature of the sample speech, the age category corresponding to the second coding feature of the sample speech, the age label, the original acoustic features of the sample speech, and the reconstructed acoustic features of the first sample.
[0034] Based on the trained initial encoder, trained initial decoder, trained initial acoustic model, and initial age-aware module, construct the speech model to be trained.
[0035] The speech model to be trained is trained based on the sample speech, the original acoustic features, and the sample text to obtain the trained speech model.
[0036] The speech synthesis model is constructed based on the trained speech model.
[0037] According to the present invention, a speech synthesis method is provided, wherein training a speech model to be trained based on the sample speech, the original acoustic features, and the sample text to obtain a trained speech model includes:
[0038] The speech features of the sample speech and the sample text are input into the speech model to be trained to obtain the second sample reconstructed acoustic features;
[0039] Based on the original acoustic features and the reconstructed acoustic features from the second sample, the training age-aware module, the training decoder, and the training acoustic model in the training speech model are iteratively trained to obtain the trained speech model.
[0040] The present invention also provides a speech synthesis system, characterized in that it comprises:
[0041] The feature extraction unit is used to obtain the speech features of the source speech of the target object;
[0042] The encoding unit is used to input the speech features of the source speech into the encoder in the speech synthesis model to obtain the first encoding feature and the second encoding feature of the source speech; the first encoding feature is an age-related feature, and the second encoding feature is an age-independent feature;
[0043] The processing unit is configured to input the first encoded feature of the source speech into the age perception module in the speech synthesis model to obtain the target age feature of the first age data, and obtain the target age feature of the second age data based on the target age feature of the first age data; the first age data is the age data corresponding to the source speech, and the second age data is the age data corresponding to the speech synthesis requirements of the target object;
[0044] A synthesis unit is used to input the target age features of the second age data, the second coding features of the source speech, and the target text to be synthesized into the speech synthesis module in the speech synthesis model to obtain the target speech of the target object under the second age data.
[0045] The speech synthesis model is trained based on sample speech of the sample object, as well as the age label and sample text corresponding to the sample speech.
[0046] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the above-described speech synthesis methods.
[0047] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech synthesis method as described above.
[0048] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the above-described speech synthesis methods.
[0049] The speech synthesis method, system, electronic device, and storage medium provided by this invention separate age-related and age-independent information of the target object from the inherent speech features of the target object's source speech through an encoder. This allows for flexible control of age and speaker attributes. An age-aware module extracts and stretches age-related features to obtain the target age features corresponding to the target object's speech synthesis requirements. Based on this, the target speech corresponding to the target object's age data is synthesized. The entire process requires only a small amount of the target object's source speech to accurately capture the voice features of the target object at a specific age. It also learns other age-independent features of the target object in detail, enabling the synthesized speech to more realistically reflect the voice characteristics of the target object at different ages. This significantly improves the quality and naturalness of speech synthesis, thereby reducing the cost and complexity of data acquisition while increasing the synthesis accuracy of speech at a specific age. This provides users with highly customized, natural, fluent, and realistic speech synthesis services, enhancing the user experience. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0051] Figure 1 This is one of the flowcharts of the speech synthesis method provided by the present invention.
[0052] Figure 2 This is the second flowchart of the speech synthesis method provided by the present invention.
[0053] Figure 3 This is a schematic diagram of the training process of the speech synthesis model provided by the present invention.
[0054] Figure 4 This is a schematic diagram of the speech synthesis system provided by the present invention.
[0055] Figure 5This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0057] In recent years, with the rapid advancement of artificial intelligence technology, speech synthesis technology has gradually become an indispensable core component of human-computer interaction. Its applications cover a wide range of scenarios, from automated telephone response systems and intelligent navigation assistance to virtual personal assistants, and it has shown broad application prospects in education, entertainment, and many other fields. Traditional speech synthesis technology mainly focuses on converting input text into natural-sounding speech output. In this process, elements such as sound quality, speech rate, and intonation are key to determining the user experience.
[0058] To achieve high-precision synthesis of a speaker's voice at a specific age, theoretically, a large-scale speech dataset covering all stages of the speaker's lifespan is needed. Constructing such a dataset faces numerous challenges in reality, such as the difficulty of data acquisition and insufficient sample representativeness. Furthermore, when simulating the voice characteristics of different age groups, existing speech synthesis technologies can often only roughly divide them into coarse age groups such as youth, middle-aged, and elderly, lacking large-scale speech datasets covering all stages of the lifespan. Consequently, they lack the ability to accurately capture the voice characteristics of a specific age, which to some extent restricts the personalized performance and accuracy of speech synthesis technology, making it difficult to achieve high-precision synthesis of voice at a specific age.
[0059] In response, this application provides a speech synthesis method that adaptively synthesizes speech of different ages through age decoupling, thereby improving the high-precision synthesis of speech of a specific age.
[0060] It should be noted that the execution entity of this method can be a speech synthesis system, such as a smart speaker or a speech synthesizer; this implementation does not specifically limit this. The system is similar to a time machine for sound, allowing users to freely travel through their own vocal journey and experience a speech journey spanning time.
[0061] Figure 1 This is one of the flowcharts illustrating the speech synthesis method provided by the present invention; for example... Figure 1 As shown, the method includes steps 110, 120, 130 and 140.
[0062] Step 110: Obtain the speech features of the source speech of the target object.
[0063] The target audience here can be the speaker (or subject) for whom speech synthesis at a specific age is required. The source speech here is the registered speech with age data tags input by the user when speech synthesis is required for the target audience; the source speech will also be referred to as the registered speech below. The number of registered speeches can be one or more, and this embodiment does not specifically limit this.
[0064] The user input referred to may be information input through command line interface, graphical interface, voice input, etc., and this embodiment does not specifically limit it.
[0065] Optionally, after obtaining the source speech of the target object, feature extraction can be performed on the source speech to obtain its speech features.
[0066] Feature extraction here can involve first extracting acoustic features from the source speech to obtain raw acoustic features, and then further extracting features from the raw acoustic features to obtain speech features that can represent the speaker's speech characteristics. For example, Mel spectrum features can be extracted from the raw waveform of the source speech as acoustic features, and then the Mel spectrum features can be input into a deep learning network, such as an Integrable Deep Neural Network (IDNN), to further extract features from the raw acoustic features, obtaining an embedding vector X-vector that can represent speaker information, which can also be denoted as X. This yields the speech features that can represent the speaker's speech characteristics.
[0067] The feature extraction steps of IDNN include: after processing the extracted features through a series of deep neural network (DNN) layers, the network typically includes a special layer to aggregate temporal information, ensuring that the output feature vector has a fixed length. This aggregation layer can be a statistical pooling layer, which extracts statistical information, such as the mean and standard deviation, from a series of frame-level features to generate a fixed-length vector, regardless of the length of the input speech. The aggregated features are then fed into deeper network layers, ultimately forming a fixed-length embedding vector X-vector.
[0068] Step 120: Input the speech features of the source speech into the encoder of the speech synthesis model to obtain the first coding feature and the second coding feature of the source speech; the first coding feature is an age-related feature, and the second coding feature is an age-independent feature; wherein, the speech synthesis model is trained based on sample speech of the sample object, as well as the age label and sample text corresponding to the sample speech.
[0069] Optionally, before performing step 120, a speech synthesis model can be obtained through training. Specifically, an initialization model can be constructed first. This initialization model may include an initial encoder, an initial age perception module, and an initial speech synthesis module. The initial encoder may be a model prepared for age and timbre decoupling coding after parameter initialization, or it may be a pre-trained model with age and timbre decoupling coding functions. The initial age perception module may be a model prepared for age feature perception of a specific age after parameter initialization, or it may be a pre-trained model with age feature perception functions of a specific age. The initial speech synthesis module may be a model prepared for speech feature decoding and acoustic feature reconstruction after parameter initialization, or it may be a pre-trained model with speech feature decoding and acoustic feature reconstruction functions. This embodiment of the invention does not specifically limit the specifics of these limitations.
[0070] In addition, a large number of sample speakers (also called sample objects) at different ages, along with their age tags and sample text, can be collected to construct a dataset. For example, sample speech can be collected and age tags obtained through manual annotation; alternatively, age tags can be collected first, and then sample speech corresponding to the age tags can be obtained through manual recording or speech synthesis. The age tags provided by each sample speaker can be the same or different; this embodiment does not specifically limit this. The sample text here is the content description text of the sample speech.
[0071] Subsequently, the initialization model can be iteratively trained based on the sample speech, corresponding age labels, and sample text to construct a speech synthesis model that can accurately synthesize speech of a specific age. The training can be a holistic training of the initialization model based on the sample speech, corresponding age labels, and sample text, or a multi-task training with phased freezes based on the sample speech, corresponding age labels, and sample text. This embodiment does not specifically limit the training in this way.
[0072] In practical applications, the encoder in the trained speech synthesis model can be used to decouple age by applying the speech features of the source speech. This decoupling yields a first encoded feature related to age and a second encoded feature unrelated to age. The encoder here includes at least the bottleneck information (i.e., the first encoded feature) used for age association. Extracted age feature encoder And the bottleneck information used for age-independent processing (i.e., the second encoding feature). Age-independent feature encoder Among them, the age feature encoder Age-independent feature encoder The same model structure can be used.
[0073] Step 130: Input the first encoded feature of the source speech into the age perception module in the speech synthesis model to obtain the target age feature of the first age data, and obtain the target age feature of the second age data based on the target age feature of the first age data; the first age data is the age data corresponding to the source speech, and the second age data is the age data corresponding to the speech synthesis requirements of the target object.
[0074] Optionally, after obtaining the first and second encoded features through the encoder, the first encoded features can be... The input is fed into the age-aware module of the speech synthesis model, so that the age-aware module can apply the first encoded features. Precise control and stretching of the age attribute vector are performed to extract the vector of sound features corresponding to the first age data of the source speech, thereby obtaining the first age data. Target age characteristics .
[0075] The age-aware module here can be a linear sensing module, which obtains the corresponding target age feature by linearizing the first encoded feature.
[0076] It should be noted that, since the invertible convolutional generative flow model (Glow model for short) has been proven in computer vision to be able to stretch various attributes in an image after unsupervised training, enabling flexible control over image attributes and generating controllable and diverse images, the age perception module in this embodiment can be built and implemented based on the Glow model to precisely control and stretch the age attribute vector, thereby achieving vector extraction of sound features from different age data.
[0077] The Glow model contains multiple streaming modules; each individual streaming module consists of the following three parts:
[0078] Step (a), activate the Actnorm layer. Perform an affine transformation of the activation using the scale and bias parameters for each channel, initializing these parameters such that, given the initial data, the post-normalized activation for each channel has zero mean and unit variance. After initialization, the scale and bias are treated as regular trainable parameters independent of the data.
[0079] Step (b), reversible 1 1 convolutional layer. The output of the age feature encoder. Perform convolution. Wherein, It is the inverse of the weight matrix.
[0080] ;
[0081] in, Features are generated after processing by the convolutional layer; is the convolution kernel of the convolutional layer.
[0082] Step (c), Affine Coupling Layer. The affine coupling layer consists of the following three steps:
[0083] Step (c1), zero initialization. The last convolutional layer of each nonlinear mapping neural network is zero-initialized. Each affine coupling layer initially performs an identity function, which helps in training very deep networks.
[0084] Step (c2), splitting and merging. The channel dimensions are copied to form a new input, which is then split and stitched along the channel dimensions.
[0085] Step (c3), arrangement. Add a reversible 1 before the first two steps. 1. Convolution ensures that each dimension can influence other dimensions.
[0086] After obtaining the target age features of the first age data through the age perception module, age stretching can be performed based on the target age features of the first age data to form the second age data corresponding to the speech synthesis requirements of the target object. Target age characteristics .
[0087] Step 140: Input the target age features of the second age data, the second encoding features of the source speech, and the target text to be synthesized into the speech synthesis module in the speech synthesis model to obtain the target speech of the target object under the second age data.
[0088] Optionally, after obtaining the target age features of the second age data corresponding to the speech synthesis requirements of the target object, the target age features of the second age data can be fused with the second coding features. The fusion result and the target text are then input into the speech synthesis module in the speech synthesis model. The speech synthesis module applies the fused features of the target age features and the second coding features to simulate and synthesize the target speech of the target object under the second age data. This achieves the goal of synthesizing accurate target speech for different age stages of the target object with only a small amount of source speech, providing users with highly customized, natural, fluent, and realistic speech synthesis services to meet the needs of different scenarios and improve the user experience. The fusion here can be splicing, linear combination, etc., and this embodiment does not specifically limit it.
[0089] The method provided in this embodiment separates age-related and age-independent information of the target object from the inherent speech features of the target object's source speech through an encoder, enabling flexible control of age and speaker attributes. An age-aware module extracts and stretches age-related features to obtain the target age features corresponding to the target object's speech synthesis requirements. Based on this, the target speech corresponding to the target object's age data is synthesized. The entire process requires only a small amount of the target object's source speech to accurately capture the voice features of the target object at a specific age. It also learns other age-independent features of the target object in detail, allowing the synthesized speech to more realistically reflect the voice characteristics of the target object at different ages. This significantly improves the quality and naturalness of speech synthesis, thereby reducing the cost and complexity of data acquisition while increasing the synthesis accuracy of speech at a specific age. This provides users with highly customized, natural, fluent, and realistic speech synthesis services, enhancing the user experience.
[0090] In some embodiments, step 130 specifically includes:
[0091] In the age dictionary, find the standard age features corresponding to the first age data and the standard age features corresponding to the second age data;
[0092] The deviation between the standard age feature corresponding to the first age data and the standard age feature corresponding to the second age data is calculated to obtain the stretching parameter;
[0093] Based on the stretching parameters and the target age features of the first age data, feature transformation is performed to obtain the age transformation features corresponding to the source speech;
[0094] Based on the age conversion features corresponding to the source speech, the target age features of the second age data are obtained.
[0095] Figure 2 This is the second flowchart illustrating the speech synthesis method provided by this invention; as shown below. Figure 2 As shown, step 130, which involves obtaining the target age feature of the second age data through age stretching, specifically includes:
[0096] Based on the mapping relationship between different age data in the age dictionary and standard age features, the first age data is obtained. Corresponding standard age characteristics Second age data Corresponding standard age characteristics The age dictionary uses age data as keys and their corresponding standard age characteristics as values.
[0097] Subsequently, the standard age characteristics corresponding to the first age data are calculated. Standard age characteristics corresponding to the second age data The deviation between them yields the tensile parameters. The specific calculation formula is as follows:
[0098] ;
[0099] Then, the stretching parameters Target age characteristics of the first age data Linear stretching is performed to obtain the age conversion features corresponding to the source speech.
[0100] After obtaining the age conversion features corresponding to the source speech, it can be used directly as the target age feature of the second age data, or the age conversion features corresponding to the source speech can be processed to obtain the target age feature of the second age data. The specific method can be adaptively determined based on the number of source speech samples.
[0101] For example, in some embodiments, the step of obtaining the target age feature of the second age data specifically includes:
[0102] When the number of source voices is one, the age conversion feature corresponding to the source voice is used as the target age feature of the second age data;
[0103] When there are multiple source voices, the average value among the age conversion features corresponding to the multiple source voices is calculated to obtain the target age feature of the second age data.
[0104] Optionally, if the target object provides only one source speech, the above stretching step is performed on the source speech to obtain the corresponding age conversion feature obtained by stretching the source speech, and then the corresponding age conversion feature obtained by stretching the source speech is used as the target age feature of the second age data.
[0105] When the target object provides two or more source voices, the above stretching steps are iteratively performed on each source voice to obtain the corresponding age conversion features obtained by stretching each source voice. Then, the average of the age conversion features corresponding to all source voices is taken as the target age feature of the second age data. This allows for a more complete restoration of the age features of the second age data through multiple source voices, thereby further improving the synthesis accuracy and providing users with highly customized, natural, fluent, and realistic voice synthesis services to meet the needs of different scenarios.
[0106] The method provided in this embodiment can adaptively generate age features of a target age space from age features of a specific age space through age stretching and feature transformation. Based on the age features of the stretched target age space, target speech of the target age space is generated, which can make the synthesized speech more realistically reflect the sound characteristics of the target object's target age space, and can achieve high customization, natural fluency and strong realism.
[0107] In some embodiments, the construction steps of the age dictionary include:
[0108] In the database, obtain each of the sample voices and the age label corresponding to each of the sample voices;
[0109] The speech features of each of the sample speech samples are input into the encoder to obtain the first encoded features of each of the sample speech samples;
[0110] The first encoded feature of each of the sample speech is input into the age perception module to obtain the target age feature of the age label corresponding to each of the sample speech;
[0111] Calculate the average value of the target age features corresponding to all sample speech under each age label to obtain the standard age features corresponding to each age label;
[0112] The age dictionary is obtained by mapping and storing each age label and the corresponding standard age feature.
[0113] Optionally, the steps for constructing the age dictionary specifically include:
[0114] Data set from the database Sample speech from multiple speakers corresponding to each age label is used as a subset of the dataset. , , The number of age tags.
[0115] For each subset Each sample of speech is batch-input into the speech synthesis model to obtain that subset of the dataset. The standard age features associated with the corresponding age labels are defined in the following steps:
[0116] Extract the original acoustic features (such as Mel spectrum features) of the sample speech, and input the original acoustic features into the IDNN network to further extract the X-vector of the sample speech, that is, the speech features of the sample speech.
[0117] The speech features of the sample speech are input into the age feature encoder to obtain the first encoded features of the sample speech.
[0118] The first encoded feature of the sample speech is input into the age perception module to linearize the first encoded feature of the sample speech, thereby obtaining the target age feature of the age label corresponding to the sample speech.
[0119] Save this subset of datasets The target age features corresponding to the age labels of all sample speech data are averaged to obtain this subset of data. The standard age features corresponding to the age labels, that is, the keys for this subset of datasets. The value corresponding to the age tag.
[0120] Iterate through each subset of data using the steps described above. Each age label and its corresponding standard age feature are obtained. Each age label is used as the key and its corresponding standard age feature is used as the value to save the data, thus forming an age dictionary.
[0121] It should be noted that after obtaining the target speech based on the second age data in steps 110-140, the target speech and the corresponding second age data can be associated and stored in the database to expand the speech segments of each speaker at different age stages. This can alleviate the problem of data sparsity for each speaker at different age dimensions and help improve the performance of various speech processing tasks that require speech segments of each speaker at different age stages for model training or task processing.
[0122] In addition, the target speech can be expanded to the corresponding age label subset based on the second age data to update the age dictionary, thereby improving the accuracy of model inference and synthesis.
[0123] In some embodiments, step 140 specifically includes:
[0124] The target age feature of the second age data and the second encoded feature of the source speech are input into the decoder in the speech synthesis module to obtain the decoded features;
[0125] The decoded features and the target text are input into the acoustic model in the speech synthesis module to obtain the reconstructed acoustic features;
[0126] The reconstructed acoustic features are input into the vocoder in the speech synthesis module to obtain the target speech.
[0127] like Figure 2 As shown, the speech synthesis module includes a decoder, an acoustic model, and a vocoder connected in sequence. The decoder is used for feature decoding; the acoustic model is used for acoustic feature reconstruction; and the vocoder is used for speech synthesis.
[0128] Step 140 specifically includes:
[0129] The target age features from the second age data and the second encoded features from the source speech are fused and input into the decoder in the speech synthesis module. The decoder then decodes and outputs decoded features, which contain all the information about the speech features of the source speech. The specific calculation formula for the decoder is as follows:
[0130] ;
[0131] in, For decoding features; For the decoder's model function; The target age feature for the second age data; This is the second coding feature of the source speech.
[0132] After obtaining the decoded features output by the decoder, the decoded features and the target text can be jointly input into the acoustic model in the speech synthesis module. The acoustic model then uses the decoded features and the target text to reconstruct the corresponding reconstructed acoustic features. The target text here can be a synthesized speech description text used to describe the speech content that needs to be synthesized for the second age data, or it can be text containing marked or annotated feature information related to the second age data, etc. This embodiment does not specifically limit it.
[0133] After obtaining the reconstructed acoustic features output by the acoustic model, the reconstructed acoustic features can be input into the vocoder in the speech synthesis module, so that the vocoder can apply the reconstructed acoustic features to perform speech synthesis, thereby synthesizing the target speech of the target object under the second age data.
[0134] The method provided in this embodiment utilizes the collaboration of a decoder, an acoustic model, and a vocoder to accurately fuse the target age features of the second age data and the second coding features of the source speech, so as to reconstruct and synthesize the target speech of the target object that is highly realistic and consistent with the age features under the second age data.
[0135] In some embodiments, the speech synthesis model is constructed based on the following steps:
[0136] The speech features of the sample speech are input into the initial encoder to obtain the first and second coding features of the sample speech;
[0137] The first and second coding features of the sample speech are input into the initial decoder to obtain the sample decoding features;
[0138] The sample decoding features and the sample text are input into the initial acoustic model to obtain the first sample reconstructed acoustic features;
[0139] The initial encoder, the initial decoder, and the initial acoustic model are trained based on the age category corresponding to the first coding feature of the sample speech, the age category corresponding to the second coding feature of the sample speech, the age label, the original acoustic features of the sample speech, and the reconstructed acoustic features of the first sample.
[0140] Based on the trained initial encoder, trained initial decoder, trained initial acoustic model, and initial age-aware module, construct the speech model to be trained.
[0141] The speech model to be trained is trained based on the sample speech, the original acoustic features, and the sample text to obtain the trained speech model.
[0142] The speech synthesis model is constructed based on the trained speech model.
[0143] Figure 3 This is a schematic diagram of the training process of the speech synthesis model provided by the present invention; as shown below. Figure 3 As shown, the initialization model specifically includes an initial encoder, an initial age-aware module, an initial decoder, and an initial acoustic model.
[0144] The speech synthesis model is built by training on the initial model. To ensure the stability of the model training, the training of the speech synthesis model can be divided into two stages, namely the first stage training and the second stage training.
[0145] Optionally, the first phase of training includes the following steps:
[0146] First, we construct the dataset. This dataset contains a large number of sample voice recordings of individuals at different age stages, along with the corresponding age labels for the sample voice recordings.
[0147] Subsequently, acoustic features are extracted from the sample speech to obtain the original acoustic features of the sample speech. Then, the speaker information embedding vector is extracted from the original acoustic features of the sample speech to obtain the speech features of the sample speech. The specific implementation steps can be referred to the speech feature extraction steps of the source speech, which will not be repeated here.
[0148] Subsequently, the speech features of the sample speech are input into the initial encoder, which decouples the speech features of the sample speech by age, thereby obtaining age-related features and age-independent features from the speech features, and thus obtaining the first and second coding features of the sample speech.
[0149] The specific implementation steps for obtaining encoded features include:
[0150] Speech features of sample speech The initial age feature encoder is input into the initial encoder to obtain the first encoded feature of the sample speech output by the initial age feature encoder; the speech features of the sample speech are then input into the initial age-independent feature encoder of the initial encoder to obtain the second encoded feature of the sample speech output by the initial age-independent feature encoder. The first encoded feature... The calculation formula is as follows:
[0151] ;
[0152] in, This is the model function for the initial age feature encoder.
[0153] Second coding feature The calculation formula is as follows:
[0154] ;
[0155] in, This is the model function for the initial age-independent feature encoder.
[0156] Furthermore, the first encoded feature of the sample speech can be input into the age classifier, which then applies the first encoded feature to perform an age classification task, obtaining the age category corresponding to the first encoded feature of the sample speech. Similarly, the second encoded feature of the sample speech can be input into the age classifier, which then applies the second encoded feature to perform an age classification task, obtaining the age category corresponding to the second encoded feature of the sample speech.
[0157] Furthermore, the first and second coded features of the sample speech can be combined and input into the initial decoder. The initial decoder then applies the combined features of the first and second coded features to output all the information containing the speech features of the sample speech, i.e., the sample decoded features. The specific calculation formula is as follows:
[0158] ;
[0159] in, This is the model function for the initial decoder.
[0160] Subsequently, the sample decoding features The sample text is input into the initial acoustic model so that the initial acoustic model can apply the sample decoding features. Acoustic features are reconstructed from the sample text to obtain the reconstructed acoustic features of the first sample.
[0161] Subsequently, based on the error between the age category corresponding to the first encoded feature and the age label, the first classification loss function is calculated. Gradient flipping is performed based on the error between the age category corresponding to the second encoded feature and the age label to obtain the second classification loss function. Based on the deviation between the original acoustic features of the sample speech and the reconstructed acoustic features of the first sample, the reconstruction loss function is obtained. .
[0162] Subsequently, the first classification loss function is fused. Second classification loss function and reconstruction loss function Calculate the total loss function for the first stage of training. The fusion here can be a direct addition or a weighted addition, etc., and this embodiment does not specifically limit it. For example, the total loss function of the first stage of training... The calculation formula can be:
[0163] .
[0164] Subsequently, based on the total loss function of the first stage of training The model parameters of the initial encoder, initial decoder, and initial acoustic model in the initialization model are iteratively optimized to improve model performance. This enables the trained initial encoder to effectively decouple age, thereby reducing the mutual information between age features and age-independent features. In other words, the trained initial age feature encoder tends to retain age-related speaker features and remove redundant features in the encoded features (i.e., bottleneck information), and the trained initial age-independent feature encoder tends to retain age-independent speaker features in the encoded features (i.e., bottleneck information). Furthermore, the trained initial decoder and initial acoustic model can effectively reconstruct acoustic features.
[0165] The formula for calculating the mutual information between age-related features and age-independent features can be expressed as follows:
[0166] ;
[0167] in, This refers to the mutual information between age feature M and age-independent feature N.
[0168] Optionally, the second stage of training is a fine-tuning training of the initial model after the first stage of training as the base model. The purpose is to enable the model to learn the linear representation of age features, which is convenient for flexible age stretching. The training steps may include: based on sample speech, the original acoustic features of the sample speech, and sample text, with the goal of enabling the model to accurately learn the linear representation of age features, fine-tuning training is performed on the speech model to be trained, which is constructed from the initial encoder, the initial decoder, the initial acoustic model, and the initial age perception module in the initial model after the first stage of training, to obtain the trained speech model. Based on the trained speech model, a speech synthesis model is constructed.
[0169] The fine-tuning training here can be fine-tuning the overall parameters of the speech model to be trained, or it can be fine-tuning training by freezing some parameters of the speech model to be trained. This embodiment does not specifically limit it in this way.
[0170] For example, in some embodiments, the second phase of training specifically includes:
[0171] The speech features of the sample speech and the sample text are input into the speech model to be trained to obtain the second sample reconstructed acoustic features;
[0172] Based on the original acoustic features and the reconstructed acoustic features from the second sample, the training age-aware module, the training decoder, and the training acoustic model in the training speech model are iteratively trained to obtain the trained speech model.
[0173] Optionally, in the second stage of training, the speech features of the sample speech can be sequentially passed through the speech model to be trained. The encoder to be trained (i.e., the initial encoder after training) in the speech model to be trained extracts the first and second coding features of the sample speech. The age perception module to be trained (i.e., the initial age perception module) in the speech model to be trained performs a linear mapping on the first coding features of the sample speech to obtain the target age features corresponding to the age label of the sample speech. The decoder to be trained (i.e., the initial decoder after training) in the speech model to be trained applies the first coding features of the sample speech and the target age features corresponding to the age label of the sample speech to decode, thereby obtaining the sample decoding features of the sample speech. The acoustic model to be trained (i.e., the initial acoustic model after training) in the speech model to be trained applies the sample decoding features of the sample speech and the sample text to reconstruct the second sample reconstructed acoustic features.
[0174] Subsequently, based on the error between the reconstructed acoustic features of the second sample and the original acoustic features of the sample speech, a loss function for the second stage of training is constructed. Then, with the encoder to be trained frozen, the loss function of the second stage training is used to iteratively train the age-aware module, decoder, and acoustic model in the speech model to be trained, resulting in a trained speech model. Based on the trained speech model, a speech synthesis model that can accurately synthesize speech of a specific age is constructed. In other words, by adding a vocoder to the trained speech model, a speech synthesis model that can accurately synthesize speech of a specific age can be constructed.
[0175] The method provided in this embodiment, in the first stage of training, utilizes an initialization model to perform age classification and acoustic feature reconstruction on sample speech, constructing a composite loss function including classification loss and reconstruction loss. The model parameters are iteratively optimized, enabling the encoder to effectively distinguish and encode age-related and age-independent features, while the decoder and acoustic model accurately reconstruct acoustic features. This ensures the model can accurately capture the speaker's age characteristics, laying a solid foundation for subsequent age stretching and speech synthesis. In the second stage of training, using the model trained in the first stage as a foundation, an initial age-aware module is introduced to construct the speech model to be trained. Through fine-tuning training, the model further learns the linear representation of age features, facilitating flexible age stretching effects. This not only improves the model's generalization ability but also ensures the accuracy and naturalness of synthesized speech at specific age stages. Therefore, through a refined two-stage training strategy, a speech synthesis model capable of accurately synthesizing speech at specific age stages can be effectively constructed.
[0176] The speech synthesis system provided by the present invention is described below. The speech synthesis system described below can be referred to in correspondence with the speech synthesis method described above.
[0177] Figure 4 This is a schematic diagram of the speech synthesis system provided by the present invention; as shown below. Figure 4 As shown, the system includes:
[0178] The feature extraction unit 410 is used to obtain the speech features of the source speech of the target object;
[0179] The encoding unit 420 is used to input the speech features of the source speech into the encoder in the speech synthesis model to obtain the first encoding feature and the second encoding feature of the source speech; the first encoding feature is an age-related feature, and the second encoding feature is an age-independent feature;
[0180] The processing unit 430 is used to input the first encoded feature of the source speech into the age perception module in the speech synthesis model to obtain the target age feature of the first age data, and obtain the target age feature of the second age data based on the target age feature of the first age data; the first age data is the age data corresponding to the source speech, and the second age data is the age data corresponding to the speech synthesis requirements of the target object;
[0181] Synthesis unit 440 is used to input the target age feature of the second age data, the second coding feature of the source speech, and the target text to be synthesized into the speech synthesis module in the speech synthesis model to obtain the target speech of the target object under the second age data;
[0182] The speech synthesis model is trained based on sample speech of the sample object, as well as the age label and sample text corresponding to the sample speech.
[0183] The system provided in this embodiment separates age-related and age-independent information of the target object from the inherent speech features of the target object's source speech through an encoder. This allows for flexible control of age and speaker attributes. An age-aware module extracts and stretches age-related features to obtain the target age features corresponding to the target object's speech synthesis requirements. Based on this, the system synthesizes the target speech corresponding to the target object's age data. The entire process requires only a small amount of the target object's source speech to accurately capture the voice features of the target object at a specific age. It also learns other age-independent features of the target object in detail, enabling the synthesized speech to more realistically reflect the voice characteristics of the target object at different ages. This significantly improves the quality and naturalness of speech synthesis, thereby reducing the cost and complexity of data acquisition while increasing the synthesis accuracy of speech at specific ages. Ultimately, this provides users with highly customized, natural, fluent, and realistic speech synthesis services, enhancing the user experience.
[0184] In some embodiments, the processing unit is specifically used for:
[0185] In the age dictionary, find the standard age features corresponding to the first age data and the standard age features corresponding to the second age data;
[0186] The deviation between the standard age feature corresponding to the first age data and the standard age feature corresponding to the second age data is calculated to obtain the stretching parameter;
[0187] Based on the stretching parameters and the target age features of the first age data, feature transformation is performed to obtain the age transformation features corresponding to the source speech;
[0188] Based on the age conversion features corresponding to the source speech, the target age features of the second age data are obtained.
[0189] In some embodiments, the processing unit is further configured to:
[0190] When the number of source voices is one, the age conversion feature corresponding to the source voice is used as the target age feature of the second age data;
[0191] When there are multiple source voices, the average value among the age conversion features corresponding to the multiple source voices is calculated to obtain the target age feature of the second age data.
[0192] In some embodiments, the system further includes a dictionary building unit, specifically used for:
[0193] In the database, obtain each of the sample voices and the age label corresponding to each of the sample voices;
[0194] The speech features of each of the sample speech samples are input into the encoder to obtain the first encoded features of each of the sample speech samples;
[0195] The first encoded feature of each of the sample speech is input into the age perception module to obtain the target age feature of the age label corresponding to each of the sample speech;
[0196] Calculate the average value of the target age features corresponding to all sample speech under each age label to obtain the standard age features corresponding to each age label;
[0197] The age dictionary is obtained by mapping and storing each age label and the corresponding standard age feature.
[0198] In some embodiments, the synthesis unit is specifically used for:
[0199] The target age feature of the second age data and the second encoded feature of the source speech are input into the decoder in the speech synthesis module to obtain the decoded features;
[0200] The decoded features and the target text are input into the acoustic model in the speech synthesis module to obtain the reconstructed acoustic features;
[0201] The reconstructed acoustic features are input into the vocoder in the speech synthesis module to obtain the target speech.
[0202] In some embodiments, the system further includes a model training unit, specifically used for:
[0203] The speech features of the sample speech are input into the initial encoder to obtain the first and second coding features of the sample speech;
[0204] The first and second coding features of the sample speech are input into the initial decoder to obtain the sample decoding features;
[0205] The sample decoding features and the sample text are input into the initial acoustic model to obtain the first sample reconstructed acoustic features;
[0206] The initial encoder, the initial decoder, and the initial acoustic model are trained based on the age category corresponding to the first coding feature of the sample speech, the age category corresponding to the second coding feature of the sample speech, the age label, the original acoustic features of the sample speech, and the reconstructed acoustic features of the first sample.
[0207] Based on the trained initial encoder, trained initial decoder, trained initial acoustic model, and initial age-aware module, construct the speech model to be trained.
[0208] The speech model to be trained is trained based on the sample speech, the original acoustic features, and the sample text to obtain the trained speech model.
[0209] The speech synthesis model is constructed based on the trained speech model.
[0210] In some embodiments, the model training unit is further configured to:
[0211] The speech features of the sample speech and the sample text are input into the speech model to be trained to obtain the second sample reconstructed acoustic features;
[0212] Based on the original acoustic features and the reconstructed acoustic features from the second sample, the training age-aware module, the training decoder, and the training acoustic model in the training speech model are iteratively trained to obtain the trained speech model.
[0213] The system provided by this invention is used to execute the above-described method embodiments. For specific processes and details, please refer to the above embodiments, which will not be repeated here.
[0214] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communications bus 540. The processor 510 can call logic instructions in the memory 530 to execute a speech synthesis method, which includes: acquiring speech features of source speech of a target object; inputting the speech features of the source speech into an encoder in a speech synthesis model to obtain a first encoded feature and a second encoded feature of the source speech; the first encoded feature is an age-related feature, and the second encoded feature is an age-independent feature; inputting the first encoded feature of the source speech into an age-aware module in the speech synthesis model to obtain a target age feature of a first age data, and acquiring a target age feature of a second age data based on the target age feature of the first age data; the first age data is the age data corresponding to the source speech, and the second age data is the age data corresponding to the speech synthesis requirement of the target object; inputting the target age feature of the second age data, the second encoded feature of the source speech, and the target text to be synthesized into a speech synthesis module in the speech synthesis model to obtain the target speech of the target object under the second age data; wherein, the speech synthesis model is trained based on sample speech of a sample object, as well as the age label and sample text corresponding to the sample speech.
[0215] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0216] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech synthesis method provided by the above methods. The method includes: acquiring the speech features of the source speech of a target object; inputting the speech features of the source speech into an encoder in a speech synthesis model to obtain a first coding feature and a second coding feature of the source speech; the first coding feature is an age-related feature, and the second coding feature is an age-independent feature; inputting the first coding feature of the source speech into an age-aware module in the speech synthesis model to obtain a target age feature of a first age data, and acquiring a target age feature of a second age data based on the target age feature of the first age data; the first age data is the age data corresponding to the source speech, and the second age data is the age data corresponding to the speech synthesis requirement of the target object; inputting the target age feature of the second age data, the second coding feature of the source speech, and the target text to be synthesized into a speech synthesis module in the speech synthesis model to obtain the target speech of the target object under the second age data; wherein, the speech synthesis model is trained based on sample speech of a sample object, as well as the age label and sample text corresponding to the sample speech.
[0217] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the speech synthesis method provided by the above methods. The method includes: acquiring speech features of source speech of a target object; inputting the speech features of the source speech into an encoder in a speech synthesis model to obtain a first encoded feature and a second encoded feature of the source speech; the first encoded feature is an age-related feature, and the second encoded feature is an age-independent feature; inputting the first encoded feature of the source speech into an age-aware module in the speech synthesis model to obtain a target age feature of first age data, and acquiring a target age feature of second age data based on the target age feature of the first age data; the first age data is the age data corresponding to the source speech, and the second age data is the age data corresponding to the speech synthesis requirement of the target object; inputting the target age feature of the second age data, the second encoded feature of the source speech, and the target text to be synthesized into a speech synthesis module in the speech synthesis model to obtain the target speech of the target object under the second age data; wherein the speech synthesis model is trained based on sample speech of a sample object, and the age label and sample text corresponding to the sample speech.
[0218] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0219] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0220] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech synthesis method, characterized in that, include: The speech features of the source speech of the target object are input into the encoder in the speech synthesis model to obtain the first coding feature and the second coding feature of the source speech; the first coding feature is an age-related feature, and the second coding feature is an age-independent feature; The first encoded feature is input into the age perception module in the speech synthesis model to obtain the target age feature of the first age data. The deviation between the standard age feature corresponding to the first age data and the standard age feature corresponding to the second age data is calculated to obtain the stretching parameter. Based on the stretching parameter and the target age feature of the first age data, feature transformation is performed to obtain the age transformation feature. Based on the age transformation feature, the target age feature of the second age data is obtained. The first age data is the age data corresponding to the source speech, and the second age data is the age data corresponding to the speech synthesis requirements of the target object; The target age features of the second age data, the second encoding features, and the target text to be synthesized are input into the speech synthesis module in the speech synthesis model to obtain the target speech of the target object under the second age data. The speech synthesis model is trained based on sample speech of the sample object, as well as the age label and sample text corresponding to the sample speech; the encoder includes two feature encoders with the same model structure, which are used to extract the first encoded feature and the second encoded feature, respectively. The age perception module is a linear perception module. By linearizing the first encoded feature, the corresponding target age feature is obtained. The age perception module is built based on the Glow model and is used to control and stretch the age attribute vector to extract the vector of sound features under different age data.
2. The speech synthesis method according to claim 1, characterized in that, The step of obtaining the target age feature of the second age data based on the age conversion feature includes: When the number of source voices is one, the age conversion feature corresponding to the source voice is used as the target age feature of the second age data; When there are multiple source voices, the average value among the age conversion features corresponding to the multiple source voices is calculated to obtain the target age feature of the second age data.
3. The speech synthesis method according to claim 1, characterized in that, The standard age features corresponding to the first age data and the standard age features corresponding to the second age data are obtained by looking up in the age dictionary; The steps for constructing the age dictionary include: In the database, obtain each of the sample voices and the age label corresponding to each of the sample voices; The speech features of each of the sample speech samples are input into the encoder to obtain the first encoded features of each of the sample speech samples; The first encoded feature of each of the sample speech is input into the age perception module to obtain the target age feature of the age label corresponding to each of the sample speech; Calculate the average value of the target age features corresponding to all sample speech under each age label to obtain the standard age features corresponding to each age label; The age dictionary is obtained by mapping and storing each age label and the corresponding standard age feature.
4. The speech synthesis method according to any one of claims 1-3, characterized in that, The step of inputting the target age features of the second age data, the second encoding features, and the target text to be synthesized into the speech synthesis module of the speech synthesis model to obtain the target speech of the target object under the second age data includes: The target age feature of the second age data and the second encoded feature of the source speech are input into the decoder in the speech synthesis module to obtain the decoded features; The decoded features and the target text are input into the acoustic model in the speech synthesis module to obtain the reconstructed acoustic features; The reconstructed acoustic features are input into the vocoder in the speech synthesis module to obtain the target speech.
5. The speech synthesis method according to any one of claims 1-3, characterized in that, The speech synthesis model is constructed based on the following steps: The speech features of the sample speech are input into the initial encoder to obtain the first and second coding features of the sample speech; The first and second coding features of the sample speech are input into the initial decoder to obtain the sample decoding features; The sample decoding features and the sample text are input into the initial acoustic model to obtain the first sample reconstructed acoustic features; The initial encoder, the initial decoder, and the initial acoustic model are trained based on the age category corresponding to the first coding feature of the sample speech, the age category corresponding to the second coding feature of the sample speech, the age label, the original acoustic features of the sample speech, and the reconstructed acoustic features of the first sample. Based on the trained initial encoder, trained initial decoder, trained initial acoustic model, and initial age-aware module, construct the speech model to be trained. The speech model to be trained is trained based on the sample speech, the original acoustic features, and the sample text to obtain the trained speech model. The speech synthesis model is constructed based on the trained speech model.
6. The speech synthesis method according to claim 5, characterized in that, The step of training the speech model to be trained based on the sample speech, the original acoustic features, and the sample text to obtain the trained speech model includes: The speech features of the sample speech and the sample text are input into the speech model to be trained to obtain the second sample reconstructed acoustic features; Based on the original acoustic features and the reconstructed acoustic features from the second sample, the training age-aware module, the training decoder, and the training acoustic model in the training speech model are iteratively trained to obtain the trained speech model.
7. A speech synthesis system, characterized in that, include: The feature extraction unit is used to obtain the speech features of the source speech of the target object; The encoding unit is used to input the speech features of the source speech into the encoder in the speech synthesis model to obtain the first encoding feature and the second encoding feature of the source speech; the first encoding feature is an age-related feature, and the second encoding feature is an age-independent feature; The processing unit is configured to input the first encoded feature into the age perception module in the speech synthesis model to obtain the target age feature of the first age data, calculate the deviation between the standard age feature corresponding to the first age data and the standard age feature corresponding to the second age data to obtain the stretching parameter, perform feature transformation based on the stretching parameter and the target age feature of the first age data to obtain the age transformation feature, and obtain the target age feature of the second age data based on the age transformation feature. The first age data is the age data corresponding to the source speech, and the second age data is the age data corresponding to the speech synthesis requirements of the target object; A synthesis unit is used to input the target age features of the second age data, the second encoding features, and the target text to be synthesized into the speech synthesis module in the speech synthesis model to obtain the target speech of the target object under the second age data. The speech synthesis model is trained based on sample speech of the sample object, as well as the age label and sample text corresponding to the sample speech; the encoder includes two feature encoders with the same model structure, which are used to extract the first encoded feature and the second encoded feature, respectively. The age perception module is a linear perception module. By linearizing the first encoded feature, the corresponding target age feature is obtained. The age perception module is built based on the Glow model and is used to control and stretch the age attribute vector to extract the vector of sound features under different age data.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the speech synthesis method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the speech synthesis method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Method, device, equipment and medium for generating audio
CN112652292A
Age-based sound generation method and device
CN116030787A