Digital human timbre adaptive matching method and system based on voiceprint feature migration
Through the combination of deep vocabulary decoupling network and diffusion stream matching neural vocoder, multi-scale tone features are extracted and combined, the limitations of adaptive tone matching in the existing technology are solved, and high-fidelity and personalized digital human tone synthesis are realized, which improves the naturalness and reality of speech synthesis.
Patent Information
- Application Number
- CN202510779082.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing voice synthesis technology has limitations in achieving highly realistic, fine, controllable and pure digital human tone adaptive matching, and it is difficult to independently and finely adjust and customize tone details, limiting the flexibility and personalization of tone shaping.
The deep voiceprint decoupling network is used to extract pure tone identity embedding and multi-scale tone characteristics from the voice samples through a multi-task adversarial learning strategy, and customized voiceprint feature combination and speech synthesis are combined with a diffusion stream matching neural vocoder to achieve accurate and flexible matching of tone characteristics.
It significantly improves the sound quality fidelity and auditory reality of digital human voice, enhances the richness and customization of tone performance, and makes digital human voice closer to real human voices and adapts to different application scenarios and user needs.
Smart Images

Figure CN120356474B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech synthesis technology, and in particular to a method and system for adaptively matching digital human timbre based on voiceprint feature migration. Background Art
[0002] As a crucial vehicle for human-computer interaction and virtual content presentation, the realism and expressiveness of digital humans are crucial to the user experience. Among numerous perceptual dimensions, speech synthesis technology empowers digital humans with the ability to speak, and the timbre of synthesized speech is a key element in shaping their unique personality, identity, and emotions. Current speech synthesis technologies, particularly text-to-speech systems based on deep neural networks, have made significant progress in improving the naturalness and similarity of synthesized speech. Mainstream approaches employ a user encoder to extract the target user's voiceprint features and incorporate this information as conditional information into acoustic models (such as Tacotron, FastSpeech, and VITS) or neural vocoders to generate speech with a user-specific timbre. These approaches aim to learn and transfer timbre information from target speech samples, enabling personalized speech synthesis and voice cloning. Some research also focuses on feature decoupling techniques, attempting to separate timbre features from other acoustic properties such as speech content and prosody during the synthesis process, in order to obtain a purer timbre representation for controlling speech synthesis.
[0003] While existing speech synthesis technologies have achieved some success in timbre control, they still face limitations in achieving highly realistic, finely controlled, and pure adaptive matching of digital human voices. Existing methods have limited control over timbre details, and most offer global timbre embedding. This makes it difficult to independently and finely adjust and customize the macroscopic and microscopic details of timbre, limiting the flexibility and personalization of timbre shaping.
[0004] Therefore, a digital human voice adaptive matching method and system based on voiceprint feature migration is proposed. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for adaptively matching the voice of a digital human. By deeply decoupling and customizing the combined voiceprint features, combined with advanced vocoders, accurate and flexible personalized voice synthesis can be achieved, thereby enhancing realism and expressiveness.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] The digital human voice adaptive matching method based on voiceprint feature migration includes:
[0008] A deep voiceprint decoupling network is used to process user voice samples. The deep voiceprint decoupling network uses a multi-task adversarial learning strategy to extract a pure timbre identity embedding from the voice sample that is decoupled from content information and general acoustic feature information, and extracts multi-scale timbre features that characterize the macro characteristics and micro details of the user's timbre;
[0009] According to user configuration, select and combine the multi-scale timbre features and assign speech synthesis weights to form a customized voiceprint feature set;
[0010] The pure timbre identity embedding and customized voiceprint feature set are used as conditions and injected into the diffusion stream matching neural vocoder; after receiving the content information to be synthesized, the diffusion stream matching neural vocoder generates a digital human speech synthesis waveform with the user's unique timbre and corresponding to the content through probability stream conversion and step-by-step denoising and refinement.
[0011] Preferably, the deep voiceprint decoupling network includes at least one user identity encoding backbone network, at least one content information discriminator and at least one general acoustic feature information discriminator; the general acoustic feature information includes the average fundamental frequency range, basic rhythmic pattern and general acoustic expression of common emotion types.
[0012] The multi-task adversarial learning strategy includes:
[0013] Minimize the user classification loss of timbre identity embedding by training the user identity encoding backbone network;
[0014] Maximizing the accuracy of calculating content information from timbre identity embedding by training the content information discriminator, and minimizing the accuracy of calculating content information from timbre identity embedding by training the user identity encoding backbone network;
[0015] By training the universal acoustic feature information discriminator, the accuracy of calculating the universal acoustic feature information from the timbre identity embedding is maximized, and at the same time, the user identity encoding backbone network is trained to minimize the accuracy of calculating the universal acoustic feature information from the timbre identity embedding.
[0016] Preferably, the multi-scale timbre features include:
[0017] The global timbre features extracted from the deep network layer of the deep voiceprint decoupling network characterize the global static characteristics of the timbre; and the local timbre features extracted from the shallow network layer of the deep voiceprint decoupling network characterize the local dynamic details of the timbre; the global timbre features include the average range and basic sound quality parameters, and the local timbre features include the resonance peak dynamic parameters and the acoustic pattern of the vocalization habit.
[0018] Preferably, the step of selecting a combination from the multi-scale timbre features and assigning speech synthesis weights according to user configuration includes:
[0019] Receive user input, including at least one user configuration parameter, the user configuration parameter including the timbre style label of the target digital person, the desired degree of realism, and specific application scenario information; parse the user configuration parameter and map it to a target area in a predefined timbre feature space; preliminarily screen out a feature subset most relevant to the target area from the extracted global timbre features and local timbre features; set initial baseline contribution weights for the screened global timbre feature subset and local timbre feature subset respectively; and dynamically adjust the adaptive speech synthesis weight based on the user configuration parameter to calculate the final speech synthesis weight of each selected timbre feature in the customized voiceprint feature set.
[0020] Preferably, the processing of the diffusion flow matching neural vocoder specifically includes:
[0021] Receiving the pure voice identity embedding, the customized voiceprint feature set, and the content information to be synthesized as input conditions; using a conditional probability flow model, mapping and generating a preliminary acoustic representation of predefined Gaussian noise through a reversible transformation, wherein the preliminary acoustic representation is output as a Mel-spectrogram;
[0022] The preliminary acoustic representation is used as the starting state of the diffusion process; a denoising network is initialized, and the denoising network uses the pure timbre identity embedding, the customized voiceprint feature set and the speech content information to be synthesized as conditional inputs; iterative denoising is performed through a plurality of preset time steps, and at each time step, the denoising network gradually removes noise and enhances the details and harmonic structure of the speech based on the acoustic representation of the current time step and the injected conditional information; this denoising step is repeated until a preset termination condition is reached to obtain a denoised acoustic representation; and the denoised acoustic representation is converted into a final digital human speech waveform through a vocoder decoding module.
[0023] Another aspect of the present invention is a digital human voice adaptive matching system based on voiceprint feature migration, comprising:
[0024] A voiceprint feature processing module, which uses a deep voiceprint decoupling network to process user voice samples. The deep voiceprint decoupling network uses a multi-task adversarial learning strategy to extract a pure timbre identity embedding from the voice sample that is decoupled from content information and general acoustic feature information, and extracts multi-scale timbre features that characterize the macro characteristics and micro details of the user's timbre;
[0025] A customized feature set generation module is used to select and combine the multi-scale timbre features and assign speech synthesis weights according to user configuration to form a customized voiceprint feature set;
[0026] The speech synthesis module is used to embed the pure timbre identity and the customized voiceprint feature set as conditions and inject them into the diffusion flow matching neural vocoder; after receiving the content information to be synthesized, the diffusion flow matching neural vocoder generates a digital human speech synthesis waveform with the user's unique timbre and corresponding to the content through probability stream conversion and step-by-step denoising and refinement.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] 1. By employing a deep voiceprint decoupling network and a multi-task adversarial learning strategy, this invention can extract a pure timbre identity embedding from user voice samples that is significantly decoupled from content information and general acoustic feature information. This overcomes the current problem of timbre features being easily affected by content or other acoustic factors, resulting in impure timbre transfer. This allows the timbre identity of the digital human voice to more accurately match the target user, and the basic timbre is more stable and reliable.
[0029] 2. This invention not only extracts the global timbre identity but also further extracts multi-scale timbre features that characterize both the macroscopic characteristics and microscopic details of the user's timbre. It also allows for the selective combination and weighting of these different levels of features based on user configuration. This overcomes the limitations of existing technologies, which mostly provide global timbre embedding but struggle to independently adjust timbre details. It empowers users to shape the digital human's timbre in a more refined and personalized way, significantly enhancing the richness and customization of timbre performance.
[0030] 3. This invention uses a pure timbre identity embedding and a customized multi-scale voiceprint feature set as conditions, injecting them into an advanced diffusion flow matching neural vocoder for speech synthesis. This vocoder combines the precision of probability stream conversion with the detail restoration capabilities of gradual denoising and refinement, and can more effectively integrate complex timbre characteristics into the final generated speech waveform. Compared with existing technologies, this not only ensures the effective transfer of timbre characteristics, but also significantly improves the overall naturalness, sound quality fidelity, and auditory realism of the synthesized speech, making the digital human voice closer to the real human voice. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a flowchart of the method and system for adaptive digital human voice matching based on voiceprint feature migration proposed in an embodiment of the present invention;
[0032] Figure 2 This is a structural diagram of the deep voiceprint decoupling network proposed in an embodiment of the present invention;
[0033] Figure 3 This is a system structure diagram of the digital human voice adaptive matching method and system based on voiceprint feature migration proposed in an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0035] The present invention provides a method and system for adaptively matching digital human voice based on voiceprint feature migration. For the specific method and system flow chart, please refer to Figure 1 and Figure 3 .
[0036] Example 1
[0037] See Figure 1 The present invention provides a method for adaptively matching digital human timbre based on voiceprint feature migration. The technical solution is as follows:
[0038] A deep voiceprint decoupling network is used to process user voice samples. The deep voiceprint decoupling network uses a multi-task adversarial learning strategy to extract a pure timbre identity embedding from the voice sample that is decoupled from content information and general acoustic feature information, and extracts multi-scale timbre features that characterize the macro characteristics and micro details of the user's timbre;
[0039] According to user configuration, select and combine the multi-scale timbre features and assign speech synthesis weights to form a customized voiceprint feature set;
[0040] The pure timbre identity embedding and customized voiceprint feature set are used as conditions and injected into the diffusion stream matching neural vocoder; after receiving the content information to be synthesized, the diffusion stream matching neural vocoder generates a digital human speech synthesis waveform with the user's unique timbre and corresponding to the content through probability stream conversion and step-by-step denoising and refinement.
[0041] Specifically, Figure 2 As shown, the deep voiceprint decoupling network includes at least one user identity encoding backbone network, at least one content information discriminator and at least one general acoustic feature information discriminator; the general acoustic feature information includes the average fundamental frequency range, basic rhythmic pattern and general acoustic expression of common emotion types.
[0042] This network collaborative working mechanism makes the extracted timbre identity embedding purer and more focused on reflecting the unique voice essence of the target user, providing a high-quality feature foundation for subsequent precise timbre migration and adaptive matching, and helping to generate digital human voices with distinct timbre characteristics that are not easily affected by content or general vocal style.
[0043] The user identity encoding backbone network can be based on advanced speaker recognition model architectures, such as ECAPA-TDNN or a ResNet-based x-vector structure. These networks can effectively extract highly discriminative speaker identity representations from speech. The content information discriminator and general acoustic feature information discriminator can adopt relatively lightweight convolutional neural network (CNN) or recurrent neural network (RNN) structures, which can efficiently classify or regress specific attribute information from a given embedding vector.
[0044] To effectively process universal acoustic feature information, this method first clearly defines and quantifies these features, then uses them as training data to feed the corresponding universal acoustic feature information discriminator. For example, the average fundamental frequency range is obtained by calculating statistics such as the fundamental frequency mean and standard deviation within the effective utterance segment of the speech signal. Basic prosodic patterns are analyzed by analyzing the energy profile, duration information, and pitch curve morphology of the speech, and then combining them with pattern recognition methods to classify them into typical categories. The universal acoustic representation of common emotional types can be obtained by analyzing the statistical distribution or typical patterns of acoustic features corresponding to each emotion (such as pitch, energy, speaking rate, and MFCC) based on a reference dataset with emotion labels, thus forming a universal acoustic profile for that emotion. When training the discriminator, these extracted and quantified universal acoustic features serve as the discriminator's learning targets or labels. The discriminator attempts to predict or identify these universal features from the timbre identity embedding output by the user identity encoding backbone network, thereby guiding the timbre identity embedding to remove these universal acoustic components that are weakly correlated with individual identity but affect the listening experience.
[0045] This technology quantifies fundamental frequency statistics, classifies prosodic patterns, and constructs emotional acoustic profiles to clearly characterize universal acoustic features. Combined with an adversarial training mechanism, it forces the timbre identity embedding to be unable to predict these universal features from the discriminator, thereby stripping away universal acoustic components that are weakly correlated with individual identity. This significantly improves the purity and identity differentiation of the timbre embedding, laying the foundation for high-fidelity timbre transfer.
[0046] Furthermore, the multi-task adversarial learning strategy includes:
[0047] Minimize the user classification loss of timbre identity embedding by training the user identity encoding backbone network;
[0048] Maximizing the accuracy of calculating content information from timbre identity embedding by training the content information discriminator, and minimizing the accuracy of calculating content information from timbre identity embedding by training the user identity encoding backbone network;
[0049] By training the universal acoustic feature information discriminator, the accuracy of calculating the universal acoustic feature information from the timbre identity embedding is maximized, and at the same time, the user identity encoding backbone network is trained to minimize the accuracy of calculating the universal acoustic feature information from the timbre identity embedding.
[0050] This strategy can significantly improve the decoupling degree and representation purity of the extracted timbre identity embedding, making the timbre characteristics more focused on individual uniqueness, and providing key technical support for the subsequent realization of highly realistic and personalized digital human speech synthesis that is not interfered by content or other universal acoustic factors.
[0051] To implement this multi-task adversarial learning strategy, the main loss functions can be defined as follows: the user classification loss usually adopts the standard cross-entropy loss function to measure the prediction accuracy of the timbre identity embedding for the user category; the content information discriminator loss in the content information adversarial loss also adopts the cross-entropy loss to maximize its recognition accuracy of the content information in the embedding (such as phoneme category), while the corresponding backbone network adversarial loss aims to minimize this accuracy, for example, through the gradient reversal layer; similarly, the discriminator loss and the backbone network adversarial loss in the general acoustic feature information adversarial loss can adopt standard forms such as cross-entropy loss or mean square error loss (MSE) according to their specific task types (such as sentiment classification or fundamental frequency range regression), and set opposite optimization goals.
[0052] During model optimization, the total loss of the user identity encoding backbone network is composed of the weighted sum of the aforementioned user classification loss and discriminator loss. The user classification loss weight and the weights of each discriminator loss are adjustable coefficients that balance the contribution of each task to the backbone network parameter update, ensuring a synergistic improvement in decoupling and timbre identity representation. Each discriminator is independently optimized based on its own loss function to continuously enhance its ability to discriminate specific non-timbre identity information.
[0053] The training is iterated using an alternating optimization approach: in one or more training steps, the parameters of the user identity encoding backbone network are first fixed, and the parameters of the content information discriminator and the general acoustic feature information discriminator are updated to minimize their respective losses; then, the parameters of all discriminators are fixed, and the parameters of the user identity encoding backbone network are updated to minimize their total loss.
[0054] This optimization strategy significantly improves the purity and identity distinction of timbre embedding by dynamically balancing the weights of timbre representation and feature decoupling. The alternating training mechanism enables the discriminator and backbone network to mutually reinforce each other, ensuring the effective extraction of common acoustic features while strengthening the ability to extract the essence of individual timbre, laying the foundation for high-quality timbre transfer.
[0055] Furthermore, the multi-scale timbre features include:
[0056] The global timbre features extracted from the deep network layer of the deep voiceprint decoupling network characterize the global static characteristics of the timbre; and the local timbre features extracted from the shallow network layer of the deep voiceprint decoupling network characterize the local dynamic details of the timbre; the global timbre features include the average range and basic sound quality parameters, and the local timbre features include the resonance peak dynamic parameters and the acoustic pattern of the vocalization habit.
[0057] This differentiated extraction of multi-scale features provides a rich and structured information foundation for the subsequent refined, layered control and customization of the digital human's timbre, making it possible not only to match the target user's overall timbre impression, but also to reproduce the more personalized dynamic details and vocal habits in their voice, thereby significantly improving the realism and expressiveness of the synthesized timbre.
[0058] To extract multi-scale timbre features, when the user identity encoding backbone network adopts an architecture such as ResNet, global timbre features can be generated from the activation maps of the deeper residual blocks at the back end of the network through statistical pooling or attention-weighted pooling along the temporal dimension. Local timbre features can be extracted from the activation maps of the relatively shallow residual blocks at the front end of the network. Temporal dynamics can be captured by segmenting the feature maps or further modeling the shallow feature sequences using a temporal convolutional network. If the backbone network is based on the ECAPA-TDNN architecture, its final speaker embedding vector can be directly used as the global timbre feature, while local timbre features can be obtained and aggregated from the outputs of different SE-Res2Block modules within it or from frame-level features before attention pooling, ensuring that structured timbre representations are captured and formed from different levels and time scales.
[0059] The acoustic patterns of specific vocalization habits can be specific types of terminal intonation changes: for example, some speakers tend to use a specific pattern of pitch drops or rises at the end of declarative sentences, or a unique intonation curve in interrogative sentences. Alternatively, the acoustic patterns of specific speech sounds, such as bubbly speech, exhibit very low and irregular fundamental frequencies, with the glottis closed for part of each cycle. These acoustic patterns can manifest in periodic bursts of low-frequency energy, specific shapes in the harmonic structure, and specific traces of Mel-Frequency Cepstral Coefficients (MFCCs).
[0060] Furthermore, the step of selecting a specific combination from the multi-scale timbre features and assigning speech synthesis weights according to user configuration includes:
[0061] Receive user input, including at least one user configuration parameter, the user configuration parameter including the timbre style label of the target digital person, the desired degree of realism, and specific application scenario information; parse the user configuration parameter and map it to a target area in a predefined timbre feature space; preliminarily screen out a feature subset most relevant to the target area from the extracted global timbre features and local timbre features; set initial baseline contribution weights for the screened global timbre feature subset and local timbre feature subset respectively; and dynamically adjust the adaptive speech synthesis weight based on the user configuration parameter to calculate the final speech synthesis weight of each selected timbre feature in the customized voiceprint feature set.
[0062] This method achieves a high degree of customization and flexible control of the digital human's voice, so that the final synthesized voice can not only match the user's core voice, but can also be precisely adjusted according to the style, realism and scene requirements set by the user, greatly improving the personalization of the digital human's voice, adaptability to application scenarios and user satisfaction.
[0063] During the parsing and mapping phase of user-configured parameters, the system first receives user input for a specific timbre style label (such as "sweet" or "magnetic"), the desired level of realism (such as "highly realistic" or "stylized"), and specific application scenario information. These relatively subjective or advanced user instructions are then converted and quantified using built-in parsing logic. For example, a preset rule library can be used to map the "sweet" style to a target region in the timbre feature space corresponding to a higher average fundamental frequency and smoother formant transitions. Alternatively, a pre-trained text encoder can be used to convert the user's style label and scenario description into a configuration embedding vector. The similarity between this vector and the prototype vectors representing different typical timbre regions in the timbre feature space is then calculated to determine the most suitable target region.
[0064] After determining the target area in the timbre feature space, the system will perform feature subset screening accordingly. It selects the global and local timbre features that are most relevant to the target timbre characteristics defined by the current user configuration. For example, the system can pre-define the correlation scores of various multi-scale timbre features (such as average range parameters, formant dynamic parameters, etc.) with various dimensions in the timbre feature space or preset style prototypes; when the user configuration is mapped to a target area, the system will automatically filter out those feature subsets whose correlation scores with the core dimensions in the area are higher than the preset threshold.
[0065] The relevance score is as follows:
[0066] For each user configuration parameter value, its association weight with different timbre feature sub-items is pre-set.
[0067] User configuration parameters (For example, It's a style tag. is the degree of realism);
[0068] Tone Characteristics sub-item (For example, is the average range, is the formant dynamic parameter);
[0069] Preset a weight matrix , indicating configuration parameters A value pair feature The importance or relevance of each feature. There is a baseline contribution weight that is independent of user configuration .
[0070] For label parameters (such as style), you can map them to a set of values. For example, if the style tag If the word “sweet” is sweet, then its corresponding numerical vector may be [0.8, 0.2, ...], indicating its similarity to some predefined style prototypes.
[0071] For continuous parameters (such as the degree of realism, assuming the range is 0-1), they are used directly.
[0072] For each feature , according to user configuration Calculate an adjustment factor This adjustment factor is for Specifically, Is a configuration parameter For example, if the user selects the style "sweet" and the sense of reality "high", then Will combine "sweet" The impact of "high realism" on impact.
[0073] Final relevance score .
[0074] This correlation scoring mechanism dynamically quantifies the mapping between user preferences and timbre characteristics through a weight matrix, converting subjective style labels and continuous parameters into computable adjustment factors. By leveraging the synergy of multi-dimensional configuration parameters, it precisely controls the contribution weights of different timbre characteristics, enabling fine-grained adjustment of timbre style and realism, ensuring that synthesized speech is highly aligned with the user's personalized needs.
[0075] When assigning speech synthesis weights, the system first sets initial baseline contribution weights for the selected global and local timbre feature subsets. These initial weights are based on the average contribution of various timbre features to timbre perception and differentiation, derived from statistical analysis of a large amount of speech data. This ensures a relatively balanced or generally effective initial contribution across all features, absent specific user preferences. Subsequently, the adaptive speech synthesis weights are dynamically adjusted. The core logic of this stage is to fine-tune these initial weights based on specific user-entered configuration parameters (such as style tags, realism, and application scenarios). For example, if a user selects the "Professional Broadcasting" style, the system automatically increases the weights of global timbre features related to articulation clarity and pitch stability, while suppressing the weights of local detail features that may introduce excessive variation. Conversely, if a user selects the "Lively Anime" style, the system significantly increases the weights of local timbre features related to pitch fluctuation, exaggerated formant variations, and specific vocalization habits, and may adjust global timbre features to match the character's vocal range. At the same time, the realism parameter will also guide the weight adjustment towards full reproduction (high realism) or selective exaggeration / simplification (stylization), and finally calculate the final speech synthesis weight of each selected timbre feature in the customized voiceprint feature set to achieve personalized timbre output that is highly in line with user needs.
[0076] Furthermore, the processing of the diffusion flow matching neural vocoder specifically includes:
[0077] Receiving the pure voice identity embedding, the customized voiceprint feature set, and the content information to be synthesized as input conditions; using a conditional probability flow model, mapping and generating a preliminary acoustic representation of predefined Gaussian noise through a reversible transformation, wherein the preliminary acoustic representation is output as a Mel-spectrogram;
[0078] The preliminary acoustic representation is used as the starting state of the diffusion process; a denoising network is initialized, and the denoising network uses the pure timbre identity embedding, the customized voiceprint feature set and the speech content information to be synthesized as conditional inputs; iterative denoising is performed through a plurality of preset time steps, and at each time step, the denoising network gradually removes noise and enhances the details and harmonic structure of the speech based on the acoustic representation of the current time step and the injected conditional information; this denoising step is repeated until a preset termination condition is reached to obtain a denoised acoustic representation; and the denoised acoustic representation is converted into a final digital human speech waveform through a vocoder decoding module.
[0079] This diffuse flow matching neural vocoder fully utilizes the precise mapping capability of the flow model and the detail refinement advantage of the diffuse model. Under the guidance of strong timbre conditions, it can generate digital human speech waveforms with rich details, pure sound quality, and highly restored timbre characteristics of the target user, significantly improving the naturalness and fidelity of the synthesized speech.
[0080] The diffusion flow matching neural vocoder consists of a conditional probability flow generation module and a conditional diffusion denoising module, which work in series and share conditional inputs. The conditional probability flow model module can adopt an architecture based on normalized flow; the denoising network in the conditional diffusion denoising module can have an architecture of U-Net structure or a sequence-to-sequence model based on Transformer. The U-Net structure can effectively capture multi-scale features and reconstruct fine details through its symmetrical encoder-decoder path and jump connections. The Transformer structure, with its self-attention mechanism, has advantages in modeling long sequence dependencies and is suitable for processing acoustic feature sequences.
[0081] The preliminary acoustic representation is designed as a structured intermediate representation that contains core speech information. Specifically, it is a mel-spectrogram that is temporally aligned with the input content information (such as a phoneme sequence). The frame rate and frequency dimensions adopt common standards in the field of speech synthesis. It encodes basic speech content, prosodic contours, and macro-timbre characteristics guided by a clean voice identity embedding and a customized voiceprint feature set.
[0082] This embodiment uses a deep voiceprint decoupling network and a multi-task adversarial learning strategy to extract pure timbre identity embedding and multi-scale timbre features from user voices; customizes and weights the multi-scale timbre features based on user configuration; injects these features into a diffusion flow matching neural vocoder, and combines them with content information to generate a personalized digital human voice waveform. This can generate a digital human voice with a highly personalized timbre that is highly consistent with the timbre characteristics of the target user, while ensuring high naturalness and high fidelity of the synthesized speech. Through refined control of the macroscopic characteristics and microscopic details of the timbre, as well as the introduction of user-defined configurations, the flexibility and expressiveness of the digital human's timbre shaping are greatly enhanced, enabling it to better adapt to different application scenarios and user needs, significantly improving the realism and user experience of the digital human as a carrier of human-computer interaction and virtual content presentation.
[0083] Example 2
[0084] This embodiment further describes an operation mode of the digital human voice adaptive matching system based on voiceprint feature migration in actual application scenarios, such as Figure 3 The system can be integrated or used in various applications or platforms that require personalized digital human voice.
[0085] In a typical application process, an external application (hereinafter referred to as the "caller") interacts with this system to achieve adaptive matching of the user's timbre and speech synthesis:
[0086] Voiceprint feature processing module:
[0087] The caller first collects a voice sample from the target user and related user profile information. This profile may include parameters such as the desired specific voice style label, level of realism, or application scenario. Upon receiving the voice sample, the internal deep voiceprint decoupling network is activated. Using a multi-task adversarial learning strategy, it extracts a clean voice identity embedding and multi-scale voice features from the voice sample. This data is then sent to the customized feature set generation module via a pre-defined interface.
[0088] Customized feature set generation module:
[0089] After receiving user configuration information and multi-scale timbre features, a customized voiceprint feature set that can reflect the user's specific needs is formed according to preset algorithms or rules.
[0090] Speech synthesis module:
[0091] The caller sends the text content (or its processed linguistic feature representation) to be read aloud by the digital human to the speech synthesis module of this system. This module also receives the clean voice identity embedding and customized voiceprint feature set generated by the previous steps.
[0092] The diffusion-matched neural vocoder within the speech synthesis module uses the received pure voice identity embedding and customized voiceprint feature set as the core timbre control conditions, combined with the content information to be synthesized. Subsequently, through a process of probabilistic stream conversion and gradual denoising and refinement, a digital human speech synthesis waveform is synthesized with the target user's unique timbre (or a timbre adjusted according to the user's configuration) and accurately corresponds to the input content.
[0093] The resulting speech waveform data is returned to the caller via an interface. The caller can then use this high-quality, personalized digital human voice in specific application scenarios, such as integrating it into video content, using it as a response voice for a virtual assistant, or using it in a real-time interactive digital human system, significantly enhancing the realism, personality, and user experience of their digital human character.
[0094] This example shows Figure 3 The system shown here serves as a fully functional back-end service, providing core digital human voice adaptive matching and speech synthesis capabilities for various front-end applications by receiving input data, collaborative processing among internal modules, and outputting the final results. It fully embodies the extraction, migration, customization of voiceprint features, and their ultimate adaptive application in speech synthesis.
[0095] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A digital human voice adaptive matching method based on voiceprint feature migration, characterized in that: The following steps are involved: A deep voiceprint decoupling network is used to process user voice samples. The deep voiceprint decoupling network uses a multi-task adversarial learning strategy to extract a pure timbre identity embedding from the voice sample that is decoupled from content information and general acoustic feature information, and extracts multi-scale timbre features that characterize the macro characteristics and micro details of the user's timbre; According to user configuration, select and combine the multi-scale timbre features and assign speech synthesis weights to form a customized voiceprint feature set; The pure timbre identity embedding and customized voiceprint feature set are used as conditions and injected into the diffusion stream matching neural vocoder; after receiving the content information to be synthesized, the diffusion stream matching neural vocoder generates a digital human speech synthesis waveform with the user's unique timbre and corresponding to the content through probability stream conversion and step-by-step denoising and refinement.
2. The method for adaptively matching digital human timbre based on voiceprint feature migration according to claim 1, characterized in that: The deep voiceprint decoupling network includes at least one user identity encoding backbone network, at least one content information discriminator and at least one general acoustic feature information discriminator; the general acoustic feature information includes the average fundamental frequency range, basic rhythmic pattern and general acoustic expression of common emotion types.
3. The method for adaptively matching digital human timbre based on voiceprint feature migration according to claim 1, characterized in that: The multi-task adversarial learning strategy includes: Minimize the user classification loss of timbre identity embedding by training the user identity encoding backbone network; Maximizing the accuracy of calculating content information from timbre identity embedding by training the content information discriminator, and minimizing the accuracy of calculating content information from timbre identity embedding by training the user identity encoding backbone network; By training the universal acoustic feature information discriminator, the accuracy of calculating the universal acoustic feature information from the timbre identity embedding is maximized, and at the same time, the user identity encoding backbone network is trained to minimize the accuracy of calculating the universal acoustic feature information from the timbre identity embedding.
4. The method for adaptively matching digital human timbre based on voiceprint feature migration according to claim 1, characterized in that: The multi-scale timbre features include: The global timbre features extracted from the deep network layer of the deep voiceprint decoupling network characterize the global static characteristics of the timbre; and the local timbre features extracted from the shallow network layer of the deep voiceprint decoupling network characterize the local dynamic details of the timbre; the global timbre features include the average range and basic sound quality parameters, and the local timbre features include the resonance peak dynamic parameters and the acoustic pattern of the vocalization habit.
5. The method for adaptively matching digital human timbre based on voiceprint feature migration according to claim 1, characterized in that: The step of selecting a combination from the multi-scale timbre features according to user configuration and assigning speech synthesis weights includes: Receive user input, including at least one user configuration parameter, the user configuration parameter including the timbre style label of the target digital person, the desired degree of realism, and specific application scenario information; parse the user configuration parameter and map it to a target area in a predefined timbre feature space; preliminarily screen out a feature subset most relevant to the target area from the extracted global timbre features and local timbre features; set initial baseline contribution weights for the screened global timbre feature subset and local timbre feature subset respectively; and dynamically adjust the adaptive speech synthesis weight based on the user configuration parameter to calculate the final speech synthesis weight of each selected timbre feature in the customized voiceprint feature set.
6. The method for adaptively matching digital human timbre based on voiceprint feature migration according to claim 1, characterized in that: The processing process of the diffusion flow matching neural vocoder specifically includes: Receiving the pure voice identity embedding, the customized voiceprint feature set, and the content information to be synthesized as input conditions; using a conditional probability flow model, mapping and generating a preliminary acoustic representation of predefined Gaussian noise through a reversible transformation, wherein the preliminary acoustic representation is output as a Mel-spectrogram; The preliminary acoustic representation is used as the starting state of the diffusion process; a denoising network is initialized, and the denoising network takes the pure timbre identity embedding, the customized voiceprint feature set and the content information to be synthesized as conditional inputs; iterative denoising is performed through a plurality of preset time steps, and at each time step, the denoising network gradually removes noise and enhances the details and harmonic structure of the speech based on the acoustic representation of the current time step and the injected conditional information; this denoising step is repeated until a preset termination condition is reached to obtain a denoised acoustic representation; and the denoised acoustic representation is converted into a final digital human speech waveform through a vocoder decoding module.
7. The digital human voice adaptive matching system based on voiceprint feature migration is characterized by: The system comprises: A voiceprint feature processing module, configured to process user voice samples using a deep voiceprint decoupling network. The deep voiceprint decoupling network, through a multi-task adversarial learning strategy, extracts a pure timbre identity embedding from the voice sample that is decoupled from content information and general acoustic feature information, and extracts multi-scale timbre features that characterize the macroscopic characteristics and microscopic details of the user's timbre; A customized feature set generation module is used to select and combine the multi-scale timbre features and assign speech synthesis weights according to user configuration to form a customized voiceprint feature set; The speech synthesis module is used to embed the pure timbre identity and the customized voiceprint feature set as conditions and inject them into the diffusion flow matching neural vocoder; after receiving the content information to be synthesized, the diffusion flow matching neural vocoder generates a digital human speech synthesis waveform with the user's unique timbre and corresponding to the content through probability stream conversion and step-by-step denoising and refinement.
Citation Information
Patent Citations
Voice style migration system and method for tone and style deep decoupling
CN117912446A
Tone generation method based on voice conversion
CN118197329A