Digital human tone adaptive matching method and system based on voiceprint feature migration
Through the combination of deep voiceprint decoupling network and diffusion stream matching neural vocoder, the combined tone characteristics are extracted and customized, which solves the limitations of adaptive tone matching in the existing technology, and realizes highly realistic and personalized tone synthesis, which enhances the naturalness and reality of digital human voice.
Patent Information
- Application Number
- CN202510779082.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-12
AI Technical Summary
The existing voice synthesis technology has limitations in achieving highly realistic, fine, controllable and pure digital human tone adaptive matching, and it is difficult to make independent and fine adjustments and user-defined combinations of tone details, limiting the flexibility and personalization of tone shaping.
A deep voiceprint decoupling network is used to extract pure tone identity embedding and multi-scale tone characteristics from the speech samples through a multi-task adversarial learning strategy, and customized combinations and speech synthesis are combined with diffusion stream matching neural vocoder to generate digital human voices with user unique tone and corresponding to the content.
It realizes accurate and flexible personalized tone synthesis, improves the naturalness, sound quality fidelity and auditory realism of digital human voice, and enhances the richness and customization of tone performance.
Smart Images

Figure CN120356474A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech synthesis, and specifically to a method and system for digital human voice color adaptive matching based on voiceprint feature migration. Background Art
[0002] As an important carrier for human-computer interaction and virtual content presentation, the realism and expressiveness of digital humans are crucial for the user experience. Among many perception dimensions, speech synthesis technology endows digital humans with the ability to "speak", and the voice color of the synthesized speech is a key factor in shaping the unique personality, identity, and emotion of digital humans. Current speech synthesis technologies, especially text-to-speech systems based on deep neural networks, have made remarkable progress in improving the naturalness and similarity of synthesized speech. Mainstream methods extract the voiceprint features of the target user by introducing a user encoder and incorporate them as conditional information into acoustic models (such as Tacotron, FastSpeech, VITS, etc.) or neural vocoders to generate speech with the specific voice color of the user. These methods are committed to learning and migrating voice color information from target speech samples to achieve personalized speech synthesis and voice cloning. Some studies also focus on feature decoupling technology, attempting to separate voice color features from other acoustic attributes such as speech content and prosody during the synthesis process, in order to obtain a more pure voice color representation for controlling speech synthesis.
[0003] Although existing speech synthesis technologies have achieved certain results in voice color control, there are still limitations in achieving highly realistic, finely controllable, and pure digital human voice color adaptive matching. The existing methods have limited control accuracy for voice color details, and most provide global voice color embeddings, making it difficult to independently and finely adjust the macroscopic characteristics and microscopic details of the voice color and customize combinations by users, which limits the flexibility and personalization of voice color shaping.
[0004] Therefore, a method and system for digital human voice color adaptive matching based on voiceprint feature migration are proposed. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for digital human voice color adaptive matching, which realizes precise and flexible personalized voice color synthesis and improves the realism and expressiveness by deeply decoupling and customizing the combination of voiceprint features and combining with an advanced vocoder.
[0006] To achieve the above purpose, the present invention provides the following technical solutions: A method for digital human voice color adaptive matching based on voiceprint feature migration, including: A deep voiceprint decoupling network is used to process the user's voice samples. The deep voiceprint decoupling network extracts a pure tone color identity embedding decoupled from the content information and general acoustic feature information from the voice samples through a multi-task adversarial learning strategy, and extracts multi-scale tone color features for characterizing the macroscopic characteristics and microscopic details of the user's tone color; According to the user configuration, a combination is selected from the multi-scale tone color features and voice synthesis weights are assigned to form a customized voiceprint feature set; Taking the pure tone color identity embedding and the customized voiceprint feature set as conditions, they are injected into a diffusion flow matching neural vocoder; after receiving the content information to be synthesized, the diffusion flow matching neural vocoder generates a digital human voice synthesis waveform with the user's unique tone color and corresponding to the content through probability flow conversion and step-by-step denoising refinement.
[0007] Preferably, the deep voiceprint decoupling network includes at least one user identity encoding backbone network, at least one content information discriminator, and at least one general acoustic feature information discriminator; the general acoustic feature information includes the average fundamental frequency range, the basic prosody pattern, and the general acoustic performance of common emotional types.
[0008] The multi-task adversarial learning strategy includes: By training the user identity encoding backbone network, minimizing the user classification loss of the tone color identity embedding; By training the content information discriminator, maximizing the accuracy of calculating the content information from the tone color identity embedding, and at the same time training the user identity encoding backbone network, minimizing the accuracy of calculating the content information from the tone color identity embedding; By training the general acoustic feature information discriminator, maximizing the accuracy of calculating the general acoustic feature information from the tone color identity embedding, and at the same time training the user identity encoding backbone network, minimizing the accuracy of calculating the general acoustic feature information from the tone color identity embedding.
[0009] Preferably, the multi-scale tone color features include: Global tone color features extracted from the deep network layer of the deep voiceprint decoupling network, which characterize the global static characteristics of the tone color; and local tone color features extracted from the shallow network layer of the deep voiceprint decoupling network, which characterize the local dynamic details of the tone color; the global tone color features include the average pitch range and the base tone quality parameters, and the local tone color features include the formant dynamic parameters and the acoustic patterns of the vocalization habits.
[0010] Preferably, the step of selecting a combination from the multi-scale tone color features and assigning voice synthesis weights according to the user configuration includes: Receive user input, which includes at least one user configuration parameter. The user configuration parameter includes the timbre style label of the target digital human, the desired degree of realism, and specific application scenario information; parse the user configuration parameter and map it to the target area in the predefined timbre feature space; initially screen out the feature subset most relevant to the target area from the extracted global timbre features and local timbre features; set initial benchmark contribution weights for the selected global timbre feature subset and local timbre feature subset respectively; and perform adaptive voice synthesis weight dynamic adjustment based on the user configuration parameter to calculate the final voice synthesis weights of each selected timbre feature in the customized voiceprint feature set.
[0011] Preferably, the processing process of the diffusion flow matching neural vocoder specifically includes: Receive the pure timbre identity embedding, the customized voiceprint feature set, and the content information to be synthesized as input conditions; use the conditional probability flow model to map and generate a preliminary acoustic representation by reversibly transforming the predefined Gaussian noise, and the preliminary acoustic representation is output as a Mel spectrogram. Use the preliminary acoustic representation as the starting state of the diffusion process; initialize a denoising network, which takes the pure timbre identity embedding, the customized voiceprint feature set, and the content information of the to-be-synthesized speech as conditional inputs; perform iterative denoising through a preset number of time steps. At each time step, the denoising network gradually removes noise and enhances the detail and harmonic structure of the speech according to the acoustic representation of the current time step and the injected conditional information; repeat this denoising step until the preset termination condition is reached to obtain the denoised acoustic representation; convert the denoised acoustic representation into the final digital human speech waveform through a vocoder decoding module.
[0012] On the other hand, a digital human timbre adaptive matching system based on voiceprint feature migration includes: A voiceprint feature processing module for processing the user's voice samples using a deep voiceprint decoupling network. The deep voiceprint decoupling network extracts a pure timbre identity embedding decoupled from the content information and general acoustic feature information from the voice samples through a multi-task adversarial learning strategy, and extracts multi-scale timbre features for characterizing the macroscopic characteristics and microscopic details of the user's timbre. A customized feature set generation module for selecting, combining, and assigning voice synthesis weights from the multi-scale timbre features according to the user configuration to form a customized voiceprint feature set. A voice synthesis module, which is used to inject the pure tone identity embedding and customized voiceprint feature set as conditions into a diffusion flow matching neural vocoder; after receiving the content information to be synthesized, the diffusion flow matching neural vocoder generates a digital human voice synthesis waveform with the user's unique tone and corresponding to the content through probability flow conversion and step-by-step denoising refinement.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By adopting a deep voiceprint decoupling network and a multi-task adversarial learning strategy, the present invention can extract a pure tone identity embedding that is significantly decoupled from content information and general acoustic feature information from user voice samples. It overcomes the problem that current tone features are easily interfered by content or other acoustic factors, resulting in impure tone migration, so that the tone identity of digital human voices can more accurately match the target user, and the basic tone is more stable and reliable.
[0014] 2. The present invention not only extracts global tone identities, but also further extracts multi-scale tone features that characterize the macroscopic characteristics and microscopic details of the user's tone, and allows selective combination and weighting of these different levels of features according to user configuration. This breaks through the limitation of most prior arts that only provide overall tone embeddings and it is difficult to independently adjust tone details, endows users with the ability to more finely and personalized shape the tone of digital humans, and significantly enhances the richness and customization of tone performance.
[0015] 3. The present invention injects the pure tone identity embedding and customized multi-scale voiceprint feature set as conditions into an advanced diffusion flow matching neural vocoder for voice synthesis. This vocoder combines the accuracy of probability flow conversion and the detail restoration ability of step-by-step denoising refinement, and can more effectively integrate complex tone features into the finally generated voice waveform. Compared with the prior art, this not only ensures the effective migration of tone features, but also significantly improves the overall naturalness, sound quality fidelity and auditory realism of the synthesized voice, making the voice of the digital human closer to real human voices. Description of the Drawings
[0016] Figure 1 It is a method flow chart of a digital human tone adaptive matching method and system based on voiceprint feature migration proposed by an embodiment of the present invention; Figure 2 It is a structural diagram of a deep voiceprint decoupling network proposed by an embodiment of the present invention; Figure 3 It is a system structural diagram of a digital human tone adaptive matching method and system based on voiceprint feature migration proposed by an embodiment of the present invention. Detailed Embodiments
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0018] The present invention provides a method and system for adaptive matching of digital human voice tones based on voiceprint feature migration. The flowcharts of the specific method and system are referred to Figure 1 and Figure 3 .
[0019] Embodiment 1 Referring to Figure 1 , the present invention provides a method for adaptive matching of digital human voice tones based on voiceprint feature migration. The technical solution is as follows: A deep voiceprint decoupling network is used to process the user's voice samples. The deep voiceprint decoupling network extracts a pure voice tone identity embedding decoupled from the content information and general acoustic feature information from the voice samples through a multi-task adversarial learning strategy, and extracts multi-scale voice tone features for characterizing the macroscopic characteristics and microscopic details of the user's voice tone; According to the user configuration, a customized voiceprint feature set is formed by selecting combinations from the multi-scale voice tone features and assigning voice synthesis weights; Taking the pure voice tone identity embedding and the customized voiceprint feature set as conditions, they are injected into a diffusion flow matching neural vocoder; after receiving the content information to be synthesized, the diffusion flow matching neural vocoder generates a digital human voice synthesis waveform with the user's unique voice tone and corresponding to the content through probability flow conversion and step-by-step denoising refinement.
[0020] Specifically, as Figure 2 shown, the deep voiceprint decoupling network includes at least one user identity encoding backbone network, at least one content information discriminator, and at least one general acoustic feature information discriminator; the general acoustic feature information includes the average fundamental frequency range, basic prosody pattern, and general acoustic manifestations of common emotional types.
[0021] This network collaborative working mechanism makes the extracted voice tone identity embedding more pure, more focused on reflecting the unique voice essence of the target user, provides a high-quality feature basis for subsequent accurate voice tone migration and adaptive matching, and helps to generate digital human voices with distinct voice tone features and not easily affected by content or general vocal styles.
[0022] The user identity encoding backbone network can be based on advanced speaker recognition model architectures, such as ECAPA-TDNN or ResNet-based x-vector structures, which can effectively extract highly discriminative speaker identity representations from speech. The content information discriminator and the general acoustic feature information discriminator can adopt relatively lightweight convolutional neural network (CNN) or recurrent neural network (RNN) structures, which can efficiently classify or regress specific attribute information from the given embedding vectors.
[0023] To achieve effective processing of general acoustic feature information, this method first clearly defines and quantifies the extraction of these features, and then inputs them as training data into the corresponding general acoustic feature information discriminator. For example: The average fundamental frequency range is obtained by calculating statistics such as the mean and standard deviation of the fundamental frequency within the effective vocalization segment of the speech signal; the basic prosodic pattern is obtained by analyzing the energy contour, duration information, and pitch curve shape of the speech, and combining pattern recognition methods to classify it into typical categories; for the general acoustic manifestations of common emotion types, based on a reference data set with emotion labels, the statistical distribution or typical pattern of the acoustic features (such as pitch, energy, speech rate, MFCC, etc.) corresponding to each emotion can be analyzed to form the general acoustic portrait of this emotion. When training the discriminator, these extracted and quantified general acoustic feature information will be used as the learning target or label of the discriminator. The discriminator needs to try to predict or identify these general features from the timbre identity embedding output by the user identity encoding backbone network, so as to guide the timbre identity embedding to remove these general acoustic components that are weakly related to the individual identity but affect the listening perception.
[0024] This technology defines the data representation of general acoustic features through quantifying fundamental frequency statistics, prosodic pattern classification, and constructing emotional acoustic portraits. Combining with the adversarial training mechanism, it forces the timbre identity embedding to be unable to be predicted by the discriminator for these general features, thereby stripping the general acoustic components that are weakly related to the individual identity, significantly improving the purity and identity distinctiveness of the timbre embedding, and laying a foundation for high-fidelity timbre transfer.
[0025] Furthermore, the multi-task adversarial learning strategy includes: By training the user identity encoding backbone network, minimizing the user classification loss of the timbre identity embedding; By training the content information discriminator, maximizing the accuracy of calculating content information from the timbre identity embedding, and at the same time training the user identity encoding backbone network, minimizing the accuracy of calculating content information from the timbre identity embedding; By training the general acoustic feature information discriminator, maximizing the accuracy of calculating general acoustic feature information from the timbre identity embedding, and at the same time training the user identity encoding backbone network, minimizing the accuracy of calculating general acoustic feature information from the timbre identity embedding.
[0026] This strategy can significantly improve the decoupling degree and representation purity of the extracted timbre identity embeddings, making the timbre features more focused on individual uniqueness, and providing key technical support for subsequent realization of highly realistic personalized digital human speech synthesis that is not interfered by content or other general acoustic factors.
[0027] To specifically implement this multi-task adversarial learning strategy, the main loss functions can be defined as follows: The user classification loss usually adopts the standard cross-entropy loss function to measure the prediction accuracy of the timbre identity embedding for user categories; the discriminator loss in the content information adversarial loss also uses the cross-entropy loss to maximize its recognition accuracy of the content information (such as phoneme categories) in the embedding, while the corresponding backbone network adversarial loss aims to minimize this accuracy, for example, through a gradient reversal layer; similarly, the discriminator loss and the backbone network adversarial loss in the general acoustic feature information adversarial loss can respectively adopt standard forms such as cross-entropy loss or mean squared error loss (MSE) according to their specific task types (such as emotion classification or fundamental frequency range regression), and set opposite optimization goals.
[0028] During the model optimization process, the total loss of the user identity encoding backbone network is composed of the above-mentioned user classification loss and discriminator loss weighted, where the weights of the user classification loss and each discriminator loss are adjustable weight coefficients, used to balance the contributions of each task to the update of the backbone network parameters, and ensure the coordinated improvement of the decoupling effect and the timbre identity representation ability. Each discriminator is independently optimized according to its own loss function to continuously enhance its discriminative ability for specific non-timbre identity information.
[0029] The training is carried out iteratively in an alternating optimization manner: in one or more training steps, first fix the parameters of the user identity encoding backbone network, and update the parameters of the content information discriminator and the general acoustic feature information discriminator to minimize their respective losses; then, fix the parameters of all discriminators and update the parameters of the user identity encoding backbone network to minimize its total loss.
[0030] This optimization strategy significantly improves the purity and identity distinctiveness of the timbre embedding by dynamically balancing the weights of timbre representation and feature decoupling. The alternating training mechanism promotes the discriminator and the backbone network to reinforce each other, ensuring the effective stripping of general acoustic features and strengthening the extraction ability of individual timbre essence, laying a foundation for high-quality timbre transfer.
[0031] Furthermore, the multi-scale timbre features include: Global timbre features extracted from the deep network layer of the deep voiceprint decoupling network, representing the global static characteristics of timbre; and local timbre features extracted from the shallow network layer of the deep voiceprint decoupling network, representing the local dynamic details of timbre; the global timbre features include the average pitch range and the base timbre parameters, and the local timbre features include the formant dynamic parameters and the acoustic patterns of vocal habits.
[0032] This differentiated extraction of multi-scale features provides a rich and structured information basis for subsequent realization of refined and hierarchical control and customization of the digital human timbre, enabling not only the matching of the overall timbre impression of the target user, but also the reproduction of more personalized dynamic details and vocal habits in their voices, thus significantly enhancing the realism and expressiveness of the synthesized timbre.
[0033] To achieve the extraction of multi-scale timbre features, when the user identity encoding backbone network adopts an architecture such as ResNet, the global timbre features can be generated from the activation maps of deeper residual blocks at the backend of the network by performing statistical pooling or attention-weighted pooling on the time dimension; while the local timbre features can be extracted from the activation maps of relatively shallower residual blocks at the front end of the network, and the feature maps are segmented or the shallow feature sequences are further modeled using a temporal convolutional network to capture the time dynamics. If the backbone network is based on the ECAPA-TDNN architecture, its final speaker embedding vector can be directly used as the global timbre feature, and the local timbre features can be obtained and aggregated from the outputs of different SE-Res2Block modules or the frame-level features before attention pooling inside it, thus ensuring the capture and formation of structured timbre representations from different levels and time scales.
[0034] The acoustic patterns of specific vocal habits can be specific types of intonation end variations: for example, some speakers are accustomed to using specific pitch drop or rise patterns at the end of declarative sentences, or using unique intonation curves in interrogative sentences. Or the acoustic patterns of special voices, such as vocal fry, which is characterized by a very low and irregular fundamental frequency, and the glottis is partially closed for some time in each cycle. Its acoustic patterns may be reflected in the periodic bursts of low-frequency energy, the special morphology of the harmonic structure, and the specific trajectories of Mel Frequency Cepstral Coefficients (MFCCs).
[0035] Furthermore, the step of selecting a specific combination from the multi-scale timbre features according to the user configuration and assigning speech synthesis weights includes: Receive user input, which includes at least one user configuration parameter. The user configuration parameter includes the timbre style label of the target digital human, the desired degree of realism, and specific application scenario information; Parse the user configuration parameter and map it to the target area in the predefined timbre feature space; Initially screen out the feature subset most relevant to the target area from the extracted global timbre features and local timbre features; Set initial benchmark contribution weights for the selected global timbre feature subset and local timbre feature subset respectively; And perform adaptive speech synthesis weight dynamic adjustment based on the user configuration parameter, and calculate the final speech synthesis weights of each selected timbre feature in the customized voiceprint feature set.
[0036] This method realizes highly customized and flexible control of the digital human's timbre, making the finally synthesized voice not only able to match the user's core timbre, but also can be precisely adjusted according to the style, realism and scenario requirements set by the user, greatly improving the personalization degree, application scenario adaptability and user satisfaction of the digital human's voice.
[0037] In the stage of parsing and mapping the user configuration parameter, the system first receives the specific timbre style label input by the user (such as "sweet" or "magnetic"), the desired degree of realism (such as "highly realistic" or "stylized"), and specific application scenario information. Subsequently, these relatively subjective or high-level user instructions are converted and quantified through the built-in parsing logic: For example, the "sweet" style can be mapped to the target area in the timbre feature space corresponding to a relatively high average fundamental frequency and a smoother formant transition by using a preset rule library; Or, use a pre-trained text encoder to convert the user's style label and scenario description into a configuration embedding vector, and calculate the similarity between this vector and the prototype vectors representing different typical timbre areas in the timbre feature space, so as to determine the most matching target area.
[0038] After determining the target area in the timbre feature space, the system will perform feature subset screening accordingly. Select the global timbre features and local timbre features that are most relevant to the target timbre characteristics defined by the current user configuration. For example, the system can pre-define the correlation score of various multi-scale timbre features (such as average pitch range parameters, formant dynamic parameters, etc.) with each dimension or preset style prototype in the timbre feature space; When the user configuration is mapped to a certain target area, the system will automatically screen out the feature subset whose correlation score with the core dimension in this area is higher than the preset threshold.
[0039] Among them, the correlation score is specifically: For each value of the user configuration parameter, preset its correlation weight with different timbre feature sub-items.
[0040] User configuration parameter (For example, is a style label, is the degree of realism); sub-item of timbre feature (for example, is the average pitch range, is the formant dynamic parameter); Preset a weight matrix , indicating a certain value of the configuration parameter for the importance or relevance of the feature . Each feature has a baseline contribution weight that does not depend on user configuration .
[0041] For label-type parameters (such as style), they can be mapped to a set of numerical values. For example, if the style label is "sweet", then its corresponding numerical vector may be [0.8, 0.2,...], indicating its similarity to some predefined style prototypes.
[0042] For continuous parameters (such as the degree of realism, assuming a range of 0 - 1), they are directly used.
[0043] For each feature , calculate an adjustment factor according to the user configuration . This adjustment factor is a certain aggregation of . Specifically, is the numerical representation of the configuration parameter ; for example, if the user selects the style "sweet" and the realism "high", then will synthesize the influence of "sweet" on and the influence of "high realism" on .
[0044] Final relevance score .
[0045] This relevance scoring mechanism dynamically quantifies the mapping relationship between user preferences and timbre features through a weight matrix, unifying subjective style labels and continuous parameters into computable adjustment factors. Based on the synergistic effect of multi-dimensional configuration parameters, it precisely controls the contribution weights of different timbre features, realizes fine-grained adjustment of timbre style and degree of realism, and ensures that the synthesized speech highly meets the user's personalized needs.
[0046] When allocating speech synthesis weights, the system first sets initial benchmark contribution weights for the selected global timbre feature subset and local timbre feature subset. These initial weights are based on the average contribution degrees of various timbre features to timbre perception and distinctiveness obtained from statistical analysis of a large amount of speech data, ensuring that each feature has a relatively balanced or generally effective initial contribution in the absence of specific user preferences. Subsequently, it enters the adaptive speech synthesis weight dynamic adjustment stage. The core logic of this stage is to finely adjust these initial weights according to the specific configuration parameters input by the user (such as style labels, degree of realism, application scenarios). For example, if the user selects the "professional broadcasting" style, the system will automatically increase the weights of global timbre features related to pronunciation clarity and pitch stability, while suppressing the weights of local detail features that may introduce excessive variations. Conversely, if the "lively anime" style is selected, the weights of local timbre features related to pitch fluctuations, exaggerated formant variations, and specific vocal habits will be significantly increased, and the global timbre features may be adjusted to match the vocal range setting of the character. At the same time, the realism degree parameter will also guide the weights to adjust in the direction of full reproduction (high realism) or selective exaggeration / simplification (stylization). Finally, the final speech synthesis weights of each selected timbre feature in the customized voiceprint feature set are calculated to achieve a personalized timbre output that highly conforms to the user's needs.
[0047] Furthermore, the processing process of the diffusion flow matching neural vocoder specifically includes: Receiving the pure timbre identity embedding, the customized voiceprint feature set, and the content information to be synthesized as input conditions; using the conditional probability flow model to map and generate a preliminary acoustic representation by reversibly transforming a predefined Gaussian noise, and the preliminary acoustic representation is output as a Mel spectrogram; Taking the preliminary acoustic representation as the starting state of the diffusion process; initializing a denoising network, which takes the pure timbre identity embedding, the customized voiceprint feature set, and the content information of the speech to be synthesized as conditional inputs; performing iterative denoising through a preset number of time steps. At each time step, the denoising network gradually removes noise and enhances the detail and harmonic structure of the speech according to the acoustic representation at the current time step and the injected conditional information; repeating this denoising step until a preset termination condition is reached to obtain a denoised acoustic representation; converting the denoised acoustic representation into the final digital human speech waveform through a vocoder decoding module.
[0048] This diffusion flow matching neural vocoder makes full use of the accurate mapping ability of the flow model and the detail refinement advantage of the diffusion model, and can generate a digital human speech waveform with rich details, pure sound quality, and highly restored target user timbre characteristics under the guidance of strong timbre conditions, significantly improving the naturalness and fidelity of the synthesized speech.
[0049] The diffusion flow matching neural vocoder includes a conditional probability flow generation module and a conditional diffusion denoising module, which work in series and share the conditional input. The conditional probability flow model module can adopt an architecture based on normalizing flow; the denoising network in the conditional diffusion denoising module can have an architecture of U-Net structure or a sequence-to-sequence model based on Transformer. The U-Net structure can effectively capture multi-scale features and perform fine-grained detail reconstruction through its symmetric encoder-decoder path and skip connections. The Transformer structure, with its self-attention mechanism, has an advantage in modeling long-range sequence dependencies and is suitable for processing acoustic feature sequences.
[0050] The initial acoustic representation is designed in specifications as a structured intermediate representation containing core speech information. Specifically, the initial acoustic representation is a Mel spectrogram temporally aligned with the input content information (such as a phoneme sequence), with the frame rate and frequency dimension both adopting common standards in the field of speech synthesis, encoding basic speech content, prosodic contours, and macroscopic timbre characteristics guided by a pure timbre identity embedding and a customized voiceprint feature set.
[0051] In this embodiment, by using a deep voiceprint decoupling network and a multi-task adversarial learning strategy, a pure timbre identity embedding and multi-scale timbre features are extracted from the user's speech; the multi-scale timbre features are customized and weighted according to the user's configuration; these features are injected into the diffusion flow matching neural vocoder to generate a personalized digital human speech waveform in combination with the content information. It can generate digital human speech with highly personalized timbre and highly consistent with the target user's timbre characteristics, while ensuring high naturalness and high fidelity of the synthesized speech; through the refined control of the macroscopic characteristics and microscopic details of the timbre, and the introduction of user-defined configurations, the flexibility and expressiveness of digital human timbre shaping are greatly enhanced, enabling it to better adapt to different application scenarios and user needs, and significantly improving the realism and user experience of the digital human as a carrier for human-computer interaction and virtual content presentation.
[0052] Embodiment 2 This embodiment further elaborates on an operation mode of the digital human timbre adaptive matching system based on voiceprint feature transfer in an actual application scenario, such as Figure 3 shown. The system can be integrated or invoked in various application programs or platforms that require personalized digital human speech.
[0053] In a typical application process, an external application program (hereinafter referred to as the "caller") interacts with this system to achieve adaptive matching of the user's timbre and speech synthesis: Voiceprint feature processing module: The caller first collects the target user's voice sample and related user configuration information. The user configuration information may include parameters such as the desired specific timbre style label, degree of realism or application scenario. After receiving the voice sample, the internal deep voiceprint decoupling network is started to extract pure timbre identity embedding and multi-scale timbre features from the voice sample through a multi-task adversarial learning strategy. These data are sent to the customized feature set generation module through a preset interface.
[0054] Customized feature set generation module: After receiving user configuration information and multi-scale timbre features, a customized voiceprint feature set that can reflect the user's specific needs is formed according to preset algorithms or rules.
[0055] Speech synthesis module: The caller sends the text content (or its processed linguistic feature representation) that needs to be read by the digital human to the speech synthesis module of this system. The module also receives the pure voice identity embedding and customized voiceprint feature set generated by the previous steps.
[0056] The diffusion flow matching neural vocoder inside the speech synthesis module embeds the received pure timbre identity and customized voiceprint feature set as the core timbre control conditions, and combines the content information to be synthesized. Subsequently, through the process of probability stream conversion and step-by-step denoising and refinement, a digital human speech synthesis waveform with the target user's unique timbre (or the timbre adjusted according to the user's configuration) and accurately corresponding to the input content is synthesized.
[0057] The final generated speech waveform data is returned to the caller through the interface. The caller can use this high-quality, personalized digital human voice in its specific application scenarios, such as integrating it into video content, as a response voice for a virtual assistant, or using it in a real-time interactive digital human system, thereby significantly improving the realism, personality, and user experience of its digital human character.
[0058] This example demonstrates Figure 3 The system shown here serves as a fully functional backend service, providing core digital human voice adaptive matching and speech synthesis capabilities for various front-end applications by receiving input data, coordinating internal module processing, and outputting final results. It fully embodies the extraction, migration, customization of voiceprint features, and their ultimate adaptive application in speech synthesis.
[0059] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A digital human voice color adaptive matching method based on voiceprint feature migration, characterized in that Including the following steps: Processing the user's voice sample with a deep voiceprint decoupling network, which extracts a pure tone color identity embedding decoupled from the content information and general acoustic feature information from the voice sample through a multi-task adversarial learning strategy, and extracts multi-scale tone color features for characterizing the macroscopic characteristics and microscopic details of the user's tone color; Selecting, combining and assigning voice synthesis weights from the multi-scale tone color features according to the user configuration to form a customized voiceprint feature set; Injecting the pure tone color identity embedding and the customized voiceprint feature set as conditions into a diffusion flow matching neural vocoder; after receiving the content information to be synthesized, the diffusion flow matching neural vocoder generates a digital human voice synthesis waveform with the user's unique tone color and corresponding to the content through probability flow conversion and step-by-step denoising refinement.
2. The digital human voice color adaptive matching method based on voiceprint feature migration according to claim 1, characterized in that The deep voiceprint decoupling network includes at least one user identity encoding backbone network, at least one content information discriminator and at least one general acoustic feature information discriminator; the general acoustic feature information includes the average fundamental frequency range, the basic prosody pattern and the general acoustic performance of common emotion types.
3. The digital human voice color adaptive matching method based on voiceprint feature migration according to claim 1, wherein The multi-task adversarial learning strategy includes: By training the user identity encoding backbone network, minimizing the user classification loss of the tone color identity embedding; By training the content information discriminator, maximizing the accuracy of calculating content information from the tone color identity embedding, and at the same time training the user identity encoding backbone network, minimizing the accuracy of calculating content information from the tone color identity embedding; By training the general acoustic feature information discriminator, maximizing the accuracy of calculating general acoustic feature information from the tone color identity embedding, and at the same time training the user identity encoding backbone network, minimizing the accuracy of calculating general acoustic feature information from the tone color identity embedding.
4. The digital human voice color adaptive matching method based on voiceprint feature migration according to claim 1, characterized in that The multi-scale tone color features include: Global tone color features extracted from the deep network layer of the deep voiceprint decoupling network, which characterize the global static characteristics of the tone color; and local tone color features extracted from the shallow network layer of the deep voiceprint decoupling network, which characterize the local dynamic details of the tone color; the global tone color features include the average pitch range and the base tone quality parameters, and the local tone color features include the formant dynamic parameters and the acoustic patterns of the vocalization habits.
5. The digital human voice color adaptive matching method based on voiceprint feature migration according to claim 1, wherein: The step of selecting, combining and assigning voice synthesis weights from the multi-scale tone color features according to the user configuration includes: Receiving user input, including at least one user configuration parameter, the user configuration parameter including the tone color style label of the target digital human, the desired degree of realism and the specific application scenario information; parsing the user configuration parameter and mapping it to the target area in the predefined tone color feature space; initially screening out the feature subset most relevant to the target area from the extracted global tone color features and local tone color features; setting initial benchmark contribution weights for the screened global tone color feature subset and local tone color feature subset respectively; and making an adaptive dynamic adjustment of the voice synthesis weights based on the user configuration parameter, and calculating the final voice synthesis weights of each selected tone color feature in the customized voiceprint feature set.
6. The method for adaptive matching of the digital human voice color based on voiceprint feature migration according to claim 1, wherein The processing process of the diffusion flow matching neural vocoder specifically includes: Receive the pure tone identity embedding, the customized voiceprint feature set, and the content information to be synthesized as input conditions; use a conditional probability flow model to map and generate a preliminary acoustic representation by reversibly transforming a predefined Gaussian noise, and the preliminary acoustic representation is output as a Mel spectrogram; Use the preliminary acoustic representation as the starting state of the diffusion process; initialize a denoising network, which takes the pure tone identity embedding, the customized voiceprint feature set, and the content information of the speech to be synthesized as conditional inputs; perform iterative denoising through a preset number of time steps. At each time step, the denoising network gradually removes noise and enhances the detail and harmonic structure of the speech according to the acoustic representation at the current time step and the injected conditional information; repeat this denoising step until a preset termination condition is reached to obtain a denoised acoustic representation; convert the denoised acoustic representation into a final digital human speech waveform through a vocoder decoding module.
7. The digital human voice color adaptive matching system based on voiceprint feature migration is characterized in that The system includes: A voiceprint feature processing module for processing a user's speech sample using a deep voiceprint decoupling network. The deep voiceprint decoupling network extracts a pure tone identity embedding decoupled from the content information and general acoustic feature information from the speech sample through a multi-task adversarial learning strategy, and extracts multi-scale tone features for characterizing the macroscopic characteristics and microscopic details of the user's tone; A customized feature set generation module for selecting and combining from the multi-scale tone features according to user configuration and assigning speech synthesis weights to form a customized voiceprint feature set; A speech synthesis module for injecting the pure tone identity embedding and the customized voiceprint feature set as conditions into a diffusion flow matching neural vocoder; after receiving the content information to be synthesized, the diffusion flow matching neural vocoder generates a digital human speech synthesis waveform with the user's unique tone and corresponding to the content through probability flow conversion and gradual denoising refinement.
Citation Information
Patent Citations
Voice style migration system and method for tone and style deep decoupling
CN117912446A
Tone generation method based on voice conversion
CN118197329A
Interpretable multi-dimensional voice style control method
CN120183379A
Artificial intelligence music generation model and method for configuring the same
US20250054473A1
Cited By
Sound cloning method and system, electronic equipment and medium
CN121983069A
Sound cloning methods, systems, electronic devices and media
CN121983069B
TTS model-based anthropomorphic speech synthesis method, device and system
CN122290568A