Interpretable multi-dimensional voice style control method
Through hierarchical decoupling network and dynamic mapping technology, accurate analysis of natural language descriptions and independent regulation of multidimensional speech style characteristics are achieved, solving the problems of single regulation methods and cumbersome operation processes in the existing technology, and improving the personalization and user experience of speech synthesis.
Patent Information
- Application Number
- CN202510463171.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-20
AI Technical Summary
The existing voice style control technology has problems such as single regulation methods, cumbersome operational processes, and poor user experience.
The hierarchical decoupling network is adopted to realize the accurate analysis of the user's natural language description and independent regulation of multidimensional speech style characteristics through the speech synthesis backbone network, semantic encoder and dynamic mapping module.
It significantly improves the personalization and flexibility of speech synthesis, simplifies operation steps, improves user experience, and has the ability to generate zero-sample styles.
Smart Images

Figure CN120183379A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech synthesis, and specifically proposes an interpretable multi-dimensional speech style control method. Background Art
[0002] Existing speech style control technologies mainly adopt a dual-channel design framework of "backbone model + modulation module". Their technical implementation is usually based on general speech synthesis architectures (such as FastSpeech2, VITS, etc.), and personalized speech output is achieved by adding independent style modulation components. The core technical implementation of this system mainly includes the following two key links:
[0003] 1. Style semantic feature extraction
[0004] The key point of this link is to accurately capture semantic information related to the speech style from the text instructions input by the user, while filtering out the noise caused by expression differences. Existing solutions mostly use pre-trained text models (such as T5, BERT, etc.) as basic feature extractors, and adapt the general text understanding ability to the style feature extraction task through parameter fine-tuning.
[0005] 2. Cross-modal feature integration
[0006] This stage needs to effectively integrate the text style features into the speech synthesis process. Currently, there are mainly the following implementation methods:
[0007] (1) Attention-based fusion: Adopt cross-attention mechanism to establish dynamic association between text and speech features
[0008] (2) Operation-based fusion: Include basic operation methods such as direct splicing in the time series / feature dimension and feature superposition that maintains dimension consistency
[0009] In recent years, speech synthesis technology has been deeply integrated into many scenarios such as intelligent assistants, navigation systems, and digital content production. However, the current system still faces significant challenges in speech style regulation (including dimensions such as tone, emotion, rhythm, pitch, etc.): problems such as single regulation method and cumbersome operation process lead to poor user experience and affect the enthusiasm for use. In response to this situation, this research innovatively developed a speech style regulation scheme based on natural language instructions. This technology allows users to directly input natural language descriptions such as "Please use a sad tone" and "Adopt a low voice", and the intelligent parsing module realizes precise adjustment of speech features. This method not only greatly simplifies the operation steps, but also can flexibly meet the user's needs for diverse speech expressions, effectively improving product stickiness. From the perspective of the technological evolution trend, natural language interaction will surely become the mainstream direction of speech style control, which makes the academic value and application prospect of this research particularly prominent. Summary of the Invention
[0010] This method aims to solve the problems existing in the existing voice style control technologies, such as single control method, cumbersome operation process, and poor user experience. By introducing a hierarchical decoupling network, the present invention realizes the accurate parsing of the natural language description input by the user and the independent control of multi-dimensional voice style features, significantly improving the personalization and flexibility of speech synthesis.
[0011] To achieve the above object, the present invention adopts the following technical solutions:
[0012] An interpretable multi-dimensional voice style control method, comprising the following steps:
[0013] S1: Data preprocessing and feature decoupling, decoupling the acoustic features of multi-speaker voice data, and extracting independent control parameters of fundamental frequency, energy, and Mel spectrogram;
[0014] S2: Hierarchical network training, generating style parameters through a speech synthesis backbone network, a semantic encoder, and a dynamic mapping module;
[0015] S3: User interaction and real-time optimization, completing dynamic voice style control based on natural language instructions and physical parameter mapping rules, and providing an interpretable interaction interface.
[0016] Further, in the present invention, the S1 includes:
[0017] Constructing a multi-style voice data set covering 10 emotion types and 5 role types, annotating the fundamental frequency (F0), short-time energy (Energy), and Mel-spectrogram parameters and performing normalization processing;
[0018] Adopting the YIN algorithm to extract the fundamental frequency trajectory, calculating the energy and spectrum through the short-time Fourier transform (STFT), and introducing an adversarial discriminator network to optimize feature independence.
[0019] Further, in the present invention, the speech synthesis backbone network in the S2 is the FastSpeech2 model, generating speech features through the teacher-student distillation strategy, and the training loss function includes L1 loss and cross-entropy loss, and the naturalness MOS of the target speech ≥ 4.0.
[0020] Further, in the present invention, the dynamic mapping module in the S2 adopts a diffusion model to generate acoustic parameters, performs 1000 steps of denoising during training, and compresses the latent space dimension through a variational autoencoder (VAE) to reduce the computational complexity.
[0021] Further, in the present invention, the semantic encoder in the S2 is implemented based on a CLIP-like model, and the text-style alignment data set is fine-tuned through contrastive learning to optimize the mapping relationship between text embeddings and style vectors.
[0022] Further, in the present invention, S3 includes:
[0023] Construct an instruction-parameter mapping table to parse natural language instructions into quantization adjustment parameters of fundamental frequency and energy;
[0024] Dynamically update the mapping rules through reinforcement learning and design a heatmap interface to visualize the parameter regulation intensity.
[0025] Compared with the prior art, the present invention has the following significant advantages:
[0026] Multi-dimensional style control: Achieve multi-dimensional independent regulation of speech style through a hierarchical decoupling network, avoiding the problem of strong style coupling in traditional models.
[0027] Strong natural language understanding ability: Adopt a CLIP-like model and a Diffusion model to improve the system's parsing ability for natural language descriptions input by users, and can accurately map to acoustic parameters.
[0028] Zero-shot style generation ability: The system can generate new-style speech without relying on a large amount of labeled data or reference audio, reducing data dependence.
[0029] Good interaction interpretability: Users can intuitively understand or change the physical meaning of style parameters by adjusting high-level semantic vectors, improving the interaction experience. Description of the Drawings
[0030] Figure 1 It is a schematic diagram of the overall structure of the system in the present invention. Detailed Embodiment
[0031] The following will elaborate on the detailed embodiments of each part.
[0032] An interpretable multi-dimensional speech style control method includes the following steps:
[0033] S1: Data preprocessing and feature decoupling
[0034] Dataset construction: Collect multi-speaker speech data covering 10 emotion types (such as "angry", "calm") and 5 role types (such as "teacher", "customer service");
[0035] Annotate the fundamental frequency (F0), short-time energy (Energy), and Mel-spectrogram parameters, and perform normalization processing.
[0036] Feature extraction and decoupling:
[0037] Use the YIN algorithm to extract the fundamental frequency trajectory, and use the short-time Fourier transform (STFT) to calculate the energy and spectrum;
[0038] Introduce an adversarial discriminator network to optimize the feature decoupling ability of the underlying encoder (Formula 1):
[0039] L adv = E[logD(x)] + E[log(1 - D(G(z)))] (Formula 1)
[0040] where D is the discriminator, G is the underlying encoder, x is the real feature, and z is the latent vector.
[0041] S2: Hierarchical network training process
[0042] Backbone network training:
[0043] Use FastSpeech2 as the speech synthesis backbone and generate high-quality speech features through the teacher-student distillation strategy (the training loss function includes L1 loss and cross-entropy loss);
[0044] Pre-train on the LibriTTS dataset with the target speech naturalness MOS ≥ 4.068.
[0045] Semantic encoder optimization:
[0046] Construct a text-style alignment dataset and fine-tune the CLIP-like model through contrastive learning (Formula 2):
[0047]
[0048] where s i is the text embedding and t i is the corresponding style vector.
[0049] Dynamic mapping module implementation:
[0050] Adopt a diffusion model to generate acoustic parameters, gradually add noise during training and learn the inverse process (the denoising step T = 1000);
[0051] Compress the latent space dimension through a variational autoencoder (VAE) to reduce the computational complexity.
[0052] S3: User interaction and real-time optimization
[0053] Instruction parsing and mapping:
[0054] Construct an instruction-parameter mapping table (for example, "sadness" corresponds to a 10% decrease in fundamental frequency and a 15% decrease in energy);
[0055] Dynamically adjust the mapping rules through reinforcement learning to adapt to the expression habits of different users.
[0056] Interpretability interface design:
[0057] Show the regulation intensity of parameters such as pitch and energy in the form of a heat map;
[0058] Users can fine-tune the parameters through natural language instructions or sliders.
[0059] This method realizes the multi-dimensional independent control of speech styles through a hierarchical decoupling network and dynamic mapping technology, and solves the defects of traditional methods in decoupling degree, interaction interpretability, and data dependence. Specific implementation cases show that this method has significant technical advantages and commercial potential in scenarios such as intelligent hardware and barrier-free assistance.
[0060] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. An explainable multi-dimensional speech style control method, characterized in that: The following steps are involved: S1: Data preprocessing and feature decoupling: decoupling the acoustic features of multi-speaker speech data and extracting independent control parameters of fundamental frequency, energy and Mel spectrum; S2: Hierarchical network training, generating style parameters through the speech synthesis backbone network, semantic encoder and dynamic mapping module; S3: User interaction and real-time optimization, completes dynamic control of voice style based on natural language instructions and physical parameter mapping rules, and provides an explainable interactive interface.
2. The interpretable multi-dimensional voice style control method according to claim 1, characterized in that: The S1 includes: Construct a multi-style speech dataset covering 10 emotion types and 5 role types, annotate the fundamental frequency (F0), short-term energy (Energy), and Mel-spectrogram parameters, and perform normalization processing; The YIN algorithm is used to extract the fundamental frequency trajectory, the energy and spectrum are calculated through short-time Fourier transform (STFT), and the adversarial discriminator network is introduced to optimize feature independence.
3. The interpretable multi-dimensional voice style control method according to claim 1, characterized in that: The speech synthesis backbone network in S2 is the FastSpeech2 model, which generates speech features through the teacher-student distillation strategy. The training loss function includes L1 loss and cross entropy loss, and the target speech naturalness MOS ≥ 4.
0.
4. The interpretable multi-dimensional speech style control method according to claim 1, characterized in that: The dynamic mapping module in S2 uses a diffusion model to generate acoustic parameters, performs a 1000-step denoising process during training, and compresses the latent space dimension through a variational autoencoder (VAE) to reduce computational complexity.
5. The interpretable multi-dimensional voice style control method according to claim 1, characterized in that: The semantic encoder in S2 is implemented based on the CLIP-like model, and the mapping relationship between text embedding and style vector is optimized through contrastive learning and fine-tuning the text-style alignment dataset.
6. The interpretable multi-dimensional voice style control method according to claim 1, characterized in that: The S3 includes: Construct an instruction-parameter mapping table to parse natural language instructions into quantized adjustment parameters of base frequency and energy; The mapping rules are dynamically updated through reinforcement learning, and a heat map interface is designed to visualize the parameter regulation intensity.
Citation Information
Cited By
Digital human tone adaptive matching method and system based on voiceprint feature migration
CN120356474A