Voice generation method and device based on voice style adaptation, equipment and medium
By extracting and fusion of the features of the target text, speaker's voice and reference style speech, acoustic encodings containing the target style characteristics are generated, which solves the problems of insufficient pronunciation style modeling and insufficient rhythm control in the prior art, and achieves more natural and personalized speech generation.
Patent Information
- Application Number
- CN202510651564.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-22
AI Technical Summary
The existing pronunciation synthesis technology lacks the ability to model the pronunciation style in specific scenarios, lacks pronunciation control, and has poor generalization ability when dealing with the speaker without seeing it, resulting in a lack of naturalness and personalized adaptability in the generated pronunciation.
By obtaining the target text, the target speaker's voice and reference style speech, the phoneme feature sequence, acoustic features and rhythm encoding information are extracted, the feature fusion is fusion-generated, and the acoustic encoding containing the target style features is generated through the style adaptation module. Finally, the Mel spectral features are generated and the pre-trained vocoder is input to generate a speech waveform.
It improves the naturalness and expressiveness of the speech synthesis system in terms of tone matching, rhythm adjustment and emotional expression, enhances the adaptability to specific scenes and the generalization ability of the unknown speaker, and improves the clarity and stability of speech generation.
Smart Images

Figure CN120526752A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech synthesis technology, and in particular to a speech generation method, apparatus, device and storage medium based on speech style adaptation. Background Art
[0002] Current zero-shot speech synthesis technology has made significant progress in multi-speaker adaptation and multi-style speech synthesis, enabling the generation of highly natural, personalized speech with limited training data. However, existing technologies still face several limitations when it comes to speech generation for specific contexts. This is particularly true in debate scenarios, fintech, and healthcare, where the demand for voice style, prosody control, and adaptation to unseen speakers remains unmet.
[0003] In the healthcare sector, speech synthesis technology is widely used in telemedicine, health consultation, and voice-assisted systems, such as smart health assistants, doctor-assisted voice recording, and patient health status feedback. However, existing TTS systems struggle to ensure accuracy and comprehensibility when generating medical terminology. This is especially true when describing symptoms, treatment plans, or medical advice, where the voice expression often lacks sufficient professionalism and credibility. Furthermore, voice interaction in the medical field requires adjusting the voice style based on the patient's psychological state and health status. For example, when a patient is anxious or uneasy, a more soothing voice should be used, while when providing urgent health advice, clarity and authority should be enhanced. However, existing systems struggle to flexibly adjust voice style to meet the needs of diverse medical scenarios. Furthermore, when faced with voice input from an unseen patient, existing systems have limited adaptability and may not accurately match the patient's language style, impacting the fluency of voice interaction and the accuracy of information transmission.
[0004] In the FinTech sector, intelligent customer service, investment advisors, and voice-interactive financial assistants require highly adaptable, personalized voice for diverse scenarios. For example, during investment advisory sessions, voice assistants need to adjust their voice tone based on market fluctuations and user investment preferences to enhance user trust and understanding. However, existing TTS (Text-to-Speech) systems lack sufficient contextual understanding when processing specialized financial terminology, resulting in monotonous voice and inability to effectively distinguish between key information and supporting explanations. Furthermore, financial risk warning systems typically require a serious and authoritative voice style, but existing systems lack the ability to adapt to this style, making it difficult for voice expressions to convey the urgency and authority of risk warnings. Furthermore, the financial sector has a broad user base encompassing clients of varying ages and professional backgrounds. Existing systems have poor generalization capabilities for unseen speakers or language styles, which can lead to discrepancies between voice expressions and user expectations, impacting user experience and the effectiveness of information delivery.
[0005] In terms of speech synthesis in debate scenarios, existing technologies are mainly optimized for general speech synthesis scenarios and lack accurate modeling of speech features in specific contexts. During a debate, the rebuttal speech not only needs to reflect the timbre characteristics of the target speaker, but also needs to adapt to the tone, rhythm, and prosody of the opponent's voice. However, current speech synthesis systems are generally unable to effectively utilize the opponent's voice style information, resulting in the generated rebuttal speech lacking the unique tone of confrontation and logical fluency of debate. In addition, the existing system's rhythmic control capabilities are limited, making it difficult to accurately adjust the speaking speed, stress, or emotional expression, resulting in a significant gap between the naturalness and expressiveness of the generated speech and the real debate speech. Summary of the Invention
[0006] The main purpose of the present invention is to provide a speech generation method, device, equipment and storage medium based on speech style adaptation, aiming to solve the technical problems that the existing technology lacks speech style modeling for specific scenarios, has insufficient prosody control capabilities, and has poor generalization capabilities when dealing with unseen speakers, resulting in the generated speech lacking naturalness and personalized adaptability.
[0007] To achieve the above object, the present invention provides a speech generation method based on speech style adaptation, comprising:
[0008] Obtain target text, target speaker voice, and reference style voice;
[0009] Extracting a phoneme feature sequence of the target text;
[0010] Extracting acoustic features of the target speaker's speech;
[0011] extracting prosodic coding information of the reference style speech;
[0012] Performing feature fusion on the phoneme feature sequence, acoustic features and prosody coding information to generate a fusion coding vector;
[0013] Processing the fused code vector through a style adaptation module to generate an acoustic code containing target style features;
[0014] Mel-spectrogram features are generated according to the acoustic coding, and the Mel-spectrogram features are input into a pre-trained vocoder to generate a speech waveform.
[0015] Furthermore, to achieve the above-mentioned object, the present invention provides a speech generation device based on speech style adaptation, comprising:
[0016] A data acquisition module is used to acquire target text, target speaker speech, and reference style speech;
[0017] A text processing module, configured to extract a phoneme feature sequence of the target text;
[0018] an acoustic feature extraction module, configured to extract acoustic features of the target speaker's speech;
[0019] A prosody analysis module, configured to extract prosody coding information of the reference style speech;
[0020] A feature fusion module, configured to fuse the phoneme feature sequence, acoustic features, and prosodic coding information to generate a fused coding vector;
[0021] a style adaptation module, configured to process the fused code vector through the style adaptation module to generate an acoustic code containing target style features;
[0022] The speech generation module is used to generate Mel-spectrogram features according to the acoustic coding, and input the Mel-spectrogram features into a pre-trained vocoder to generate a speech waveform.
[0023] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a speech generation program based on speech style adaptation stored in the memory and runnable on the processor, wherein the speech generation program based on speech style adaptation, when executed by the processor, implements the steps of the speech generation method based on speech style adaptation as described above.
[0024] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a speech generation program based on speech style adaptation is stored. When the speech generation program based on speech style adaptation is executed by a processor, the steps of the speech generation method based on speech style adaptation as described above are implemented.
[0025] Beneficial effects: The present invention relates to the field of speech synthesis technology and can be applied to business scenarios such as medical health, financial technology and debate. It discloses a speech generation method based on speech style adaptation, including: obtaining a target text, a target speaker's speech and a reference style speech, extracting the phoneme feature sequence of the target text, the acoustic features of the target speaker's speech and the prosodic coding information of the reference style speech, performing feature fusion on the phoneme feature sequence, acoustic features and prosodic coding information to generate a fused coding vector; processing the fused coding vector through a style adaptation module to generate an acoustic code containing the target style features, generating a Mel spectrum feature based on the acoustic code, and inputting the mel spectrum feature into a pre-trained vocoder to generate a speech waveform. The present invention integrates the phoneme features of the target text, the acoustic features of the target speaker, and the prosodic coding information of the reference style speech to make the generated speech more natural and expressive in terms of timbre matching, prosodic adjustment, and emotional expression. The style adaptation module optimizes the speech style mapping, making the generated speech adaptable to specific scenario requirements and improving the generalization ability of unseen speakers. The processing based on Mel-spectrogram features and pre-trained vocoder makes the speech generation have higher clarity and stability, thereby improving the adaptability and practical value of the speech synthesis system. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0027] Figure 1 Schematic diagram of an application environment of a speech generation method based on speech style adaptation in one embodiment of the present invention;
[0028] Figure 2 1 is a flow chart of an embodiment of a method for generating speech based on speech style adaptation according to the present invention;
[0029] Figure 3 Schematic diagram of functional modules of a preferred embodiment of a speech generation device based on speech style adaptation of the present invention;
[0030] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0031] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0033] The speech generation method based on speech style adaptation provided by the embodiment of the present invention can be applied in Figure 1In an application environment, the user end communicates with the server end through a network. The server end can obtain the target text, the target speaker's voice and the reference style voice through the user end, extract the phoneme feature sequence of the target text, the acoustic features of the target speaker's voice and the prosody coding information of the reference style voice, perform feature fusion on the phoneme feature sequence, acoustic features and prosody coding information to generate a fused coding vector; process the fused coding vector through the style adaptation module to generate an acoustic code containing the target style features, generate a Mel spectrum feature based on the acoustic code, and input the pre-trained vocoder to generate a speech waveform. The present invention makes the generated speech more natural and expressive in terms of timbre matching, prosody adjustment and emotional expression by fusing the phoneme features of the target text, the acoustic features of the target speaker and the prosody coding information of the reference style voice; optimizes the speech style mapping through the style adaptation module so that the generated speech can adapt to the needs of specific scenarios and improves the generalization ability of unseen speakers; based on the processing of the Mel spectrum features and the pre-trained vocoder, the speech generation has higher clarity and stability, thereby improving the adaptability and practical value of the speech synthesis system. The user end may be, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server end may be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below using specific embodiments.
[0034] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of a method for speech generation based on speech style adaptation provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0035] like Figure 2 As shown, the speech generation method based on speech style adaptation proposed by the present invention includes the following steps:
[0036] S10, obtaining target text, target speaker voice, and reference style voice;
[0037] In this embodiment, the process of acquiring the target text, target speaker voice, and reference style speech involves the combined input of voice and text data, ensuring consistency in content, timbre, and style of the generated speech. The target text is the textual content of the speech to be synthesized and provides the semantic information required for speech synthesis. The target speaker voice contains the speaker's timbre characteristics, pitch patterns, and personalized prosodic features, ensuring that the generated speech matches the target speaker's timbre. The reference style speech provides control parameters for the speech style, including information such as speech rate, emotion, and tone, ensuring that the generated speech has the desired expressiveness in specific scenarios.
[0038] The target text can be obtained through speech-to-text (ASR) systems, manual input, database access, or preprocessing with natural language processing systems. The target text typically consists of complete sentences, phrases, or word sequences, requiring preprocessing steps such as text normalization, word segmentation, and stop word removal to improve speech synthesis accuracy. Text data can come from news releases, text input for dialogue systems, and user-provided instruction text. In some applications, sentiment analysis models can be incorporated to determine the target text's tone and further optimize the voice style adaptation process.
[0039] The target speaker's voice is typically acquired from pre-recorded audio samples, which can be sourced from a cloud database, user audio recorded by a voice assistant, or recorded in real time using a voice acquisition device. The acquired target speaker's voice data undergoes preprocessing, such as noise removal, audio alignment, and phoneme segmentation, to ensure the stability and reliability of the input voice features. The source of the target speaker's voice may vary in different application scenarios. For example, in intelligent customer service scenarios, the target speaker's voice may be derived from historical user voice interaction records. In the field of speech synthesis research, existing large-scale speech datasets can be used as the source of the target speaker's voice.
[0040] The selection of reference-style speech can be based on a speech style library, historical speech data, or user-provided audio samples. Speech style characteristics include, but are not limited to, rhythm, prosody, emotion, intonation, and energy variation. The reference-style speech undergoes prosodic encoding, converting its style information into a quantifiable encoding vector. This ensures that this style information can be used to control the expressiveness of the synthesized speech during speech generation. A style analysis model can be used to extract subjective emotional characteristics of the reference-style speech, such as anger, joy, and calmness. These characteristics can be manipulated during speech synthesis, ensuring that the generated speech's style is more natural and tailored to the application scenario.
[0041] Different approaches can be used to obtain the target text, target speaker voice, and reference style voice in different technical environments. In cloud-based data processing platforms, the target text can be automatically converted from user voice input via a speech recognition system, or the user can directly enter the text. In mobile devices, the target text can be extracted from instant messaging apps, emails, or text messages. In smart home systems, the target text can be derived from voice commands.
[0042] The method for acquiring the target speaker's voice can be adjusted based on different application scenarios. In intelligent voice assistant systems, the target speaker's voice can be derived from historical user interaction records, stored and optimized with user authorization. In teleconferencing systems, the target speaker's voice can be collected through a real-time recording module and enhanced with neural network-based denoising to improve voice quality. For scenarios with multiple target speakers, speaker separation technology can be used to distinguish different speech sources through a deep learning network and independently extract the timbre characteristics of each target speaker.
[0043] The selection of reference speech styles can be dynamically adjusted based on user preferences and application scenarios. In personalized speech synthesis systems, based on the user's historical speech input, the historical speech closest to the current text content can be automatically extracted as the reference speech style. In emotional speech synthesis systems, based on text sentiment classification algorithms, a speech database with corresponding styles can be automatically matched and appropriate speech samples can be selected as the reference speech style. In film and television dubbing or game character speech synthesis systems, existing character speech samples can be combined to extract the speech style characteristics of specific characters and use them as reference speech styles to achieve personalized dubbing.
[0044] Example: In the healthcare sector, target text can be used to generate voice output for intelligent health assistants. For example, in a telemedicine consultation system, a doctor's diagnosis can serve as the target text, which is converted by a speech synthesis system into voice output and provided to the patient. During this process, the target speaker's voice can be a pre-recorded sample of the doctor's speech to ensure consistency in speech style. The reference style can be derived from a previous recording of the patient's speech, allowing the generated speech to adjust its speed and intonation to suit the patient's emotional state, enhancing the comfort and credibility of the voice interaction.
[0045] In the financial sector, target text acquisition can be used in voice interaction systems for smart investment advisors. When users inquire about financial products, the system can obtain target text, such as market analysis reports or investment recommendations, and synthesize it with the target speaker voice of a financial analyst, ensuring professional and credible voice output. Reference voice styles can be selected from audio samples of financial reports under different market conditions. For example, during periods of high market volatility, a more cautious and relaxed voice style can be selected, making it easier for users to accept financial decision-making advice.
[0046] In a debate scenario, target text is obtained to generate rebuttal speech. The target speaker voice ensures that the speech output matches the timbre characteristics of a specific debater. The reference style speech is derived from the opponent's voice, allowing the generated speech to mimic the opponent's tone, rhythm, and accent adjustment, thereby enhancing the confrontational and persuasive power of the debate speech. For example, in a legal debate training system, the input target text is a legal provision or case analysis. The target speaker voice can be a lawyer's voice sample, and the reference style speech can be the opponent's debate speech. This allows the generated speech to automatically adjust the rhythm and emotional expression of the opponent's tone, improving the realism and effectiveness of the debate training.
[0047] By acquiring the target text, the target speaker's voice, and a reference style speech, the speech synthesis process not only preserves the target speaker's timbre characteristics but also adaptively adjusts the speech style in specific scenarios, enhancing the naturalness and expressiveness of the synthesized speech. The target text provides the core semantic information for speech synthesis, the target speaker's voice ensures timbre matching, and the reference style speech further enhances the personalized characteristics of the synthesized speech, making the generated speech more adaptable to specific contextual needs.
[0048] S20, extracting a phoneme feature sequence of the target text;
[0049] In this embodiment, the process of extracting the phoneme feature sequence of the target text involves linguistic analysis and acoustic conversion of the text, so that the input text can be mapped to the corresponding phoneme sequence, so that the subsequent speech synthesis process can accurately restore the pronunciation characteristics of the speech. Phoneme is the smallest pronunciation unit of speech, which determines the basic pronunciation method of speech synthesis. The phoneme sets of different languages are different, so language adaptability needs to be considered when extracting phonemes. The phoneme feature sequence not only includes phoneme symbols, but may also contain prosodic information, such as syllable duration, stress pattern, linking rules, etc., to ensure the naturalness and fluency of the generated speech.
[0050] The extraction of phoneme features is usually completed by a phoneme segmentation model, which can be implemented based on statistical methods (such as hidden Markov models) or neural network methods (such as Transformer, BiLSTM+CRF, etc.). The text is first preprocessed, including removing punctuation, standardizing numbers and abbreviations, processing foreign words, etc., in order to improve the accuracy of phoneme conversion. Then, the target text is converted into the corresponding phoneme symbol sequence through the pre-trained phoneme segmentation model. For example, in English text processing, the word "cat" may be converted into the phoneme sequence / k / / t / , while in Chinese, the pinyin "zhong" may be converted to
[0051] The phoneme symbol sequence needs to be further converted into a phoneme embedding vector so that subsequent deep learning models can more efficiently process the phoneme features. Embedding vectors are typically generated through word embedding or character-level vectorization, allowing the phoneme sequence to preserve the relationships between speech information in a high-dimensional vector space. Furthermore, to better process the temporal information of speech, the phoneme embedding vectors need to be positionally encoded so that the phoneme sequence is correctly arranged in the time dimension and the relative positions of the syllables are preserved.
[0052] The positionally encoded phoneme sequence is then fed into a bidirectional long short-term memory (BiLSTM) network to extract contextual features. BiLSTM leverages information from both forward and backward passes to perform fine-grained modeling of phoneme features, ensuring that the representation of each phoneme is not only based on its own characteristics but also incorporates contextual information from preceding and following phonemes. For example, in connected speech, the pronunciation of certain phonemes may vary due to the influence of preceding and following phonemes. BiLSTM can capture these variations, improving the naturalness of speech synthesis.
[0053] Finally, the phoneme context-related features are concatenated with the phoneme embedding vectors to form a more complete phoneme feature representation. Layer normalization is then performed to ensure stable data distribution and reduce the gradient vanishing or gradient exploding problems during model training, thereby generating high-quality phoneme feature sequences.
[0054] The method for extracting phoneme feature sequences can vary in different application scenarios. In systems based on multilingual speech synthesis, a multilingual phoneme mapping model can be used to automatically match corresponding phoneme sets to text in different languages. For example, in intelligent translation systems, the input text may contain multiple languages, and the phoneme segmentation model needs to be compatible with different phoneme systems and perform cross-language phoneme conversion.
[0055] In real-time speech synthesis systems, phoneme extraction must meet low latency requirements. Lightweight deep learning models, such as MobileBERT or the quantized Transformer model, can be used to improve computational efficiency. In such systems, the phoneme extraction process can adopt a streaming approach, performing phoneme segmentation and feature conversion simultaneously with text input. This allows the entire system to respond more quickly to user input and improve the interactive experience.
[0056] In adaptive speech synthesis systems, phoneme embeddings can be personalized based on the target speaker's timbre characteristics. For example, in a personalized navigation system, different users may prefer different pronunciation styles, such as American English or British English. The system can adjust the weights of phoneme features based on the user's speech history, making the synthesized speech more consistent with the user's daily pronunciation habits.
[0057] By extracting the target text's phoneme feature sequences, the text information can be accurately converted into the basic pronunciation units required for speech. Contextual information is then incorporated to optimize the phoneme representation, improving the clarity and fluency of speech synthesis. By using positional encoding and a bidirectional long short-term memory network, the phoneme features preserve the temporal information of speech, enhancing the naturalness of prosodic features such as connected speech and stress. Furthermore, layer normalization ensures the stability of the phoneme features, improving the quality of subsequent speech generation and the model's generalization capabilities.
[0058] S30, extracting acoustic features of the target speaker's speech;
[0059] In this embodiment, extracting the acoustic features of the target speaker's speech involves a multi-dimensional analysis of the target speaker's audio data to extract the core timbre, prosody, and vocal characteristics used for speech synthesis. Acoustic features are an important foundation for building natural, personalized speech, ensuring that the generated speech matches the timbre of the target speaker and maintains their unique voice style. Acoustic feature extraction typically includes multiple dimensions of feature information, such as fundamental frequency, mel-spectrogram envelope, and formant frequency distribution, to provide a comprehensive description of the speech.
[0060] Fundamental frequency (F0) reflects the pitch variation of the target speaker and plays a decisive role in the tone, emotion, and style of the speech. Fundamental frequency features can be extracted through methods such as Linear Predictive Coding (LPC) or Kalman Filter. The LPC method estimates the variation of the fundamental frequency by analyzing the harmonic structure of short-term speech signals, while the Kalman Filter can optimize the fundamental frequency curve in more complex contexts, reduce noise interference, and make the fundamental frequency information smoother and more stable. For example, when simulating the intonation of the target speaker, high-frequency fundamental frequency variations are often used to express excitement or tension, while lower fundamental frequency variations are used to express calmness or composure.
[0061] The Mel Spectral Envelope (Mel Spectral Envelope) is used to describe the spectral characteristics of the target speaker's speech, ensuring that the generated speech matches the target speaker's timbre distribution. Mel Spectral Envelope is a spectral analysis method optimized for human auditory perception, which better aligns with the human ear's perception of speech timbre. The Short-Time Fourier Transform (STFT) is a commonly used method for calculating the Mel Spectral Envelope (Mel Spectral Envelope) of the target speaker's speech. By performing time-domain decomposition of audio data through a sliding window, the STFT captures the frequency distribution of speech at different time segments and maps it to the Mel scale, improving the ability to model timbre characteristics. For example, when synthesizing the target speaker's speech, the Mel Spectral Envelope (Mel Spectral Envelope) can ensure that the timbre of the synthesized speech closely matches the target speaker's natural vocalization habits without excessively distorting frequency shifts.
[0062] Formant frequency distribution is an important determinant of speech timbre and is mainly used to describe the resonance characteristics of the target speaker's speech. Formant is a key frequency peak determined by the shape of the vocal tract and plays a decisive role in the personalization of speech of different speakers. Through Linear Predictive Analysis (LPA) or Cepstral Analysis, the formant frequency distribution of the target speaker can be extracted and its stability characteristics can be calculated. For example, the formant characteristics of different people usually show a fixed pattern. For example, the first formant (F1) usually corresponds to the degree of opening of the vowel, while the second formant (F2) corresponds to the position of the tongue. By analyzing the formant distribution, it can be ensured that the timbre of the synthesized speech is more in line with the natural pronunciation characteristics of the target speaker.
[0063] In order to further improve the stability of acoustic features, it is necessary to perform tensor splicing on the extracted fundamental frequency features, Mel-spectrogram envelope, and formant frequency distribution to form a complete acoustic feature representation. Tensor splicing can be done by channel splicing or time dimension splicing, so that different types of acoustic features can complement each other and improve the expressiveness of the speech synthesis model. Subsequently, the spliced acoustic features are input into a pre-trained acoustic encoder to generate a high-dimensional joint acoustic feature matrix. The pre-trained acoustic encoder can be based on a variational autoencoder (VAE) or a deep residual network (ResNet). By learning the statistical distribution of large-scale speech data, the acoustic features of the target speaker can be better generalized to different speech content, thereby improving the stability and naturalness of the synthesized speech.
[0064] By extracting the acoustic features of the target speaker's speech, the generated speech accurately matches the target speaker's timbre, intonation, and prosody, enhancing the personalization of speech synthesis. Extracting fundamental frequency features ensures pitch consistency in speech synthesis, calculating the Mel-spectrogram envelope enhances the natural timbre of the speech, and analyzing the formant distribution optimizes the target speaker's individual voice expression. Through tensor concatenation and acoustic encoder processing, the stability and generalizability of the acoustic features are significantly improved, enhancing the speech synthesis system's adaptability to different contexts and making the resulting speech synthesis clearer, more natural, and more realistic.
[0065] S40, extracting prosody coding information of the reference style speech;
[0066] In this embodiment, extracting prosodic encoding information for the reference-style speech involves extracting prosodic features related to speech style, rhythm, stress, and emotion from the input speech data and encoding these features for subsequent use in controlling the prosodic performance of speech synthesis. Prosodic encoding information is a key factor influencing the naturalness and expressiveness of speech, ensuring that the generated speech not only matches the timbre of the target speaker but also accurately reflects the tone and emotional characteristics of the reference-style speech.
[0067] Prosodic encoding information typically includes key features such as fundamental frequency trajectory, phoneme duration distribution, and energy fluctuation characteristics. The fundamental frequency trajectory describes the pitch variation pattern of speech and is a core parameter for measuring speech prosody. The extracted fundamental frequency trajectory is typically smoothed using a short-time Fourier transform (STFT) or a Kalman filter to remove noise and jitter from the speech signal and make the fundamental frequency variations more coherent. For example, in emphatic sentences, the fundamental frequency trajectory typically has steep rises and falls, while in declarative sentences, the fundamental frequency trajectory is more gradual.
[0068] Phoneme duration distribution represents the duration of pronunciation of different phonemes and has a significant impact on speech rate and rhythm control. Phoneme duration can be calculated using dynamic time warping (DTW) or adversarial predictive coding (APC) to ensure that duration characteristics remain consistent across different speech rates. For example, in fast-paced speech, phoneme duration is generally short, while in emotional reading, the pronunciation duration of key phonemes is relatively extended to enhance expressiveness.
[0069] Energy fluctuation features describe the volume fluctuation trend of speech. They can be extracted through frame-level energy calculation or the Mel Energy Envelope. Speech energy fluctuations directly influence the emotional expression of speech. For example, angry speech typically exhibits large energy fluctuations, while calm speech exhibits more uniform energy variations. Furthermore, speech energy patterns vary in different contexts. For example, speech energy is typically low in telephone conversations, while it is higher and fluctuates more significantly in public speeches.
[0070] To ensure compatibility of the prosodic encoding information with the subsequent speech synthesis process, temporal alignment of the fundamental frequency trajectory, phoneme duration distribution, and energy fluctuation characteristics is required. Temporal alignment typically employs the Dynamic Time Warping (DTW) method to ensure that the prosodic features match those of the target text and target speaker's speech. Furthermore, temporal alignment can be achieved using a Temporal Alignment Network (TAN) based on an attention mechanism to improve alignment accuracy and ensure that the prosodic encoding information accurately guides speech generation.
[0071] Finally, the time-aligned prosodic features are fed into a pre-trained prosodic encoder for hierarchical encoding to extract a high-dimensional prosodic feature vector. The prosodic encoder can be trained using self-supervised learning to ensure it consistently extracts stable prosodic patterns across different speech styles. For example, a Transformer-based prosodic encoder can learn prosodic patterns across different speech styles through a multi-head attention mechanism and convert them into a unified encoding vector for use in the subsequent speech synthesis module.
[0072] By extracting the prosodic encoding information from the reference speech style, the generated speech accurately matches the prosodic features of the reference speech style, improving the expressiveness and personalized adaptability of speech synthesis. Extracting the fundamental frequency trajectory ensures accurate reproduction of the pitch variation pattern of the speech. Analysis of phoneme duration distribution optimizes speech rate control, and calculation of energy fluctuation characteristics enhances emotional expression. Timeline alignment and optimized prosodic encoder processing ensure that the prosodic encoding information maintains stability and generalization capabilities in different contexts, improving the overall quality of the speech synthesis system and making the generated speech more natural, fluent, and expressive.
[0073] S50, performing feature fusion on the phoneme feature sequence, acoustic features, and prosody coding information to generate a fused coding vector;
[0074] In this embodiment, phoneme feature sequences, acoustic features, and prosodic coding information are the three core feature inputs in the speech synthesis process, representing the textual content, timbre, and prosodic style of the speech, respectively. To generate natural and fluent speech output, these three features are fused to generate a fused coding vector. This allows the speech synthesis model to simultaneously consider text, timbre, and prosodic information, improving the naturalness and expressiveness of the generated speech.
[0075] The first step in feature fusion is to align feature data from different sources. Since the distribution of phoneme feature sequences, acoustic features, and prosodic coding information in the time dimension may be different, time axis alignment is required. The alignment method can use the Dynamic Time Warping (DTW) method or the Alignment Network based on the attention mechanism to ensure that the temporal relationship between different features can match. For example, when processing long text speech, the length of the phoneme sequence may be much larger than the time step of the acoustic feature, so alignment needs to be performed through interpolation or repeated padding.
[0076] After time alignment, the phoneme feature sequence and acoustic features are interactively modeled using a cross-attention mechanism to generate cross-features of phonemes and acoustics. The cross-attention mechanism establishes a correlation between phoneme and acoustic features, enabling the same phoneme to adapt to different pronunciation patterns in different acoustic contexts. For example, the same phoneme may correspond to different timbre characteristics depending on the speaking speed or emotional expression. The cross-attention mechanism can effectively capture these variations, improving the flexibility of speech synthesis.
[0077] Next, the cross-attention phoneme and acoustic features generated by the cross-attention are fused with the prosodic encoding information through a multi-head weighted fusion. The multi-head attention mechanism can simultaneously focus on prosodic information at different levels, allowing the fused features to preserve the matching relationship between phonemes and timbre while also incorporating prosodic information for adjustment. For example, when reading a sentence, certain words require specific prosodic emphasis. The attention mechanism can automatically adjust the acoustic features of these words, making the synthesized speech more consistent with natural human speech patterns.
[0078] During the fusion process, a residual connection module is used to add the attention-enhanced features to the original acoustic features to enhance feature stability. Residual connections effectively prevent information loss during feature fusion and enhance model training stability. For example, during long sentence speech synthesis, residual connections can preserve the target speaker's timbre while fine-tuning prosodic information to ensure consistent speech style.
[0079] Finally, the fused residual features are layer-normalized to ensure the stability of the feature distribution and eliminate scale differences between different feature data. Layer normalization improves the model's generalization capabilities, allowing the speech synthesis system to maintain stable sound quality and prosodic expression in different contexts. This generates high-quality fused encoding vectors, providing standardized input for subsequent speech synthesis.
[0080] By integrating features from phoneme feature sequences, acoustic features, and prosodic encoding information, the speech synthesis system can simultaneously consider text, timbre, and prosodic information, improving the naturalness and expressiveness of speech synthesis. Timeline alignment optimizes the matching relationship between different feature data, cross-attention enhances the interactive modeling of phoneme and acoustic features, and the multi-head attention mechanism enables prosodic information to dynamically adjust the intonation, rhythm, and emotional expression of speech. Residual connections and layer normalization enhance the stability of feature fusion, ensuring a closer match between the timbre, prosody, and phonemes in the generated speech, improving the quality of speech synthesis and user experience.
[0081] S60, processing the fused code vector through a style adaptation module to generate an acoustic code containing target style features;
[0082] In this embodiment, the role of the style adaptation module is to introduce target style features based on the fused coding vector, so that the final generated speech not only matches the timbre characteristics of the target speaker, but also meets the requirements of the target style in terms of intonation, rhythm, emotion, etc. The fused coding vector contains multimodal information such as phoneme features, acoustic features, and prosodic coding information. However, the fusion of these features only provides basic speech synthesis information and still lacks personalized style adaptation capabilities. Therefore, it is necessary to further adjust the speech style through the style adaptation module to make it more natural and consistent with the target context.
[0083] The first step in the style adaptation module is to predefine multiple style weight matrices during the training phase. Each style weight matrix corresponds to a specific speech style, such as formal, casual, excited, calm, and so on. The style weight matrices can be used to learn the statistical characteristics of different speech styles through pre-trained models, allowing the system to flexibly adapt to different speech synthesis requirements during the inference phase. For example, in customer service speech synthesis, a formal style may require a slower speaking rate, clear pronunciation, and a stable pitch, while a casual style may include a faster speaking rate, richer intonation, and natural linking.
[0084] Next, the cosine similarity between the fused encoding vector and each style weight matrix is calculated to generate a style matching score. Cosine similarity is a mathematical method for measuring the similarity between two vectors and can effectively determine the degree of match between the fused encoding vector and a predefined style. By calculating cosine similarity, we can determine which style the input speech is most similar to, providing a reference for the subsequent style adaptation process. For example, in a debate scenario, if the characteristics of the fused encoding vector best match the style weight matrix of the "strong confrontation" style, the system may choose to enhance the accent and speed variation in the speech to improve the effectiveness of the rebuttal.
[0085] After obtaining the style match score, the system selects the style weight matrix with the highest match as the dominant style parameter and uses it as the primary style control vector for the current speech generation. Because different application scenarios may have different requirements for speech style, the selection of dominant style parameters depends not only on the speech data itself but can also be adjusted based on user input, environmental factors, or contextual information. For example, in a healthcare voice assistant, the dominant style parameters can automatically adjust based on the patient's emotional state, using a gentler voice style when the patient is anxious and a more authoritative voice style when providing important health advice.
[0086] To apply the dominant style parameters to the fused coding vector, a gating mechanism is used for weighted fusion. This mechanism dynamically adjusts the weights of the input signals, adapting the fused coding vector to different stylistic characteristics so that the generated speech accurately reflects the target style. For example, in news broadcast scenarios, the gating mechanism can enhance tonal stability, making the synthesized speech more consistent with the style of news broadcasts. In entertainment dubbing, the gating mechanism can introduce greater pitch fluctuations and speech rate variations to enhance expressiveness.
[0087] Finally, the weighted fused features are processed using a nonlinear activation function to ensure that the output acoustic encoding is suitable for the subsequent speech synthesis process. Nonlinear activation functions (such as GELU or ReLU) can enhance the model's expressiveness, ensuring that the final acoustic encoding not only incorporates the target style characteristics but also maintains the effective information of the input features, improving the overall quality of speech synthesis.
[0088] By processing the fused encoding vectors through the style adaptation module, speech synthesis not only matches the target speaker's timbre but also flexibly adjusts the speech's stylistic characteristics to suit different application scenarios. A predefined style weight matrix improves the operability of style control, cosine similarity calculation ensures the accuracy of style selection, a gating mechanism enhances the stability of style adaptation, and a nonlinear activation function improves the model's expressiveness. Overall, this solution improves the naturalness, personalization, and adaptability of speech synthesis, enabling synthesized speech to achieve more desirable expressiveness in a variety of scenarios.
[0089] S70: Generate mel-spectrogram features according to the acoustic coding, and input the mel-spectrogram features into a pre-trained vocoder to generate a speech waveform.
[0090] In this embodiment, the process of generating mel-spectrogram features based on acoustic coding and inputting them into a pre-trained vocoder to generate a speech waveform involves spectral analysis of the speech signal and neural network-driven waveform synthesis. Mel-spectrogram features are a speech representation method based on the human auditory characteristics, which can improve the naturalness and clarity of speech synthesis. The pre-trained vocoder is responsible for converting discrete spectral information into a continuous audio signal to generate the final speech waveform.
[0091] In the process of generating Mel-spectrogram features, the acoustic code must first be transformed into a time-frequency transform to extract the characteristics of the speech signal at different frequency components. The time-frequency transform can use the short-time Fourier transform (STFT) to convert the time series data into a frequency domain representation for subsequent Mel-scale mapping. The short-time Fourier transform decomposes the input signal through a sliding window, allowing the signal to be represented in both the time domain and the frequency domain. The spectral energy is then mapped to the Mel scale through a Mel filter bank (MelFilterBank) to match the human ear's perception of different frequencies. For example, in the low-frequency range, the human ear is more sensitive to frequency changes, so the Mel filter has a higher resolution, while in the high-frequency range, the Mel filter has a relatively low resolution.
[0092] The generated Mel-spectrogram features contain multiple time frames, each of which corresponds to spectral information within a short time window. To ensure that the Mel-spectrogram features can accurately reflect the timbre, rhythm, and emotional information of the speech, they need to be normalized to eliminate differences in volume and spectral dynamic range between different speech samples. Feature normalization methods can include mean normalization, log transformation, or batch normalization, which make the dynamic range of the Mel-spectrogram more adaptable to the input distribution of the neural network, thereby improving the stability and generalization ability of the speech synthesis system.
[0093] After obtaining the Mel-spectrogram features, they are fed into a pre-trained vocoder to generate the final speech waveform. A vocoder is a model used to convert discrete spectral data into a continuous audio signal. Common vocoders include WaveNet, WaveGlow, and HiFi-GAN. Different vocoder models vary in computational efficiency and synthesis quality. For example, WaveNet uses an autoregressive structure, resulting in high synthesis quality but slow computational speed, while HiFi-GAN uses a generative adversarial network (GAN) structure, offering fast synthesis speed and near-natural speech quality.
[0094] The input to a vocoder typically includes Mel-spectrogram features and additional acoustic control parameters such as pitch, energy, and duration. During inference, the vocoder predicts the amplitude and phase information of the speech waveform based on the Mel-spectrogram features and generates a high-quality waveform signal through an inverse Fourier transform or direct neural network operations. For example, in a WaveGlow-based vocoder, the input Mel-spectrogram features are transformed through a series of normalizing flows, making the output audio waveform smoother and more natural.
[0095] To improve the synthesis performance of the vocoder, fine-tuning can be performed based on the target speaker's timbre characteristics, making the resulting speech closer to the target speaker's natural pronunciation in timbre and prosody. For example, in a voice cloning task, transfer learning can be performed on a pre-trained vocoder using a small amount of the target speaker's speech data, making it more adaptable to the target speaker's timbre characteristics.
[0096] Example: In a healthcare voice interaction system, to assist patients whose speech abilities are impaired due to illness or surgery with speech rehabilitation, the system first obtains the text content entered by the patient, the patient's own historical speech data, and standardized reference speech provided by medical experts. The system then extracts the phoneme feature sequence of the text to ensure that the speech synthesis process matches the patient's language habits. At the same time, acoustic features are extracted from the patient's historical speech data to ensure that the generated speech matches the patient's timbre, speaking speed, and vocalization style. Furthermore, the system extracts prosodic coding information from the reference speech to ensure that the synthesized speech conforms to natural speech habits in terms of pronunciation rhythm, stress placement, and other aspects.
[0097] After feature extraction, the system fuses features from these different sources to generate a fused coding vector. This process includes timeline alignment, cross-attention modeling, multi-head attention fusion, and residual connections to ensure that phoneme, acoustic, and prosodic features remain consistent in the temporal dimension and can complement each other to optimize the fluency of the synthesized speech. Subsequently, the system processes the fused coding vector through a style adaptation module to generate an acoustic code containing the target style features. In this process, the system matches the style weight matrix that best matches the patient's individual voice characteristics, and performs dynamic weighted fusion through a gating mechanism to ensure that the generated speech not only meets the patient's voice characteristics, but also draws on the expression methods of standardized medical speech.
[0098] Finally, the system generates Mel-spectrogram features based on acoustic coding, and generates speech waveforms through a pre-trained vocoder, so that patients can hear speech feedback that matches their own timbre. After multiple interactions with the patient, the system will also optimize the synthesized speech style to improve the match between the generated speech and the patient's natural voice. For example, if the system detects that the generated speech differs from the patient's original speech style in terms of speaking speed or stress position, the system can calculate the similarity index of the style vector and use the style adaptation loss to optimize the style weight matrix to gradually improve the personalization of speech generation. This method can help patients adapt to speech rehabilitation training more quickly and improve rehabilitation efficiency.
[0099] In a robo-advisory voice broadcast system, financial institutions need to provide investors with personalized market analysis and reports to enhance the user experience and improve investment decision-making efficiency. The system can generate personalized robo-advisory voice broadcasts tailored to the preferences of different investors. For example, the system first obtains the target text for the broadcast, including the latest market trends, industry news, and investment advice. It also collects historical broadcast content heard by the user and standardized financial analyst voice data as a reference style.
[0100] The system then extracts the text's phoneme feature sequences to ensure the speech synthesis system accurately represents professional terminology and key values. It also extracts acoustic features from the user's historical broadcast data, such as their preferred speaking speed and timbre, to generate speech output that better suits their habits. Furthermore, prosodic coding information is extracted from standard speech data of financial analysts to ensure that the synthesized speech, in terms of emotional expression and intonation, conforms to the style of professional financial broadcasts.
[0101] During the feature fusion phase, the system uniformly processes the aforementioned features to generate a fused encoding vector. The system then uses a style adaptation module to match the reporting style most suited to user preferences. For investors seeking robust market analysis, the system can prioritize a more stable voice style weight matrix, ensuring a moderate speech speed and smooth intonation, thereby enhancing user trust in the market analysis. For investors preferring short-term trading and focusing on market trends, the system can select a more compact and rhythmic voice style, adjusting it through a gating mechanism to ensure that the reporting content highlights market fluctuations and enhances the urgency of information delivery.
[0102] After the speech broadcast is complete, the system also calculates the style vector of the generated speech and compares it with the user's historical preferences to optimize the adaptability of the broadcast style. For example, if a user has long preferred an analyst voice with more emotional fluctuations, while the currently generated speech is relatively stable, the system will optimize the weight parameters of the style adaptation module using the style adaptation loss, making the subsequent broadcast more consistent with the user's expectations, thereby improving the user's interactive experience.
[0103] In speech synthesis debate training systems, this technology can be used to synthesize more realistic rebuttal speech, enabling the training system to simulate different debate styles and help debaters improve their on-the-spot reaction and debating skills. For example, the system first obtains the target debate text, including the point to be refuted and the debater's pre-set response. It also collects historical speech data from debaters and professional debaters as stylistic references.
[0104] During the feature extraction phase, the system extracts the text's phoneme feature sequences to ensure that the synthesized speech accurately conveys the debate content. Simultaneously, acoustic features are extracted from the debaters' historical speech data to ensure that the synthesized speech retains the debaters' timbre. Furthermore, prosodic coding information is extracted from professional debaters' speech data to learn common speech features such as speech rate changes and stress distribution during debates, ensuring that the generated speech is more realistic and authentic to real debates.
[0105] During the feature fusion phase, the system deeply integrates phoneme features, acoustic features, and prosodic encoding information, and uses a style adaptation module to match the speech style that best suits the current debate strategy. For example, when making targeted rebuttals, the system can choose a faster speech style with prominent accents to enhance aggressiveness; when strategically guiding, the system can choose a style with a moderate speech rate and steady intonation to enhance persuasiveness. Furthermore, a gating mechanism adjusts the degree of style adaptation, allowing the system to flexibly switch between different debate strategies, improving the adaptability of training.
[0106] After speech generation is complete, the system calculates the style vector of the generated speech and compares it with the target debater's style vector to optimize the speech style match. For example, if a debater prefers a more rational and analytical debate style, while the system-generated speech tends to be more emotional, the system will optimize the weight parameters of the style adaptation module using style adaptation loss, making the subsequent synthesized speech more consistent with the debater's personal style and improving the personalized effect of debate training.
[0107] By generating Mel-spectrogram features based on acoustic coding and inputting them into a pre-trained vocoder to generate speech waveforms, the speech synthesis system can efficiently convert discrete speech features into continuous audio signals, improving the naturalness and sound quality of the synthesized speech. The time-frequency analysis of the Mel-spectrogram features ensures timbre stability and clarity, while the use of a pre-trained vocoder significantly improves the fluency and computational efficiency of speech synthesis, enabling the system to achieve high-quality speech output in a variety of application scenarios.
[0108] The present invention relates to the field of speech synthesis technology and can be applied in business scenarios such as healthcare, financial technology, and debate. A speech generation method based on speech style adaptation is disclosed, comprising: obtaining a target text, a target speaker's speech, and a reference style speech; extracting a phoneme feature sequence of the target text, acoustic features of the target speaker's speech, and prosodic coding information of the reference style speech; fusing the phoneme feature sequence, acoustic features, and prosodic coding information to generate a fused coding vector; processing the fused coding vector through a style adaptation module to generate an acoustic code containing the target style features; generating mel-spectrogram features based on the acoustic code, and inputting the mel-spectrogram features into a pre-trained vocoder to generate a speech waveform. By fusing the phoneme features of the target text, the acoustic features of the target speaker, and the prosodic coding information of the reference style speech, the present invention makes the generated speech more natural and expressive in terms of timbre matching, prosodic adjustment, and emotional expression; optimizing the speech style mapping through the style adaptation module so that the generated speech can adapt to specific scenario requirements and improve generalization capabilities for unseen speakers; and processing based on the mel-spectrogram features and the pre-trained vocoder to achieve higher clarity and stability in speech generation, thereby improving the adaptability and practical value of the speech synthesis system.
[0109] In one embodiment, the above step S20 includes:
[0110] S201, segmenting the target text into phoneme symbol sequences using a pre-trained phoneme segmentation model;
[0111] S202, inputting the phoneme symbol sequence into an embedding layer to generate a phoneme embedding vector;
[0112] S203, performing position encoding processing on the phoneme embedding vector to generate a time series vector;
[0113] S204. Input the timing vector into a bidirectional long short-term memory network to generate phoneme context-related features;
[0114] S205. Perform tensor concatenation on the phoneme context-related features and the phoneme embedding vectors to generate a concatenated feature vector;
[0115] S206. Perform layer normalization on the concatenated feature vector to generate the phoneme feature sequence.
[0116] In this embodiment, extracting the phoneme feature sequence of the target text is a key step in the speech synthesis system, which converts text information into a phoneme-level representation that can be used for speech generation. Phonemes are the smallest units of pronunciation in speech, and the accuracy of the phoneme feature sequence determines the clarity and naturalness of the final synthesized speech. This step involves text preprocessing, phoneme segmentation, feature embedding, temporal information modeling, and feature optimization to ensure that the phoneme feature sequence can fully express the text content and adapt to different speech styles and contexts.
[0117] First, the target text needs to be processed by a pre-trained phoneme segmentation model to convert the text into a sequence of phoneme symbols. The role of the phoneme segmentation model is to divide continuous text into corresponding phoneme sequences according to the phonetic rules of the language. For example, in English, the word "speech" can be converted into / s / / p / And in Chinese pinyin, "中国" can be converted into / kuo / . The phoneme segmentation model usually adopts a neural network model, such as a sequence annotation model based on Transformer or a BiLSTM-CRF structure, to ensure the accuracy and consistency of phoneme segmentation.
[0118] After phoneme segmentation is completed, the sequence of phoneme symbols is input into the embedding layer to generate phoneme embedding vectors. The role of the embedding layer is to map discrete phoneme symbols into a high-dimensional vector space, so that the relationships between different phonemes can be learned by a deep learning model. For example, phonemes with similar pronunciations can be close in the vector space to improve the model's adaptability to homophones or liaison phenomena. Phoneme embedding can use trained static embeddings (such as the Word2Vec method) or dynamic context-aware embeddings (such as Transformer Embeddings) to enhance its applicability in different contexts.
[0119] To preserve the timing information of phonemes in a sequence, positional encoding is performed on the phoneme embedding vector to generate a time sequence vector. Since the phoneme sequence in text input is discrete, positional encoding is used to incorporate information about the relative position of phonemes in a sentence. For example, in a Transformer-based speech synthesis system, sine and cosine positional encoding can be used to enhance the temporal information of the sequence, enabling the model to identify which phonemes should be read together and which require pauses.
[0120] Next, the time series vector is fed into a bidirectional long short-term memory network (BiLSTM) to generate phoneme contextual features. BiLSTM uses bidirectional propagation to enable the representation of each phoneme to incorporate contextual information, thereby enhancing speech fluency. For example, in English, the word "read" may be pronounced differently in different contexts ( or / rεd / ), BiLSTM can more accurately predict its phoneme representation by learning the relationship between the previous and next phonemes, thereby improving the accuracy of the synthesized speech.
[0121] The phoneme contextual features are then tensor-concatenated with the original phoneme embedding vector to generate a concatenated feature vector. This operation aims to preserve the original phoneme representation while integrating contextual information, ensuring that the generated phoneme features contain both basic pronunciation information and adaptability to the context. For example, in Chinese speech synthesis, variations in tone may require additional contextual information to ensure speech accuracy, and tensor concatenation can enhance this learning capability.
[0122] Finally, layer normalization is performed on the concatenated feature vectors to generate the final phoneme feature sequence. Layer normalization standardizes different speech inputs, ensuring a stable dynamic range of phoneme features and reducing the impact of varying text inputs on the model. For example, during training, input texts of varying lengths may lead to an uneven distribution of phonemes. Layer normalization can reduce this imbalance, improving the stability and generalization of speech synthesis.
[0123] This embodiment extracts the target text's phoneme feature sequence, enabling efficient conversion of text information into the basic phoneme units required for speech synthesis. It also optimizes phoneme representation by incorporating contextual information, improving the accuracy and naturalness of speech synthesis. The phoneme segmentation model ensures accurate text-to-phoneme conversion, while the phoneme embedding vector enhances phoneme representation. BiLSTM provides contextual awareness, and layer normalization improves the stability of phoneme features, resulting in a more coherent and smooth speech synthesis process.
[0124] In one embodiment, the above step S30 includes:
[0125] S301, extracting the fundamental frequency features of the target speaker's speech through a linear prediction coding module;
[0126] S302, determining the Mel spectrum envelope of the target speaker's speech by short-time Fourier transform;
[0127] S303, detecting the formant frequency distribution of the target speaker's speech;
[0128] S304, performing tensor concatenation on the fundamental frequency feature, the Mel spectrum envelope, and the formant frequency distribution to generate a concatenated acoustic tensor;
[0129] S305: Input the concatenated acoustic tensor into a pre-trained acoustic encoder to generate a joint acoustic feature matrix.
[0130] In this embodiment, extracting the acoustic features of the target speaker's speech is a key step in the speech synthesis process, directly determining the timbre, rhythm, and quality of the synthesized speech. Acquiring acoustic features involves multiple layers of audio signal processing, including fundamental frequency analysis, spectral modeling, and formant extraction. Ultimately, these features are represented through a deep learning model, enabling the speech synthesis system to accurately simulate the target speaker's speech characteristics.
[0131] First, the fundamental frequency characteristics of the target speaker's speech need to be extracted through the Linear Predictive Coding (LPC) module. Fundamental frequency (F0) is an important parameter that determines the pitch of speech and has a significant impact on the naturalness, emotional expression, and timbre of speech. LPC performs linear prediction analysis on the speech signal, establishes a linear relationship between the current speech frame and the historical speech frames, and calculates the fundamental frequency change curve of the speech signal. For example, the fluctuation pattern of the fundamental frequency is different in different emotional expressions. For example, the fundamental frequency fluctuates greatly when angry, while it is relatively stable when calm. LPC can estimate the fundamental frequency by solving the prediction coefficient and eliminate noise interference to improve the stability of the fundamental frequency characteristics.
[0132] After obtaining the fundamental frequency features, the Mel-spectral envelope of the target speaker's speech needs to be determined through a short-time Fourier transform (STFT). STFT is a transform method used to analyze the time-varying spectral characteristics of a signal. It extracts spectral information from different time periods by dividing the speech signal into short time windows and applying a Fourier transform to each window. The Mel-spectral envelope is a spectral representation based on the Mel scale that better reflects the human ear's perception of different frequency components. The spectral information calculated by the STFT is processed by a Mel filter bank to obtain the Mel-spectral envelope of the target speaker's speech, enabling the subsequent speech synthesis model to better capture the timbre characteristics. For example, in news broadcast scenarios, the Mel-spectral envelope can ensure the clarity of the synthesized speech, while in singing synthesis applications, the Mel-spectral envelope can preserve the coherence and musicality of the timbre.
[0133] Next, the formant frequency distribution (Formant Frequency Distribution) of the target speaker's speech needs to be detected. Formants are frequency components determined by the resonant characteristics of the vocal tract and play a vital role in speech recognition and synthesis. Formants can be extracted through linear prediction analysis (LPA) or cepstral analysis. For example, the first formant (F1) is usually related to the degree of opening of the phoneme, while the second formant (F2) determines the front and back tongue position of the phoneme. The formant frequency distribution of different speakers has personalized characteristics, so in voice cloning and style transfer tasks, formant detection can ensure that the timbre of the synthesized speech matches that of the target speaker. For example, in intelligent voice navigation systems, male speakers generally have lower F1 and F2, while female speakers have relatively higher F1 and F2, so the system can use formant information for personalized synthesis.
[0134] After obtaining the fundamental frequency features, Mel-spectrogram envelope, and formant frequency distribution, these features need to be concatenated to generate a concatenated acoustic tensor. Tensor concatenation is a data fusion method that concatenates features from different sources by channel or time axis to ensure that the model can fully utilize all acoustic features. For example, in a multi-task speech synthesis model, concatenated acoustic tensors can combine timbre, rhythm, and emotional information, making the generated speech richer and more delicate.
[0135] Finally, the concatenated acoustic tensor is fed into a pre-trained acoustic encoder to generate a joint acoustic feature matrix. The acoustic encoder can employ structures such as a deep neural network (DNN), a variational autoencoder (VAE), or a residual network (ResNet) to extract high-level speech features. For example, a DNN acoustic encoder can reduce the dimensionality of the concatenated acoustic tensor and generate a stable feature representation, while a VAE acoustic encoder can improve the diversity of speech synthesis through probabilistic modeling, ensuring that speech remains natural in different scenarios.
[0136] This embodiment extracts the acoustic features of the target speaker's speech, enabling the speech synthesis system to accurately simulate the target speaker's timbre, prosody, and emotional characteristics, thereby improving the personalization and naturalness of the synthesized speech. Extraction of fundamental frequency features ensures pitch consistency, calculation of the Mel-spectrogram envelope optimizes timbre, and analysis of formant frequency distribution enhances the personalized expression of the speech. Through tensor concatenation and acoustic encoder processing, the acoustic features maintain stability and generalization capabilities in different contexts, improving the adaptability of the speech synthesis system and making the generated speech clearer, more natural, and more realistic.
[0137] In one embodiment, the above step S40 includes:
[0138] S401, detecting a fundamental frequency trajectory curve of the reference style speech;
[0139] S402, counting the start boundary point and the end boundary point of each phoneme in the reference style speech to generate a phoneme-level duration distribution;
[0140] S403, extracting sentence-level energy fluctuation features of the reference style speech;
[0141] S404, performing time axis alignment processing on the fundamental frequency trajectory curve, phoneme-level duration distribution, and sentence-level energy fluctuation characteristics;
[0142] S405 , hierarchically encoding the aligned fundamental frequency trajectory curve, phoneme-level duration distribution, and sentence-level energy fluctuation features through a pre-trained prosody encoder to generate the prosody coding information.
[0143] In this embodiment, extracting prosodic information from the reference-style speech is a crucial step in a speech synthesis system for controlling speech rhythm, stress distribution, and emotional expression. Prosodic information consists of multiple features, including fundamental frequency trajectory curves, phoneme-level duration distribution, and sentence-level energy fluctuations. The extraction and encoding of these features ensures that the generated speech not only resembles the target speaker in timbre but also meets specific stylistic requirements in intonation and rhythm.
[0144] First, it is necessary to detect the fundamental frequency trajectory curve of the reference style speech. The fundamental frequency (F0) represents the pitch change of the speech and has an important impact on the naturalness and emotional expression of the speech. The extraction of the fundamental frequency trajectory curve usually adopts the short-time Fourier transform (STFT), Kalman filter or autoregressive fundamental frequency detection algorithm (Autoregressive F0 Estimation) to remove noise and smooth the fundamental frequency changes. For example, in angry speech, the fundamental frequency trajectory curve usually has greater fluctuations, while in declarative sentences, the fundamental frequency trajectory is relatively stable. Therefore, the extraction of the fundamental frequency trajectory curve can be used to predict the emotional state of the speech and serve as a key input for prosodic coding.
[0145] Next, it is necessary to count the starting and ending boundary points of each phoneme in the reference style speech to generate a phoneme-level duration distribution. Phoneme-level duration information determines the speaking rate and rhythm, and is crucial for the coherence of synthesized speech. For example, in speech-style speech, the duration of long vowels is usually prolonged, while in daily conversations, the duration of vowels is shorter and more uniform. The extraction of phoneme duration distribution usually relies on an alignment model (AlignmentModel), such as the Transformer alignment network based on the attention mechanism or dynamic time warping (DTW), to ensure the accuracy of phoneme-level duration information.
[0146] Subsequently, it is necessary to extract the sentence-level energy fluctuation features of the reference style speech. Energy fluctuation describes the volume change trend of the speech and can be obtained by calculating the short-term energy (STE) or the Mel Energy Envelope. For example, in emotional reading, the volume of certain key words may increase, while in low-pitched expressions, the overall energy is low and the fluctuation is small. By extracting sentence-level energy fluctuation features, the system can learn the energy change patterns of different styles of speech, making the synthesized speech more natural in emotional expression.
[0147] After extracting the above-mentioned prosodic features, timeline alignment is required to ensure that the fundamental frequency trajectory curve, phoneme-level duration distribution, and sentence-level energy fluctuation characteristics are accurately matched in time sequence. Timeline alignment can use dynamic time warping (DTW) or attention-based alignment to ensure the synchronization of different prosodic features. For example, in a speech conversion task, the fundamental frequency variation of the target speech may not match that of the original speech. Timeline alignment can ensure that the converted speech is rhythmically consistent with the target style.
[0148] Finally, the aligned fundamental frequency trajectory curve, phoneme-level duration distribution, and sentence-level energy fluctuation features are input into a pre-trained prosody encoder for hierarchical encoding to generate prosody encoding information. Prosody encoders typically use variational autoencoders (VAEs), Transformer encoders, or pre-trained models based on self-supervised learning to ensure that they can generalize to different speech styles. For example, a Transformer-based prosody encoder can use a multi-head attention mechanism to learn the prosody patterns between different speech styles, allowing the encoded prosody information to be flexibly adapted to different speech synthesis tasks.
[0149] This embodiment extracts prosodic coding information from the reference style speech, enabling the speech synthesis system to better control the rhythm, prosody, and emotional expression of the speech, thereby improving the naturalness and adaptability of the synthesized speech. The extraction of the fundamental frequency trajectory curve ensures the pitch variation pattern of the speech, the statistics of the phoneme-level duration distribution optimize the speech rate control, and the calculation of the sentence-level energy fluctuation characteristics enhances the emotional expression ability. The time axis alignment and the optimization of the prosodic encoder ensure that the prosodic coding information maintains stability and generalization capabilities in different contexts, improving the overall quality of the speech synthesis system and making the generated speech more natural, fluent, and expressive.
[0150] In one embodiment, the above step S50 includes:
[0151] S501, performing time axis alignment processing on the phoneme feature sequence, acoustic features and prosody coding information;
[0152] S502, interactively modeling the aligned phoneme feature sequence and acoustic features through a cross-attention mechanism to generate phoneme and acoustic cross features;
[0153] S503, performing multi-head attention weighted fusion on the phoneme and acoustic cross features and the prosody coding information to generate an attention enhancement feature;
[0154] S504, adding the attention enhancement feature to the acoustic feature through a residual connection module to generate a residual fusion feature;
[0155] S505: Perform layer normalization processing on the residual fusion feature to generate the fusion coding vector.
[0156] In this embodiment, the purpose of fusing phoneme feature sequences, acoustic features, and prosodic coding information in a speech synthesis system is to integrate speech features at multiple levels to generate a fused coding vector with complete speech expression capabilities. This ensures that the synthesized speech not only accurately expresses the text content, but also matches the timbre characteristics of the target speaker and possesses a reasonable speech prosody and style. This step primarily involves several key technologies, including timeline alignment, cross-attention mechanism, multi-head attention weighted fusion, residual connections, and layer normalization, to ensure that the fused speech features achieve high stability and adaptability in terms of time dimension, semantic hierarchy, and style control.
[0157] First, the phoneme feature sequences, acoustic features, and prosodic coding information need to be time-aligned. Since phoneme features are extracted based on text, while acoustic features and prosodic coding information are analyzed from the target speaker's speech, their time scales may not match completely. Therefore, they need to be time-aligned through dynamic time warping (DTW), attention-based alignment, or end-to-end implicit alignment mechanisms (such as the FastSpeech structure) to ensure that features from different sources are calculated synchronously at the same time step. For example, in the case of long vowel pronunciation, the phoneme features may be shorter, while the acoustic features and prosodic information contain longer time frames. Through alignment processing, the time length of the phoneme features can be adjusted to match the actual pronunciation time.
[0158] After completing the timeline alignment, the aligned phoneme feature sequence and acoustic features need to be interactively modeled through a cross-attention mechanism to generate phoneme and acoustic cross-features. The cross-attention mechanism can learn the correlation between different features and enhance the mutual connection between phoneme information and acoustic information. For example, when synthesizing certain complex syllables, the acoustic features may contain specific formant changes, and the phoneme features themselves may not be able to capture such changes. The cross-attention mechanism can help the phoneme features better adapt to the timbre of the target speaker, thereby improving the naturalness of speech synthesis. For example, in the phenomenon of tone change in Chinese, the annotation of the phoneme may only be a single pinyin, but the acoustic features can reflect the actual tone change. The cross-attention mechanism can capture such changes, making the matching of phoneme features and acoustic features more accurate.
[0159] After obtaining the cross-features of phonemes and acoustics, they need to be fused with the prosodic encoding information through a multi-head attention weighted fusion to generate attention-enhanced features. The multi-head attention mechanism can learn different levels of information from different attention heads. For example, some attention heads can focus on the stress characteristics of speech, while others can focus on changes in speaking rate. The final fusion results in enhanced features containing richer information. For example, in emotional speech synthesis tasks, multi-head attention can help the system capture the specific speech rhythm of the target speaker, making the synthesized speech more consistent with the prosodic expression of real speech.
[0160] Next, the attention-enhanced features need to be added to the original acoustic features through the residual connection module to generate residual fusion features. The role of the residual connection is to avoid information loss during feature fusion while maintaining the stability of the original acoustic features. Especially in deep neural networks, directly superimposing features from different sources may lead to overfitting or loss of information, while the residual connection can ensure the integrity of the original features while introducing new information to improve the robustness of the fused features. For example, in the voice cloning task, if feature fusion relies entirely on the attention mechanism, the target timbre may be offset, while the residual connection can ensure the stability of the timbre so that the synthesized speech still has the characteristics of the target speaker.
[0161] Finally, layer normalization is performed on the residual fusion features to generate the final fused encoding vector. Layer normalization standardizes the features of different speech samples, ensuring a consistent dynamic range across different speech inputs. This prevents the model from being affected by specific speech characteristics and improves the system's generalization capabilities. For example, in a multi-speaker speech synthesis system, the volume and speaking rate of different speakers may vary significantly. Layer normalization can reduce the impact of these differences, allowing the speech synthesis system to more stably adapt to different speech styles.
[0162] This embodiment fuses phoneme feature sequences, acoustic features, and prosodic coding information, enabling the speech synthesis system to more precisely control the timbre, prosody, and style of speech, improving the naturalness and adaptability of synthesized speech. Timeline alignment ensures the consistency of different features across the temporal dimension, the cross-attention mechanism enhances the interactivity between phoneme and acoustic information, multi-head attention fusion improves the expressiveness of prosodic information, residual connections maintain the integrity of the original acoustic features, and layer normalization optimizes feature stability, making the synthesized speech more natural, clear, and consistent with the characteristics of the target speaker.
[0163] In one embodiment, the above step S60 includes:
[0164] S601, predefining a plurality of style weight matrices in the style adaptation module, each style weight matrix corresponding to a predefined style type;
[0165] S602, determining the cosine similarity between the fused coding vector and each style weight matrix, and generating a style matching score;
[0166] S603, selecting the style weight matrix with the highest style matching score as the dominant style parameter;
[0167] S604, performing weighted fusion processing on the dominant style parameter and the fused coding vector through a gating mechanism to generate a weighted fusion result;
[0168] S605 , performing nonlinear activation function processing on the weighted fusion result to generate an acoustic code containing target style features.
[0169] In this embodiment, processing the fused coding vector through the style adaptation module to generate an acoustic code containing the target style features is a key step in achieving speech style transfer and personalized speech synthesis in a zero-shot speech synthesis system. This process aims to select the most matching style features based on the style requirements of the target speech and optimize them in combination with the fused coding vector to ensure that the generated speech conforms to the characteristics of the target style in terms of style expression, intonation, prosody, etc. This step mainly involves key links such as pre-definition of the style weight matrix, calculation of style matching, style selection, parameter fusion, and nonlinear activation to ensure that the style features can be efficiently adapted to the target speech.
[0170] First, multiple style weight matrices are predefined in the style adaptation module, and each style weight matrix corresponds to a predefined style type. The style weight matrix is used to store feature vectors of different voice styles, such as news broadcast style, commercial advertising style, storytelling style, academic speech style, etc. The style weight matrix can be obtained through training with large-scale voice data, and a self-supervised learning-based method, such as using variational autoencoders (VAE) or generative adversarial learning (GAN) for style embedding learning, can be used to ensure that the features between different styles can be effectively distinguished. For example, in a voice assistant application, the system may predefine multiple styles, such as formal intonation, casual intonation, motivational intonation, etc. Each style weight matrix corresponds to a specific voice style, allowing the system to adapt to different user needs.
[0171] Next, we need to calculate the cosine similarity between the fused encoding vector and each style weight matrix and generate a style matching score. Cosine similarity is a common vector similarity measurement method that evaluates the similarity between two vectors by calculating the cosine value of the angle between them. The calculation formula is as follows:
[0172]
[0173] Here, x represents the fused encoding vector, and y represents the vector representation of a style weight matrix. By calculating the cosine similarity between different style weight matrices and the fused encoding vector, we can quantify the style type that the current input speech is most similar to. For example, in certain speech synthesis tasks, if the user input text contains strong emotional expressions (such as anger), the system can calculate its similarity with different style weight matrices and select the most matching emotional style for synthesis.
[0174] Then, the style weight matrix with the highest score from multiple style matching scores is selected as the dominant style parameter. The dominant style parameter represents the style information that best matches the current speech generation process. This selection process can adopt a top-1 selection method, directly selecting the style with the highest similarity; or a weighted average strategy can be used to combine the weight matrices of multiple similar styles to produce a smoother and more natural style blending effect. For example, when generating speech-style speech, the system can select "formal style" or "news broadcast style" as the dominant style parameter based on the style matching degree to ensure that the intonation of the synthesized speech is more consistent with the speech context.
[0175] After selecting the dominant style parameters, a gating mechanism is used to perform a weighted fusion of the dominant style parameters and the fused encoding vector to generate a weighted fusion result. The gating mechanism is a deep learning technique used to control information flow. It dynamically weights the input vectors using trainable parameters to ensure smooth information fusion. Specifically, the gating mechanism can use a Sigmoid function or a Softmax function to calculate the weight values:
[0176] g=σ(Wx+b)
[0177] output=g·x+(1-g)·y
[0178] Where g (gating parameter) controls the weighted ratio between the fused encoding vector x and the style weight matrix y, with a value range of (0, 1). This value determines the degree of mixing of the target style information and the original timbre information in the generated acoustic encoding.
[0179] σ (Sigmoid activation function): used to normalize the calculation results to the (0, 1) interval, so that g can be used as a dynamic weight coefficient for weighted fusion.
[0180] W (Trainable Weight Matrix): Represents the weight matrix used to calculate the gating parameters. W is a trainable matrix used to learn how to adjust the influence of various input features during the style adaptation process. It determines the importance of different style types in speech synthesis. If x has dimension d, then W is typically d×d to ensure linear transformation of the input features.
[0181] x (fused encoding vector): represents the fused feature vector of phoneme features, acoustic features, and prosodic encoding information after processing through cross-attention and residual connections. It contains information about the input text, the timbre characteristics of the target speaker, and adaptation information to the reference style.
[0182] b (bias term): used to adjust the input of the activation function so that the model can more flexibly learn adaptation strategies suitable for different speech styles.
[0183] The gating mechanism enables flexible adjustment of different style information, ensuring that the final synthesized speech retains the fundamental characteristics of the original speech while adaptively incorporating the target style. For example, in commercial speech synthesis, the gating mechanism can appropriately increase the energy variation of the speech, making it more appealing, while in academic speech synthesis, it can reduce the emotional fluctuation of the speech, making it more formal and clear.
[0184] Finally, the weighted fusion result is processed using a nonlinear activation function to generate the final acoustic encoding of the target style features. Nonlinear activation functions can use methods such as ReLU (Rectified Linear Unit), Tanh, or GeLU (Gaussian Error Linear Unit) to ensure that the model's expressive power is not limited by linear transformations. For example, GeLU, as a smooth nonlinear function, can adapt to the complex distribution of different speech features, making the generated speech more natural.
[0185] This embodiment processes the fused coding vectors through a style adaptation module, enabling the speech synthesis system to more precisely control speech style, improving speech personalization and adaptability. A predefined style weight matrix ensures the distinguishability of different styles. Cosine similarity calculation provides an efficient style matching strategy. The dominant style parameter selection mechanism enables more precise speech style adaptation. The application of a gating mechanism enhances the controllability of style features, and a nonlinear activation function optimizes the expressiveness of style features.
[0186] In one embodiment, after the above step S70, the method further includes:
[0187] S801, inputting the speech waveform into a pre-trained speech style encoder to generate a speech style vector;
[0188] S802, determining a cosine similarity index between the speech style vector and the prosody coding information of the reference style speech in a style vector space;
[0189] S803, minimizing the difference between the cosine similarity index and the target similarity threshold by using a contrast loss function to generate a style adaptation loss;
[0190] S804 , performing back-propagation gradient update on the weight matrix of the style adaptation module according to the style adaptation loss to optimize the weight parameters of the style adaptation module.
[0191] In this embodiment, the speech synthesis system performs style consistency optimization after speech synthesis to ensure that the final generated speech not only matches the target speaker's timbre but also matches the target style in terms of intonation, prosody, and emotional expression. By calculating the style vector of the generated speech and comparing it with the prosodic encoding information of the reference style speech, the weight parameters of the style adaptation module are adjusted. This allows the system to continuously optimize the style matching of the generated speech, thereby improving the overall naturalness and style consistency of the speech.
[0192] First, after acoustic coding generates mel-spectrogram features and feeds them into a pre-trained vocoder to generate a speech waveform, the resulting speech waveform needs to be fed into a pre-trained speech style encoder to generate a speech style vector. The speech style encoder is a deep neural network model that extracts stylistic features from the speech waveform, including the emotional state, prosodic patterns, speaking rate, and volume variations. This encoder can be trained based on a pre-trained model, for example through self-supervised learning, enabling it to extract the characteristic differences between different speech styles. For example, in emotional speech synthesis tasks, the speech style encoder can identify emotional features such as happiness, sadness, and anger, providing guidance for subsequent style optimization.
[0193] Next, the degree of similarity between the generated speech style vector and the prosodic encoding information of the reference speech style is determined to measure the degree of style match between the two. This similarity is calculated based on the angle between the two in the style vector space. A higher similarity indicates a closer match between the generated speech style and the reference speech style; a lower similarity indicates a significant stylistic deviation. For example, in a debate scenario, if the goal is to generate speech that aligns with a spirited rebuttal, and the generated speech style is closer to a calm statement, the similarity will be lower, suggesting that the system needs to further optimize the parameters of the style adaptation module.
[0194] The contrastive loss function is then used to calculate the style matching error to optimize the style adaptation process. Contrastive loss is calculated by calculating the difference between the generated speech and the reference style speech when the similarity falls below a set target similarity threshold. This difference is then used to generate an optimization signal, prompting the model to adjust parameters during subsequent training to reduce this difference. If the style of the generated speech has already reached the target style and the similarity reaches or exceeds the threshold, the loss is calculated as zero, indicating that no further adjustments are needed. Conversely, if the style of the generated speech still deviates from the target style, the loss is calculated as a positive number. A larger value indicates a greater gap between the generated speech and the target style, and the model needs to optimize the parameters of the style adaptation module more significantly. For example, in a voice dubbing task, if the target speech needs to have a dramatic and exaggerated style, while the current generated speech has a more subdued style, the system will generate a larger optimization signal, prompting the style adaptation module to adjust its parameters to bring the synthesized speech closer to the target style.
[0195] Finally, based on the style matching error, the system uses backpropagation to optimize the weight matrix of the style adaptation module, making the subsequently generated voice style more accurate. Specifically, the system calculates the degree of influence of each parameter of the style adaptation module on the style matching error, and adjusts the values of these parameters based on the calculation results, so that the style matching error in the future is gradually reduced. This optimization process is usually implemented through the gradient descent method, that is, with each iteration, the system reduces the error value relative to the adjustment amplitude of the parameter, so that the weight of the style adaptation module is gradually adjusted in a more optimal direction. For example, in a personalized voice assistant system, if the user prefers a more formal tone, and the voice initially generated by the system is more casual, the style adaptation module will continuously adjust based on user feedback to make the subsequent voice more in line with the user's preferences.
[0196] This embodiment calculates the style vector of the generated speech and compares it with the prosodic encoding information of the reference style speech, enabling the speech synthesis system to automatically optimize the parameters of the style adaptation module to improve the naturalness and consistency of the speech style. Using similarity to calculate the style matching degree ensures that the system can accurately quantify the style error of the generated speech, and using contrastive loss to optimize the style adaptation process and improve the controllability of the speech style. Finally, by optimizing the style adaptation module through backpropagation, the system can continuously adjust its own parameters to adapt to different speech style requirements, improving the overall quality and adaptability of speech synthesis.
[0197] In one embodiment, a speech generation device based on speech style adaptation is provided, and the speech generation device based on speech style adaptation corresponds to the speech generation method based on speech style adaptation in the above embodiment. Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of a speech generation device based on speech style adaptation according to the present invention. The modules include a data acquisition module 10, a text processing module 20, an acoustic feature extraction module 30, a prosody analysis module 40, a feature fusion module 50, a style adaptation module 60, and a speech generation module 70. Each functional module is described in detail below:
[0198] The data acquisition module 10 is used to acquire the target text, the target speaker's voice and the reference style voice;
[0199] A text processing module 20, configured to extract a phoneme feature sequence of the target text;
[0200] an acoustic feature extraction module 30 for extracting acoustic features of the target speaker's speech;
[0201] a prosody analysis module 40 for extracting prosody coding information of the reference style speech;
[0202] A feature fusion module 50 is used to fuse the phoneme feature sequence, acoustic features and prosodic coding information to generate a fused coding vector;
[0203] a style adaptation module 60 for processing the fused code vector through a style adaptation module to generate an acoustic code containing target style features;
[0204] The speech generation module 70 is configured to generate a mel-spectrogram feature according to the acoustic coding, and input the mel-spectrogram feature into a pre-trained vocoder to generate a speech waveform.
[0205] In one embodiment, the text processing module 20 is specifically configured to:
[0206] Segmenting the target text into a sequence of phoneme symbols using a pre-trained phoneme segmentation model;
[0207] Inputting the phoneme symbol sequence into an embedding layer to generate a phoneme embedding vector;
[0208] Performing position encoding processing on the phoneme embedding vector to generate a time series vector;
[0209] Inputting the time series vector into a bidirectional long short-term memory network to generate a phoneme context association feature;
[0210] Performing tensor concatenation of the phoneme context-related feature and the phoneme embedding vector to generate a concatenated feature vector;
[0211] Performing layer normalization processing on the concatenated feature vector to generate the phoneme feature sequence.
[0212] In one embodiment, the acoustic feature extraction module 30 is specifically configured to:
[0213] Extracting the fundamental frequency features of the target speaker's speech through a linear predictive coding module;
[0214] Determine the Mel spectrum envelope of the target speaker's speech by short-time Fourier transform;
[0215] detecting a formant frequency distribution of the target speaker's speech;
[0216] Performing tensor splicing on the fundamental frequency feature, the Mel spectrum envelope, and the formant frequency distribution to generate a spliced acoustic tensor;
[0217] The concatenated acoustic tensor is input into a pre-trained acoustic encoder to generate a joint acoustic feature matrix.
[0218] In one embodiment, the prosody analysis module 40 is specifically configured to:
[0219] detecting a fundamental frequency trajectory curve of the reference style speech;
[0220] Counting the starting boundary point and the ending boundary point of each phoneme in the reference style speech to generate a phoneme-level duration distribution;
[0221] extracting sentence-level energy fluctuation features of the reference style speech;
[0222] Performing time axis alignment processing on the fundamental frequency trajectory curve, phoneme-level duration distribution, and sentence-level energy fluctuation characteristics;
[0223] The aligned fundamental frequency trajectory curve, phoneme-level duration distribution, and sentence-level energy fluctuation features are hierarchically encoded through a pre-trained prosody encoder to generate the prosody coding information.
[0224] In one embodiment, the feature fusion module 50 is specifically configured to:
[0225] Performing time axis alignment processing on the phoneme feature sequence, acoustic features and prosodic coding information;
[0226] The aligned phoneme feature sequence and acoustic features are interactively modeled through a cross-attention mechanism to generate cross-features of phonemes and acoustics.
[0227] Performing multi-head attention weighted fusion on the phoneme and acoustic cross features and the prosody coding information to generate attention-enhanced features;
[0228] Adding the attention enhancement feature to the acoustic feature through a residual connection module to generate a residual fusion feature;
[0229] Perform layer normalization processing on the residual fusion feature to generate the fusion coding vector.
[0230] In one embodiment, the style adaptation module 60 is specifically configured to:
[0231] Processing the fused code vector through a style adaptation module to generate an acoustic code containing target style features includes:
[0232] Predefining a plurality of style weight matrices in the style adaptation module, each style weight matrix corresponds to a predefined style type;
[0233] Determining the cosine similarity between the fused encoding vector and each style weight matrix to generate a style matching score;
[0234] Select the style weight matrix with the highest style matching score as the dominant style parameter;
[0235] Performing weighted fusion processing on the dominant style parameter and the fusion coding vector through a gating mechanism to generate a weighted fusion result;
[0236] The weighted fusion result is processed by a nonlinear activation function to generate an acoustic code containing target style features.
[0237] In one embodiment, the speech generation module 70 is specifically configured to:
[0238] Inputting the speech waveform into a pre-trained speech style encoder to generate a speech style vector;
[0239] Determining a cosine similarity index between the speech style vector and the prosody coding information of the reference style speech in a style vector space;
[0240] Generating a style adaptation loss by minimizing the difference between the cosine similarity index and the target similarity threshold through a contrast loss function;
[0241] Back-propagation gradient update is performed on the weight matrix of the style adaptation module according to the style adaptation loss to optimize the weight parameters of the style adaptation module.
[0242] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a speech generation method based on speech style adaptation.
[0243] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a user-side method for speech generation based on speech style adaptation.
[0244] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0245] Obtain target text, target speaker voice, and reference style voice;
[0246] Extracting a phoneme feature sequence of the target text;
[0247] Extracting acoustic features of the target speaker's speech;
[0248] extracting prosodic coding information of the reference style speech;
[0249] Performing feature fusion on the phoneme feature sequence, acoustic features and prosody coding information to generate a fusion coding vector;
[0250] Processing the fused code vector through a style adaptation module to generate an acoustic code containing target style features;
[0251] Mel-spectrogram features are generated according to the acoustic coding, and the Mel-spectrogram features are input into a pre-trained vocoder to generate a speech waveform.
[0252] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0253] Obtain target text, target speaker voice, and reference style voice;
[0254] Extracting a phoneme feature sequence of the target text;
[0255] Extracting acoustic features of the target speaker's speech;
[0256] extracting prosodic coding information of the reference style speech;
[0257] Performing feature fusion on the phoneme feature sequence, acoustic features and prosody coding information to generate a fusion coding vector;
[0258] Processing the fused code vector through a style adaptation module to generate an acoustic code containing target style features;
[0259] Mel-spectrogram features are generated according to the acoustic coding, and the Mel-spectrogram features are input into a pre-trained vocoder to generate a speech waveform.
[0260] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0261] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0262] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0263] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A speech generation method based on speech style adaptation, characterized in that: The following steps are involved: Obtain target text, target speaker voice, and reference style voice; Extracting a phoneme feature sequence of the target text; Extracting acoustic features of the target speaker's speech; extracting prosodic coding information of the reference style speech; Performing feature fusion on the phoneme feature sequence, acoustic features and prosody coding information to generate a fusion coding vector; Processing the fused code vector through a style adaptation module to generate an acoustic code containing target style features; Mel-spectrogram features are generated according to the acoustic coding, and the Mel-spectrogram features are input into a pre-trained vocoder to generate a speech waveform.
2. The speech generation method based on speech style adaptation according to claim 1, characterized in that Extracting the phoneme feature sequence of the target text includes: Segmenting the target text into a sequence of phoneme symbols using a pre-trained phoneme segmentation model; Inputting the phoneme symbol sequence into an embedding layer to generate a phoneme embedding vector; Performing position encoding processing on the phoneme embedding vector to generate a time series vector; Inputting the time series vector into a bidirectional long short-term memory network to generate a phoneme context association feature; Performing tensor concatenation of the phoneme context-related feature and the phoneme embedding vector to generate a concatenated feature vector; Performing layer normalization processing on the concatenated feature vector to generate the phoneme feature sequence.
3. The speech generation method based on speech style adaptation according to claim 1, characterized in that Extracting acoustic features of the target speaker's speech, including: Extracting the fundamental frequency features of the target speaker's speech through a linear predictive coding module; Determine the Mel spectrum envelope of the target speaker's speech by short-time Fourier transform; detecting a formant frequency distribution of the target speaker's speech; Performing tensor splicing on the fundamental frequency feature, the Mel spectrum envelope, and the formant frequency distribution to generate a spliced acoustic tensor; The concatenated acoustic tensor is input into a pre-trained acoustic encoder to generate a joint acoustic feature matrix.
4. The method for generating speech based on speech style adaptation according to claim 1, wherein: Extracting prosody coding information of the reference style speech includes: detecting a fundamental frequency trajectory curve of the reference style speech; Counting the starting boundary point and the ending boundary point of each phoneme in the reference style speech to generate a phoneme-level duration distribution; extracting sentence-level energy fluctuation features of the reference style speech; Performing time axis alignment processing on the fundamental frequency trajectory curve, phoneme-level duration distribution, and sentence-level energy fluctuation characteristics; The aligned fundamental frequency trajectory curve, phoneme-level duration distribution, and sentence-level energy fluctuation features are hierarchically encoded through a pre-trained prosody encoder to generate the prosody coding information.
5. The method for generating speech based on speech style adaptation according to claim 1, wherein: The phoneme feature sequence, acoustic feature and prosody coding information are subjected to feature fusion to generate a fusion coding vector, including: Performing time axis alignment processing on the phoneme feature sequence, acoustic features and prosodic coding information; The aligned phoneme feature sequence and acoustic features are interactively modeled through a cross-attention mechanism to generate cross-features of phonemes and acoustics. Performing multi-head attention weighted fusion on the phoneme and acoustic cross features and the prosody coding information to generate attention-enhanced features; Adding the attention enhancement feature to the acoustic feature through a residual connection module to generate a residual fusion feature; Perform layer normalization processing on the residual fusion feature to generate the fusion coding vector.
6. The method for generating speech based on speech style adaptation according to claim 1, wherein: Processing the fused code vector through a style adaptation module to generate an acoustic code containing target style features includes: Predefining a plurality of style weight matrices in the style adaptation module, each style weight matrix corresponds to a predefined style type; Determining the cosine similarity between the fused encoding vector and each style weight matrix to generate a style matching score; Select the style weight matrix with the highest style matching score as the dominant style parameter; Performing weighted fusion processing on the dominant style parameter and the fusion coding vector through a gating mechanism to generate a weighted fusion result; The weighted fusion result is processed by a nonlinear activation function to generate an acoustic code containing target style features.
7. The method for generating speech based on speech style adaptation according to claim 1, wherein: After generating a mel spectrum feature according to the acoustic coding and inputting the mel spectrum feature into a pre-trained vocoder to generate a speech waveform, the method further includes: Inputting the speech waveform into a pre-trained speech style encoder to generate a speech style vector; Determining a cosine similarity index between the speech style vector and the prosody coding information of the reference style speech in a style vector space; Generating a style adaptation loss by minimizing the difference between the cosine similarity index and the target similarity threshold through a contrast loss function; Back-propagation gradient update is performed on the weight matrix of the style adaptation module according to the style adaptation loss to optimize the weight parameters of the style adaptation module.
8. A speech generation device based on speech style adaptation, characterized in that: The speech generation device based on speech style adaptation includes: A data acquisition module is used to acquire target text, target speaker voice, and reference style voice; A text processing module, configured to extract a phoneme feature sequence of the target text; an acoustic feature extraction module, configured to extract acoustic features of the target speaker's speech; A prosody analysis module, configured to extract prosody coding information of the reference style speech; A feature fusion module, configured to fuse the phoneme feature sequence, acoustic features, and prosodic coding information to generate a fused coding vector; a style adaptation module, configured to process the fused code vector through the style adaptation module to generate an acoustic code containing target style features; The speech generation module is used to generate Mel-spectrogram features according to the acoustic coding, and input the Mel-spectrogram features into a pre-trained vocoder to generate a speech waveform.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a speech generation program based on speech style adaptation stored in the memory and capable of running on the processor. When the speech generation program based on speech style adaptation is executed by the processor, the steps of the speech generation method based on speech style adaptation as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The storage medium stores a speech generation program based on speech style adaptation, which, when executed by a processor, implements the steps of the speech generation method based on speech style adaptation according to any one of claims 1 to 7.
Citation Information
Cited By
Controllable zero sample voice conversion method, device, equipment and medium
CN121034280A
Multi-round interaction emotion speech synthesis method and system based on context self-adaption
CN122024704A