Speech synthesis method and device, computer equipment and storage medium

By employing a dual-path fusion mechanism, combining the Mel spectrum generated from text conditions and voice prompts with the Mel spectrum generated from semantic tokens, the shortcomings of existing speech synthesis systems in content accuracy and natural expressiveness across speaker and emotion tasks are addressed, thereby enhancing the flexibility and applicability of speech synthesis.

CN121789634APending Publication Date: 2026-04-03PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing speech synthesis systems struggle to simultaneously achieve both content accuracy and natural expressiveness across speaker and emotion tasks, and single-path modeling methods cannot meet the flexible needs of different application scenarios.

Method used

A dual-path fusion mechanism is adopted, which performs weighted or adaptive fusion of the Mel spectrum generated by text conditions and voice prompts and the Mel spectrum generated based on semantic tokens. Combining text information, semantic information and sound condition information, the pre-trained speech processing model and large language model are used to extract speech features and fuse spectra.

Benefits of technology

The system achieves a balance between the accuracy of content expression and the naturalness of acoustic performance in speech synthesis, improving the naturalness and expressiveness of speech. The system can adjust semantic controllability and speech naturalness according to different application needs, and is suitable for complex scenarios such as cross-speaker, cross-emotion, and dialogue generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789634A_ABST
    Figure CN121789634A_ABST
Patent Text Reader

Abstract

The invention discloses a speech synthesis method and device, computer equipment and a storage medium, relates to the technical field of artificial intelligence, and can be applied to a medical speech generation scene or a financial speech generation scene. Through a semantic token path, the modeling capability of voice on a deep semantic structure is enhanced, and the continuity and definition of semantic expression are ensured; and acoustic details related to rhythm, timbre and emotion are reserved through the text condition and the voice prompt path, so that the naturalness and expressive force of the voice can be improved. And an adjustable weight or a self-adaptive mechanism is introduced in the fusion process, so that the system can realize smooth adjustment between semantic controllability and voice naturalness according to different application requirements, thereby improving the flexibility and the application range of the whole system. According to the method, the problems of voice distortion, emotion weakening or single expression caused by information loss of a single path are effectively reduced, and the method has more stable and superior synthesis performance in complex scenes such as cross-speaker, cross-emotion and dialogue generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a speech synthesis method, apparatus, computer device, and storage medium. Background Technology

[0002] Existing text-to-speech (TTS) systems can be broadly categorized into two approaches: one is based on text plus voice prompts, where text input is combined with reference timbre, emotion, or prosodic information to generate a Mel spectrum, which is then processed by a vocoder to obtain the waveform. This approach can better guarantee semantic accuracy and timbre consistency, but it falls short in terms of prosodic naturalness and emotional expressiveness. The other approach introduces semantic tokens generated by a large-scale language model (LLM), which are then used to generate the Mel spectrum through a generative model. This approach can better capture pauses, stresses, and rhythmic information in natural speech, thereby improving the naturalness of the speech and the conversational style, but it often suffers from deficiencies in semantic clarity and controllability.

[0003] Currently, most research methods tend to choose one path in practical systems, making it difficult to balance content accuracy and natural expressiveness. Especially in tasks involving multiple speakers and emotions, single-path modeling approaches struggle to meet the flexible needs of users in different application scenarios.

[0004] Therefore, how to balance content accuracy and natural expressiveness in speech synthesis has become an urgent problem to be solved in the field of speech synthesis. Summary of the Invention

[0005] The purpose of this application is to provide a speech synthesis method, apparatus, computer device, and storage medium to simultaneously ensure content accuracy and natural expressiveness during speech synthesis, thereby improving the quality of speech synthesis.

[0006] To address the aforementioned technical problems, this application provides a speech synthesis method, employing the following technical solution: A speech synthesis method, comprising: Obtain the input text sequence; The text feature vector corresponding to the input text sequence is obtained through a preset text encoder; The input sound conditions are obtained, and the corresponding sound feature vectors are obtained through a preset acoustic encoder. The first speech feature vector is obtained by concatenating the feature vector of the text and the sound feature vector. The first speech feature vector is input into the pre-trained first speech processing model to obtain the first Mel spectrum; Semantic token sequences are obtained by performing semantic recognition on the input text sequence using a pre-trained large language model. The semantic token sequence is input into a pre-trained second speech processing model to obtain the second Mel spectrum; The first Mel spectrum and the second Mel spectrum are spectrally fused to obtain the speech fused spectrum; The speech fusion spectrum is input into a preset vocoder to obtain synthesized speech.

[0007] To address the aforementioned technical problems, this application also provides a speech synthesis device, which employs the following technical solution: A speech synthesis method, comprising: The text acquisition module is used to acquire the input text sequence; The text encoding module is used to obtain the text feature vector corresponding to the input text sequence through a preset text encoder; The acoustic coding module is used to acquire input sound conditions and obtain the sound feature vector corresponding to the input sound conditions through a preset acoustic encoder. The audio-text splicing module is used to splice the text feature vector and the audio feature vector to obtain the first audio feature vector; The first spectrum generation module is used to input the first speech feature vector into the pre-trained first speech processing model to obtain the first Mel spectrum. The semantic recognition module is used to perform semantic recognition on the input text sequence using a pre-trained large language model to obtain a semantic token sequence. The second spectrum generation module is used to input the semantic token sequence into the pre-trained second speech processing model to obtain the second Mel spectrum; The spectrum fusion module is used to fuse the first Mel spectrum and the second Mel spectrum to obtain the speech fusion spectrum. The vocoder conversion module is used to input the speech fusion spectrum into a preset vocoder to obtain synthesized speech.

[0008] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution: A computer device includes a memory and a processor, the memory storing computer-readable instructions, the processor executing the computer-readable instructions to implement the steps of the speech synthesis method as described in any of the preceding claims.

[0009] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below: A computer-readable storage medium storing computer-readable instructions that, when executed by a processor, implement the steps of the speech synthesis method as described in any one of the preceding descriptions.

[0010] Compared with the prior art, the embodiments of this application have the following main advantages: This application discloses a speech synthesis method, apparatus, computer device, and storage medium, relating to the field of artificial intelligence technology, and applicable to medical or financial speech generation scenarios. This application introduces a dual-path fusion mechanism, simultaneously utilizing Mel spectra generated based on text conditions and voice prompts, as well as Mel spectra generated based on semantic tokens, and weighting or adaptively fusing the two types of spectra. This collaboratively models textual, semantic, and sound conditional information within the same acoustic representation, avoiding the insufficient information expression problem caused by relying solely on text or semantic modeling in single-path speech synthesis. This achieves a balance between the accuracy of content expression and the naturalness of acoustic performance in synthesized speech. On one hand, the semantic token path enhances the ability of speech to model deep semantic structures, ensuring the coherence and clarity of semantic expression; on the other hand, the text conditions and voice prompt path preserves prosody, timbre, and emotion-related acoustic details, which helps improve the naturalness and expressiveness of the speech. Simultaneously, the introduction of adjustable weights or adaptive mechanisms during the fusion process allows the system to smoothly adjust between semantic controllability and speech naturalness according to different application requirements, thereby improving the overall system's flexibility and applicability. This application effectively reduces the problems of speech distortion, weakened emotion, or monotonous expression caused by the lack of information in a single path, and has more stable and superior synthesis performance in complex scenarios such as cross-speaker, cross-emotion, and dialogue generation. Attached Figure Description

[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 An exemplary system architecture diagram is shown, in which this application can be applied; Figure 2 A flowchart of one embodiment of the speech synthesis method according to this application is shown; Figure 3 A schematic diagram of the structure of one embodiment of the speech synthesis apparatus according to this application is shown; Figure 4 A schematic diagram of the structure of one embodiment of a computer device according to this application is shown. Detailed Implementation

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0016] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0017] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0019] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0020] It should be noted that the speech synthesis method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the speech synthesis device is generally set in the server / terminal device.

[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative; the system can have any number of terminal devices, networks, and servers depending on implementation needs.

[0022] Continue to refer to Figure 2 A flowchart of an embodiment of the speech synthesis method according to this application is shown. The speech synthesis method includes the following steps: S201, Obtain the input text sequence; Specifically, the system receives the text content to be synthesized from the user input terminal, the upper-layer business system, or a preset text storage module, and performs structured processing on the text content to form an input text sequence. The input text sequence can be a character-level sequence, a word-level sequence, or a sub-word-level sequence, the specific form of which is determined by the input requirements of the text encoder. During the acquisition process, basic preprocessing operations can be performed on the original text, including but not limited to text cleaning, illegal character filtering, punctuation standardization, and encoding format unification, to ensure the semantic and format consistency of the text sequence. Furthermore, the text can be segmented into sentences or paragraphs as needed to ensure that the text sequence meets the model's constraints on input length or context window, thereby providing a standardized and parsable input form for text feature extraction.

[0023] The input text can be medical or insurance / finance related text, such as medical diagnosis reports, medical records, drug instructions, and other professional medical texts, or insurance policy descriptions, financial product introductions, claims guidelines, and other insurance / finance texts. Input text can also be general texts such as everyday conversations, news articles, and literary works. For input text sequences containing technical terms, specific formats, or special symbols, the preprocessing stage can also incorporate domain-specific dictionaries for terminology recognition and annotation, or perform pre-defined rule-based conversion processing on special symbols to ensure the text encoder can accurately parse the semantic information of the text.

[0024] S202, obtain the text feature vector corresponding to the input text sequence through a preset text encoder; Specifically, a text encoder maps an input text sequence into a continuous vector space to obtain a high-dimensional feature representation that characterizes the text content. The text encoder can employ a structure combining embedding layers and sequence modeling networks. For example, it maps characters, words, or subwords in the text sequence into dense vector representations and further models contextual information using recurrent neural networks, convolutional neural networks, or self-attention networks. During processing, the text encoder comprehensively considers the sequential relationships and contextual dependencies between elements in the text sequence, thereby outputting a text feature vector containing semantic and structural information. This text feature vector typically exists in the form of a time-step sequence.

[0025] A text encoder employing a pre-trained bidirectional Transformer model can effectively extract deep semantic features from an input text sequence by capturing long-distance semantic dependencies through a multi-layered self-attention mechanism. In its implementation, the pre-processed input text sequence is converted into token embeddings acceptable to the model, including word embeddings, positional embeddings, and segment embeddings, which are then fed into the stacked network layers of the Transformer encoder. Each layer contains a multi-head self-attention sublayer and a feedforward neural network sublayer. The multi-head self-attention sublayer performs weighted aggregation of information from different positions in the text sequence by computing multiple different attention weight distributions in parallel, comprehensively capturing the semantic relationships between words. The feedforward neural network sublayer performs non-linear transformations and dimensionality mappings on the aggregated features, enhancing the model's feature representation capabilities. After iterative processing by the multi-layered Transformer encoder, the final output is a sequence of text feature vectors corresponding to the length of the input text sequence, with each vector at each time step containing rich semantic information of the corresponding text segment.

[0026] S203, obtain the input sound conditions, and obtain the sound feature vector corresponding to the input sound conditions through a preset acoustic encoder; Specifically, the input sound conditions may include a reference speech signal, speaker timbre characteristics, or other acoustic information related to the generation of the target speech. Upon receiving the input sound conditions, the system first performs preprocessing operations on the original speech signal, such as resampling, endpoint detection, amplitude normalization, or denoising, to ensure the stability of the input signal. Subsequently, the processed speech signal is input into a preset acoustic encoder, which models the temporal structure and spectral characteristics of the speech, extracting sound feature vectors that reflect the speaker's timbre, vocal characteristics, or prosodic features. These sound feature vectors are represented in continuous numerical form.

[0027] Acoustic encoders can employ a hybrid architecture combining convolutional neural networks (CNNs) and recurrent neural networks (RNNs). Convolutional layers extract local spectral features from the speech signal, while recurrent layers (such as LSTM or GRU) capture the temporal dynamics of the speech. In the specific processing, the preprocessed speech signal is first converted into a spectrogram representation (such as a Mel spectrogram or speech spectrogram), which is then used as input to the acoustic encoder. Through multiple convolutional operations, spectral features are progressively extracted from low to high levels. For example, 3x3 convolutional kernels capture local texture information of the spectrum, and pooling layers reduce feature dimensionality and enhance translation invariance. Next, the feature sequence output from the convolutional layers is input to a bidirectional recurrent layer. The forward and backward temporal dependencies are used to model the prosodic rhythm and timbre variation trends of the speech, ultimately outputting a sound feature vector that comprehensively represents the core acoustic properties of the input sound. For the extraction of speaker timbre features, the acoustic encoder can also integrate a speaker recognition module. By using a pre-trained speaker embedding model (such as x-vector or d-vector), the reference speech is mapped into a fixed-dimensional speaker identity vector, and then fused with the spectral temporal features to form a sound feature vector that combines timbre individuality and speech dynamic characteristics.

[0028] S204, The feature vector of this paper and the sound feature vector are concatenated to obtain the first speech feature vector; Specifically, after obtaining the text feature vector and the sound feature vector, they need to be fused at the feature level to form a joint speech feature representation. This step typically involves concatenation operations in the feature dimension or the time dimension, enabling text information and sound conditional information to be jointly represented in the same feature space. Before concatenation, the text feature vector and the sound feature vector can be aligned or repeatedly expanded as needed to ensure consistency in the number of time steps or feature dimensions. The first speech feature vector obtained through concatenation simultaneously contains both text content information and sound conditional information.

[0029] The concatenation of text and sound feature vectors in this paper is implemented in two ways: First, if the text and sound feature vectors are sequence features of the same time length, time-dimensional alignment concatenation is used. That is, the text feature vector at each time step is concatenated with the corresponding sound feature vector at that time step to form a joint feature sequence with superimposed dimensions. Second, if the sound feature vector is a globally fixed-dimensional vector (such as a speaker embedding vector), it is copied and expanded in the time dimension to match the time step number of the text feature vector, and then concatenated point-by-point in the time step dimension. For example, when the text feature vector is a sequence of shapes (T, D_text) and the sound feature vector is a global vector of shape (D_sound), the sound feature vector is first expanded to (T, D_sound), and then concatenated to form the first speech feature vector of shape (T, D_text + D_sound). During the concatenation process, feature standardization (such as LayerNorm) can be used to eliminate differences in the numerical distribution of different modal features.

[0030] S205, input the first speech feature vector into the pre-trained first speech processing model to obtain the first Mel spectrum; Specifically, the first speech processing model maps the fused first speech feature vector to the corresponding acoustic representation. This model can employ a deep neural network-based acoustic modeling structure, such as an encoder-decoder model or an autoregressive generative model. During processing, the model predicts the time-frequency distribution of the speech at the Mel frequency scale based on the input first speech feature vector, thereby generating a first Mel spectrum. The first Mel spectrum is represented in two-dimensional matrix form, containing information in both the time and frequency dimensions, and is used to characterize the energy distribution features of the speech signal.

[0031] S206, semantic recognition is performed on the input text sequence using a pre-trained large language model to obtain a semantic token sequence; Specifically, the large language model is used for deep semantic modeling and parsing of input text sequences to obtain implicit semantic structural information within the text. After receiving the text sequence, the large language model jointly models the text context through a multi-layered self-attention mechanism, thereby generating intermediate representations reflecting semantic relationships and further outputting a corresponding sequence of semantic tokens. These semantic tokens can be discrete symbols or vectorized representations, used to characterize semantic units in the text and their relationships. In this way, the system can introduce higher-level semantic information beyond the text content, providing semantic constraints for speech generation.

[0032] S207, Input the semantic token sequence into the pre-trained second speech processing model to obtain the second Mel spectrum; Specifically, the second speech processing model generates corresponding acoustic representations based on a sequence of semantic tokens. This model is designed to focus on the mapping relationship between semantic information and acoustic features. By decoding or transforming the sequence of semantic tokens, it predicts the speech representation in the Mel frequency domain. The model can be implemented based on a sequence-to-sequence structure or a non-autoregressive generation structure. During the generation process, it comprehensively considers the temporal relationships and contextual information between semantic tokens, thereby outputting a second Mel spectrum. The second Mel spectrum primarily reflects the influence of semantics on speech generation and complements the first Mel spectrum.

[0033] The first and second speech processing models can be acoustic generation models built on different technical approaches. For example, the first speech processing model can adopt a Transformer-based non-autoregressive generation model (such as the FastSpeech series models), which achieves rapid prediction of Mel spectra through a parallel generation mechanism and can efficiently utilize the prosody and timbre information after fusing text feature vectors and sound feature vectors. The second speech processing model can adopt a generation architecture based on a diffusion probability model (such as the Diffusion Model), which generates Mel spectra that conform to semantic token sequence constraints from random noise through an iterative denoising process, possessing stronger semantic-acoustic mapping capabilities and detail generation capabilities. In addition, the first speech processing model can also use an encoder-decoder structure based on recurrent neural networks (such as the Tacotron series models), which aligns text and acoustic features through an attention mechanism to generate a first Mel spectra with natural prosodic fluctuations. The second speech processing model can also adopt a pre-trained speech-semantic multimodal model, which learns text semantic representation and acoustic feature generation simultaneously during the training phase, and can directly convert semantic token sequences into second Mel spectra rich in semantic emotional color. The two models capture the acoustic details required for speech generation from different perspectives.

[0034] S208, perform spectrum fusion on the first Mel spectrum and the second Mel spectrum to obtain the speech fusion spectrum; Specifically, after obtaining the first and second Mel spectra, a unified spectral fusion process is required to generate a fused spectral representation for final speech synthesis. The spectral fusion process can be performed in the time, frequency, or combined time-frequency dimensions, integrating spectral information from different models through weighting, concatenation, or other fusion methods. Before fusion, the two Mel spectra can be time-aligned or scale-normalized to ensure their consistency in the representation space. Through spectral fusion, the system can integrate multi-source spectral information in a single acoustic representation, providing complete input data for the vocoder.

[0035] In specific embodiments of this application, the fusion of the Mel spectrum can employ conventional weighted fusion, frequency band weighted fusion, or attention-weighted fusion schemes, specifically including the following implementation methods: Conventional weighted fusion assigns fixed or dynamically adjusted weight coefficients to the first and second Mel spectra, and then linearly sums the spectral energy values ​​at corresponding positions. For example, if the weight of the first Mel spectrum is set to 0.6 and the weight of the second Mel spectrum is set to 0.4, the calculation formula is: fused spectrum = 0.6 × first Mel spectrum + 0.4 × second Mel spectrum. The weight coefficients can be optimized offline or adaptively adjusted online according to the synthesis effect of the model on the validation set.

[0036] Frequency band weighted fusion divides the frequency axis of the Mel spectrum into multiple sub-bands (such as low frequency band 20-500Hz, mid frequency band 500-2000Hz, and high frequency band 2000-8000Hz). Independent weight coefficients are set for different sub-bands. For example, the mid frequency band, which affects speech clarity, is given a higher weight (such as 0.7), the low frequency band, which is related to timbre performance, is given a medium weight (such as 0.5), and the high frequency band, which is related to detail richness, is given a lower weight (such as 0.3). Through frequency band differentiated weighting, the spectrum information is refined and integrated.

[0037] Attention-weighted fusion introduces dynamic weight allocation based on a self-attention mechanism. The first and second Mel spectra are used as inputs to the attention module. The similarity matrix between the two spectra in the time-frequency domain is calculated. The attention weight distribution in the time step and frequency point dimensions is generated by the Softmax function, so that the fusion process can adaptively focus on the more reliable information regions in the two spectra (such as the prosodic accurate time period in the first Mel spectrum and the semantic sentiment matching frequency band in the second Mel spectrum). Finally, the fused spectrum is obtained by weighted summation. This method can adjust the fusion strategy according to the dynamic changes of the input text and sound conditions, and improve the robustness of spectrum fusion in complex scenarios.

[0038] S209: Input the speech fusion spectrum into a preset vocoder to obtain synthesized speech.

[0039] Specifically, a vocoder is used to convert the input speech fusion spectrum into a time-domain speech waveform signal. The vocoder can employ a neural network-based generative model or other parametric synthesis model. After receiving the speech fusion spectrum, it reconstructs the speech signal frame by frame based on the time-frequency information contained in the spectrum. During the generation process, the vocoder predicts the sampling point values ​​of the speech waveform based on spectral characteristics, thereby outputting a continuous audio signal. The final synthesized speech can be output in a digital audio format for playback, storage, or further processing.

[0040] This application introduces a "dual-path fusion" mechanism, which weights or adaptively fuses the Mel spectrum generated by combining text conditions with voice prompts with the Mel spectrum based on semantic tokens, balancing semantic clarity and speech naturalness. Compared with existing single-path methods, this scheme ensures content accuracy in semantic expression while significantly improving naturalness and expressiveness at the prosodic and emotional levels. Furthermore, users can dynamically adjust the weights of the two paths according to the application scenario, achieving a smooth transition from "high controllability" to "high naturalness," offering greater flexibility and scalability. This application effectively avoids speech distortion or emotional loss caused by insufficient single-path processing, thus demonstrating superior synthesis effects in scenarios involving cross-speaker, cross-emotion, and dialogue generation.

[0041] In one specific embodiment of this application, taking the voice generation of a claims guide as an example, when a user inputs the claims guide text "You can submit materials through the 'claims application' entry on the APP homepage. You need to upload the front and back of your ID card, a photo of your bank card, and a list of medical expenses," and selects "standard Mandarin voice of customer service personnel" as the reference voice condition, the system executes the following process: First, the Transformer encoder performs multi-layer processing on the text sequence. The multi-head self-attention sublayer captures the dependencies between "APP homepage" and "claims application entry," and between "submit materials" and words such as "ID card," "bank card," and "medical expense list." The feedforward neural network sublayer strengthens these dependencies. Semantic association features are used to output a sequence of text feature vectors containing the text's logical structure. After preprocessing the input customer service reference speech using an acoustic encoder, the local spectral envelope of the Mel spectrum is extracted through a convolutional layer. A bidirectional LSTM layer models the intonation fluctuations of the speech (such as the rising intonation of "you can" and the emphasis of "needs to be uploaded"). Combined with a pre-trained x-vector speaker embedding model, a sound feature vector containing the customer service representative's gentle tone and standard speaking speed is generated. The text feature vector and the sound feature vector are concatenated at time steps (the text sequence length is 15, and the sound feature vector is expanded to 15×D_sound dimension), and then processed by LayerNorm to form a fused feature vector. The FastSpeech2 model (the first speech processing model) quickly generates a first Mel spectrum with the rhythm of customer service speech based on the fused features. Its spectrogram shows energy peaks in keyword segments such as "front and back of ID card" and "bank card photo." The large language model performs semantic parsing on the text, identifying "operation guidance" semantics and generating a semantic token sequence containing semantic tags such as "instruction," "material type," and "operation entry." The diffusion model (the second speech processing model) generates a second Mel spectrum based on the semantic token sequence. This spectrum exhibits a reminder-like prosodic feature at the word "need" and a natural pause mark at the word "and." Frequency-band weighted fusion is employed. The system assigns a weight of 0.5 to the first Mel spectrum and 0.5 to the second Mel spectrum for the mid-frequency band (500-2000Hz) that affects clarity, a weight of 0.7 to the first Mel spectrum for the timbre characteristics of the low-frequency band (20-500Hz), and a weight of 0.6 to the second Mel spectrum for the detail characteristics of the high-frequency band (2000-8000Hz), generating a fused spectrum. Finally, a vocoder (such as HiFi-GAN) converts the fused spectrum into a speech waveform. The synthesized speech maintains the professional tone of the customer service personnel and, through semantic guidance and prosodic adjustment, makes the expression "needs to upload... and..." clear and has a natural guiding tone, improving the comprehensibility of the claims guidance and the user experience.

[0042] Furthermore, the step of spectral fusion of the first Mel spectrum and the second Mel spectrum to obtain the speech fusion spectrum specifically includes: Alignment processing is performed on the first Mel spectrum and the second Mel spectrum in terms of time and frequency dimensions; The first weighting coefficient corresponding to the first Mel spectrum and the second weighting coefficient corresponding to the second Mel spectrum are determined according to the preset fusion strategy, wherein the sum of the first weighting coefficient and the second weighting coefficient is 1. Based on the first weighting coefficient and the second weighting coefficient, the first Mel spectrum and the second Mel spectrum are weighted respectively to obtain the weighted first spectrum and the weighted second spectrum. The weighted first spectrum and the weighted second spectrum are superimposed to obtain the speech fusion spectrum.

[0043] In this embodiment, the fusion of the Mel spectrum is achieved using a conventional weighted fusion method. Before fusing the first and second Mel spectra, they are first aligned in both the time and frequency dimensions to ensure consistency in parameters such as frame number, time step, and Mel channel number, avoiding additional distortion caused by scale mismatch. Then, weight coefficients are determined for the first and second Mel spectra according to a preset fusion strategy. This strategy can be configured based on system preset rules or application requirements, ensuring that the sum of the first and second weight coefficients is 1 to guarantee the overall energy stability of the fused spectrum. Based on this, the first and second Mel spectra are weighted using the determined weight coefficients, allowing the two spectral information to participate in the fusion with different contribution ratios. Finally, the weighted first and second spectra are element-wise superimposed at their corresponding time-frequency positions to form a unified speech fusion spectrum representation. This method effectively integrates acoustic information from different generation paths while maintaining the stability of the fusion process.

[0044] By employing the above steps and integrating the two Mel spectra using a conventional weighted fusion method, the speech synthesis process becomes structurally more stable and reliable. This ensures the consistency of spectral energy while also considering the fusion of multi-source acoustic information, which helps improve the overall quality and consistency of synthesized speech.

[0045] Furthermore, the step of spectral fusion of the first Mel spectrum and the second Mel spectrum to obtain the speech fusion spectrum specifically includes: Alignment processing is performed on the first Mel spectrum and the second Mel spectrum in terms of time and frequency dimensions; The Mel spectrum is divided into low-frequency and high-frequency bands according to the preset frequency division dimensions; For the low-frequency band, a weighted processing method based on the first Mel spectrum is used to generate the low-frequency fused spectrum; For the high-frequency band, a weighted processing method based on the second Mel spectrum is used to generate the high-frequency fused spectrum; The low-frequency fusion spectrum is spliced ​​with the high-frequency fusion spectrum to obtain the speech fusion spectrum.

[0046] In this embodiment, the fusion of the Mel spectrum is achieved using a frequency-band weighted fusion method. Specifically, before fusing the first and second Mel spectra, the two Mel spectra are first aligned in both the time and frequency dimensions to ensure consistency in terms of time frame number, frame shift parameters, and the number of Mel channels, thus providing a unified time-frequency reference for subsequent frequency-band processing. Subsequently, the Mel spectrum is divided into frequency bands according to a preset frequency division dimension, dividing the complete Mel spectrum into low-frequency and high-frequency bands. The frequency division dimension can be set based on the Mel channel index or the corresponding actual frequency range. After the frequency band division is completed, for the low-frequency band, a weighted processing method based on the first Mel spectrum is used to generate a low-frequency fused spectrum, so that the low-frequency region mainly retains the acoustic information generated by the first path; for the high-frequency band, a weighted processing method based on the second Mel spectrum is used to generate a high-frequency fused spectrum, so that the high-frequency region mainly introduces the acoustic information provided by the second path. Finally, the obtained low-frequency fused spectrum and high-frequency fused spectrum are concatenated in the frequency dimension to form a complete speech fusion spectrum representation. This method allows for the introduction of spectral features from different sources into different frequency bands, enabling differentiated fusion processing across frequency bands.

[0047] For example, if the total number of Mel channels in the Mel spectrum is set to 80, channels 0-20 (corresponding to actual frequencies of approximately 0-500Hz) can be divided into the low-frequency band, channels 21-60 (corresponding to actual frequencies of approximately 500-2000Hz) into the mid-frequency band, and channels 61-80 (corresponding to actual frequencies of approximately 2000-8000Hz) into the high-frequency band, based on the human ear's sensitivity to low frequencies. In low-frequency processing, the weighting coefficient of the first Mel spectrum is set to 0.7, and the weighting coefficient of the second Mel spectrum is set to 0.3, i.e., low-frequency fusion spectrum = 0.7 × first Mel spectrum low-frequency band + 0.3 × second Mel spectrum low-frequency band, thus enhancing the timbre stability of the first model in the low-frequency region. In high-frequency processing, the weighting coefficient of the first Mel spectrum is set to 0.3, and the weighting coefficient of the second Mel spectrum is set to 0.6, i.e., high-frequency fusion spectrum = 0.3 × first Mel spectrum high-frequency band + 0.6 × second Mel spectrum high-frequency band, emphasizing the detail performance of the second model in the high-frequency region. A balanced weighting strategy (e.g., 0.5 for each frequency band) can be used in the mid-frequency band to balance the semantic-acoustic mapping accuracy of both. After completing the frequency band weighting, the low-frequency fusion spectrum, mid-frequency fusion spectrum, and high-frequency fusion spectrum are concatenated in the original channel order along the frequency dimension to form a speech fusion spectrum covering the entire frequency band. By using this frequency band differentiated weighting, the system can flexibly allocate model contributions based on the acoustic characteristics of different frequency bands (such as low frequency dominating timbre, high frequency dominating detail, and mid frequency dominating clarity), so that the fused spectrum can enhance high frequency details and mid frequency semantic accuracy while preserving the basic timbre. It is especially suitable for scenarios that require both timbre stability and emotional detail (such as audio reading, intelligent customer service dialogue, etc.).

[0048] In a specific embodiment of this application, a frequency band weighted fusion expression is as follows:

[0049] In the formula, For the spectrum of speech fusion, For frequency dimension, The frequency dimension variation parameter is a weighting function that changes nonlinearly with frequency f, used to dynamically adjust the first Mel spectrum. Second Mel spectrum (denoted as The contribution ratio at different frequency points The time dimension. The functional form can be pre-designed according to the acoustic characteristics of the target speech or learned through model training. For example, to enhance the stability of the low-frequency band, In the low-frequency region (e.g., f < 500Hz), the value can be 0.7-0.9, and as the frequency increases to the mid-frequency region (e.g., 500Hz ≤ f ≤ 2000Hz). The weight can be smoothly transitioned to 0.4-0.6, and further reduced to 0.2-0.4 in the high-frequency region (e.g., f>2000Hz) to highlight the advantages of the second Mel spectrum in high-frequency details. This continuous frequency-dependent weighting function, compared to the discrete weighting allocation of fixed frequency bands, can achieve a more refined smooth transition of weights between frequency bands, avoiding possible abrupt changes in spectral energy or discontinuities in features at frequency band boundaries, thereby generating a more natural and coherent fused spectrum.

[0050] Through the above steps, the Mel spectrum of different frequency regions is processed in a targeted manner by using a frequency band weighted fusion method, so that the fused spectrum is more coordinated in overall structure. This is beneficial to enhance the expression of high-frequency details while maintaining low-frequency stability, thereby improving the naturalness and clarity of synthesized speech.

[0051] Furthermore, the step of dividing the Mel spectrum into low-frequency and high-frequency bands according to a preset frequency division dimension specifically includes: Obtain the total number of Mel frequency channels in the first Mel spectrum and the second Mel spectrum; The Mel channel range corresponding to the low frequency band and the Mel channel range corresponding to the high frequency band are determined according to the preset frequency division parameters. Based on the Mel channel range, the corresponding low-frequency band spectral components and high-frequency band spectral components are extracted from the first Mel spectrum and the second Mel spectrum, respectively.

[0052] In this embodiment, the frequency band division of the Mel spectrum is implemented based on Mel frequency channels. Specifically, before performing the frequency band division operation, the total number of Mel frequency channels used in the first and second Mel spectra is first obtained to clarify the discrete division structure of the spectrum in the frequency dimension. Based on this, the Mel channel ranges corresponding to the low-frequency and high-frequency bands are determined according to preset frequency division parameters. The frequency division parameters can be set in a fixed ratio manner or flexibly configured according to actual application needs to adapt to different spectral resolutions. Subsequently, based on the determined Mel channel ranges, the corresponding low-frequency band spectral components and high-frequency band spectral components are extracted from the first and second Mel spectra, respectively, so that the spectral data of different frequency bands are structurally independent. Through the above method, while maintaining the overall structural consistency of the Mel spectrum, it is possible to achieve fine division and independent processing of different frequency regions.

[0053] Through the above steps, the frequency band division method based on Mel channel enables the spectral information of different frequency regions to be processed independently, which is beneficial to improving the accuracy and stability of the subsequent frequency band fusion process.

[0054] Furthermore, the step of spectral fusion of the first Mel spectrum and the second Mel spectrum to obtain the speech fusion spectrum specifically includes: Alignment processing is performed on the first Mel spectrum and the second Mel spectrum in terms of time and frequency dimensions; The first and second Mel spectra are used as inputs to a preset fusion network. The fusion network extracts spectral features and generates corresponding attention weights. The first Mel spectrum and the second Mel spectrum are dynamically weighted based on attention weights to obtain the first weighted spectral features and the second weighted spectral features. The first weighted spectral feature and the second weighted spectral feature are fused to obtain the speech fusion spectrum.

[0055] In this embodiment, the fusion of Mel spectra is achieved using a dynamic weighted fusion method based on an attention mechanism. Before performing the fusion operation, the first and second Mel spectra are aligned in both the time and frequency dimensions to ensure consistency in the number of time frames, frequency channels, and corresponding time-frequency positions, thereby avoiding structural mismatches that could affect the fusion effect. Subsequently, the aligned first and second Mel spectra are simultaneously input into a preset fusion network, which jointly models the two spectra. The fusion network extracts high-dimensional features from the spectral data using a feature extraction module, and then uses an attention generation module to calculate attention weights corresponding to different spectral features, characterizing the relative importance of the two spectra in different time frames or frequency channels. After obtaining the attention weights, the first and second Mel spectra are dynamically weighted based on these weights to generate a first weighted spectral feature and a second weighted spectral feature. Finally, a fusion operation is performed on the two weighted spectral features in the time-frequency dimension to obtain a unified speech fusion spectral representation. This dynamic weighting method allows the spectral fusion process to adaptively adjust according to changes in the input features.

[0056] In one specific embodiment of this application, the loss function of the fusion network is integrated. The expression is as follows:

[0057] In the formula, The Mel spectrum of the actual input speech. , , To compensate for the loss weights, CTC loss or attention alignment is used during training to ensure that the outputs of the two paths are consistent in the time dimension, thus facilitating spectral fusion. For example, the fusion network can adopt a Transformer encoder structure, where a multi-head self-attention mechanism is used to capture the correlation features between the first and second Mel spectra in the time and frequency dimensions, generating an attention weight matrix by calculating the similarity scores of different spectral regions. Assuming that the first Mel spectrum contains richer timbre fundamental frequency information and the second Mel spectrum has more accurate semantic prosodic features, the fusion network will assign a higher attention weight (e.g., 0.8) to the first Mel spectrum in voiced segments (such as vowels) to enhance the stability of low-frequency harmonics; while in unvoiced segments (such as consonants "s" and "sh"), it will increase the weight of the second Mel spectrum (e.g., 0.7) to highlight the details of high-frequency fricatives. Through this dynamic adjustment, the attention weights can adapt to the local features of the input spectrum, making the fusion process more in line with the acoustic characteristics of speech. Finally, the weighted first and second spectral features are nonlinearly mapped through a fully connected layer to output the final fused speech spectrum. Through the above steps, the dynamic weighted fusion based on the attention mechanism can fully explore the complementarity of the two Mel spectra, enabling the fused spectrum to adaptively retain key acoustic features in different speech segments, thereby further improving the expressiveness and scene adaptability of the synthesized speech.

[0058] Through the above steps, by introducing a dynamic weighted fusion method based on an attention mechanism, the spectrum fusion process becomes adaptive, which is beneficial for achieving a more balanced and stable speech synthesis performance under different speech content and conditions.

[0059] Furthermore, the steps of using the first Mel spectrum and the second Mel spectrum as input to a preset fusion network, extracting spectral features through the fusion network, and generating corresponding attention weights specifically include: The first Mel spectrum and the second Mel spectrum are subjected to feature normalization processing to obtain standardized first spectral features and second spectral features. The standardized first and second spectral features are spliced ​​together to form a joint spectral feature; The joint spectral features are input into the feature extraction module in the fusion network to extract high-dimensional spectral representation features; Based on the high-dimensional spectral representation features, the attention layers in the network are fused, and the corresponding attention weight parameters are calculated. The attention weight parameters are normalized to obtain the attention weights.

[0060] In this embodiment, the fusion network performs systematic feature processing and modeling on the input first and second Mel spectra before generating attention weights. Specifically, firstly, feature normalization is performed on the two Mel spectra to eliminate differences in amplitude range and distribution, resulting in standardized first and second spectral features, thereby improving the stability of subsequent joint modeling. After normalization, the standardized first and second spectral features are concatenated along the feature dimension to form a joint spectral feature representation containing dual-path spectral information. Subsequently, the joint spectral features are input to the feature extraction module in the fusion network, which models the structural relationship of the spectrum in the time and frequency dimensions to extract high-dimensional spectral representation features. After obtaining the high-dimensional spectral representation features, the attention layer in the fusion network maps and analyzes the high-dimensional features to calculate attention weight parameters corresponding to different spectral features. Finally, the generated attention weight parameters are normalized to meet preset constraints, thereby obtaining attention weights that can be used for subsequent dynamic weighted fusion of the spectrum. Through the above steps, unified modeling and adaptive weight generation of multi-source Mel spectral features are achieved.

[0061] Through the above steps, attention weights are automatically generated and normalized by the fusion network, enabling different spectral features to be adaptively adjusted during the fusion process, which helps to improve the stability and consistency of spectral fusion.

[0062] Furthermore, based on the high-dimensional spectral representation features, the steps of fusing attention layers in the network and calculating the corresponding attention weight parameters specifically include: The high-dimensional spectral representation features are input into the linear mapping region in the attention layer to obtain the intermediate attention feature representation; Nonlinear activation processing is applied to the intermediate attention feature representation; Based on the activated intermediate attention feature representation, attention weight parameters corresponding to the first Mel spectrum and the second Mel spectrum are generated through the weight calculation region.

[0063] In this embodiment, the generation of attention weight parameters is accomplished based on the attention layer set in the fusion network. Specifically, the high-dimensional spectral representation features obtained in the previous steps are input into the linear mapping region of the attention layer. The high-dimensional features are then adjusted and reorganized through linear transformation to obtain an intermediate attention feature representation. This intermediate attention feature representation is used to characterize the correlation information of different spectral features at the current time-frequency position. Subsequently, a nonlinear activation process is applied to the intermediate attention feature representation to enhance the expressive power of the features and introduce nonlinear modeling capabilities, thereby enabling the attention layer to capture more complex feature relationships. After completing the nonlinear activation, based on the activated intermediate attention feature representation, attention weight parameters corresponding to the first Mel spectrum and the second Mel spectrum are generated through the weight calculation region in the attention layer. The weight calculation region can output corresponding weight values ​​according to the distribution of the joint spectral features in different dimensions.

[0064] Through the above steps, the high-dimensional spectral features are mapped and weighted by the attention layer, so that different spectral information can obtain a more reasonable weight allocation in the fusion process, which is conducive to improving the adaptability and stability of spectral fusion.

[0065] In this embodiment, the speech synthesis method operates on an electronic device (e.g., Figure 1 The server shown can receive instructions or acquire data via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future wireless connection methods.

[0066] It should be emphasized that, to further ensure the privacy and security of the aforementioned input text information, the input text information can also be stored in a blockchain node.

[0067] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0068] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0069] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0070] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0071] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0072] Further reference Figure 3 As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a speech synthesis device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0073] like Figure 3 As shown, the speech synthesis method 300 described in this embodiment includes: Text acquisition module 301 is used to acquire the input text sequence; The text encoding module 302 is used to obtain the text feature vector corresponding to the input text sequence through a preset text encoder; The acoustic coding module 303 is used to acquire input sound conditions and obtain the sound feature vector corresponding to the input sound conditions through a preset acoustic encoder; The audio-text splicing module 304 is used to splice the text feature vector and the audio feature vector to obtain the first audio feature vector; The first spectrum generation module 305 is used to input the first speech feature vector into the pre-trained first speech processing model to obtain the first Mel spectrum; Semantic recognition module 306 is used to perform semantic recognition on the input text sequence through a pre-trained large language model to obtain a semantic token sequence; The second spectrum generation module 307 is used to input the semantic token sequence into the pre-trained second speech processing model to obtain the second Mel spectrum; The spectrum fusion module 308 is used to perform spectrum fusion on the first Mel spectrum and the second Mel spectrum to obtain the speech fusion spectrum; The vocode conversion module 309 is used to input the speech fusion spectrum into a preset vocoder to obtain synthesized speech.

[0074] Furthermore, the spectrum fusion module 308 specifically includes: The alignment processing submodule is used to perform time-dimensional and frequency-dimensional alignment processing on the first Mel spectrum and the second Mel spectrum; The weight configuration submodule is used to determine the first weight coefficient corresponding to the first Mel spectrum and the second weight coefficient corresponding to the second Mel spectrum according to the preset fusion strategy, wherein the sum of the first weight coefficient and the second weight coefficient is 1; The weighted processing submodule is used to perform weighted processing on the first Mel spectrum and the second Mel spectrum based on the first weight coefficient and the second weight coefficient, respectively, to obtain the weighted first spectrum and the weighted second spectrum. The spectrum superposition submodule is used to perform superposition operations on the weighted first spectrum and the weighted second spectrum to obtain the speech fusion spectrum.

[0075] Furthermore, the spectrum fusion module 308 also includes: The alignment processing submodule is used to perform time-dimensional and frequency-dimensional alignment processing on the first Mel spectrum and the second Mel spectrum; The frequency band division submodule is used to divide the Mel spectrum into low-frequency bands and high-frequency bands according to a preset frequency division dimension. The low-frequency processing submodule is used to generate a low-frequency fused spectrum by using a weighted processing method based on the first Mel spectrum for the low-frequency band. The high-frequency processing submodule is used to generate a high-frequency fusion spectrum by employing a weighted processing method based on the second Mel spectrum for the high-frequency band. The spectrum splicing submodule is used to splice the low-frequency fused spectrum with the high-frequency fused spectrum to obtain the speech fused spectrum.

[0076] Furthermore, the frequency band allocation unit specifically includes: The channel counting unit is used to obtain the total number of Mel frequency channels in the first Mel spectrum and the second Mel spectrum; The channel division unit is used to determine the Mel channel range corresponding to the low frequency band and the Mel channel range corresponding to the high frequency band according to the preset frequency division parameters. The frequency band division unit is used to extract the corresponding low-frequency band spectral components and high-frequency band spectral components from the first Mel channel spectrum and the second Mel spectrum, respectively, based on the Mel channel range.

[0077] Furthermore, the spectrum fusion module 308 also includes: The alignment processing submodule is used to perform time-dimensional and frequency-dimensional alignment processing on the first Mel spectrum and the second Mel spectrum; The attention weight submodule is used to take the first Mel spectrum and the second Mel spectrum as input to the preset fusion network, extract spectral features through the fusion network and generate corresponding attention weights. The dynamic weighting submodule is used to dynamically weight the first Mel spectrum and the second Mel spectrum based on attention weights to obtain the first weighted spectral features and the second weighted spectral features. The spectrum fusion submodule is used to perform fusion operations on the first weighted spectrum feature and the second weighted spectrum feature to obtain the speech fusion spectrum.

[0078] Furthermore, the attention weight submodule specifically includes: The feature normalization unit is used to perform feature normalization processing on the first Mel spectrum and the second Mel spectrum to obtain standardized first spectral features and second spectral features. The feature splicing unit is used to splice the standardized first spectral feature and the second spectral feature to form a joint spectral feature; The high-dimensional feature extraction unit is used to input the joint spectral features into the feature extraction module in the fusion network to extract high-dimensional spectral representation features. The attention computation unit is used to calculate the corresponding attention weight parameters by fusing attention layers in the network based on high-dimensional spectral representation features. The normalization processing unit is used to normalize the attention weight parameters to obtain the attention weights.

[0079] Furthermore, the attention computation unit specifically includes: The linear mapping subunit is used to input high-dimensional spectral representation features into the linear mapping region in the attention layer to obtain intermediate attention feature representations; Nonlinear activation subunits are used to perform nonlinear activation processing on intermediate attention feature representations; The attention weight calculation subunit is used to generate attention weight parameters corresponding to the first Mel spectrum and the second Mel spectrum, respectively, through the weight calculation area based on the activated intermediate attention feature representation.

[0080] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0081] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with memory 41, processor 42, and network interface 43 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0082] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0083] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for speech synthesis methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0084] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions for the speech synthesis method.

[0085] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0086] This application also provides an embodiment, namely, a computer device including a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the speech synthesis method described above.

[0087] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the speech synthesis method described above.

[0088] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0089] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0090] It should be noted that the software tools or components not belonging to this company that appear in the various embodiments of this application are merely illustrative examples and do not represent actual use.

[0091] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A speech synthesis method, characterized in that, include: Obtain the input text sequence; The text feature vector corresponding to the input text sequence is obtained through a preset text encoder; The input sound conditions are obtained, and the sound feature vector corresponding to the input sound conditions is obtained through a preset acoustic encoder; The first speech feature vector is obtained by concatenating the text feature vector and the speech feature vector. The first speech feature vector is input into the pre-trained first speech processing model to obtain the first Mel spectrum; The input text sequence is semantically recognized by a pre-trained large language model to obtain a semantic token sequence; The semantic token sequence is input into a pre-trained second speech processing model to obtain the second Mel spectrum; The first Mel spectrum and the second Mel spectrum are spectrally fused to obtain the speech fused spectrum; The speech fusion spectrum is input into a preset vocoder to obtain synthesized speech.

2. The speech synthesis method as described in claim 1, characterized in that, The step of fusing the first Mel spectrum and the second Mel spectrum to obtain the speech fused spectrum specifically includes: Alignment processing is performed on the first Mel spectrum and the second Mel spectrum in terms of time and frequency dimensions; The first weighting coefficient corresponding to the first Mel spectrum and the second weighting coefficient corresponding to the second Mel spectrum are determined according to the preset fusion strategy, wherein the sum of the first weighting coefficient and the second weighting coefficient is 1. Based on the first weighting coefficient and the second weighting coefficient, the first Mel spectrum and the second Mel spectrum are weighted respectively to obtain the weighted first spectrum and the weighted second spectrum. The speech fusion spectrum is obtained by superimposing the weighted first spectrum and the weighted second spectrum.

3. The speech synthesis method as described in claim 1, characterized in that, The step of fusing the first Mel spectrum and the second Mel spectrum to obtain the speech fused spectrum specifically includes: Alignment processing is performed on the first Mel spectrum and the second Mel spectrum in terms of time and frequency dimensions; The Mel spectrum is divided into low-frequency and high-frequency bands according to the preset frequency division dimensions; For the low-frequency band, a weighted processing method based on the first Mel spectrum is used to generate a low-frequency fused spectrum; For the aforementioned high-frequency band, a weighted processing method based on the second Mel spectrum is used to generate a high-frequency fused spectrum; The low-frequency fusion spectrum is spliced ​​with the high-frequency fusion spectrum to obtain the speech fusion spectrum.

4. The speech synthesis method as described in claim 3, characterized in that, The step of dividing the Mel spectrum into low-frequency and high-frequency bands according to a preset frequency division dimension specifically includes: Obtain the total number of Mel frequency channels in the first Mel spectrum and the second Mel spectrum; The Mel channel range corresponding to the low frequency band and the Mel channel range corresponding to the high frequency band are determined according to the preset frequency division parameters. Based on the Mel channel range, the corresponding low-frequency band spectral components and high-frequency band spectral components are extracted from the first Mel spectrum and the second Mel spectrum, respectively.

5. The speech synthesis method as described in claim 1, characterized in that, The step of fusing the first Mel spectrum and the second Mel spectrum to obtain the speech fused spectrum specifically includes: Alignment processing is performed on the first Mel spectrum and the second Mel spectrum in terms of time and frequency dimensions; The first Mel spectrum and the second Mel spectrum are used as inputs to a preset fusion network, which extracts spectral features and generates corresponding attention weights. Based on the attention weights, the first Mel spectrum and the second Mel spectrum are dynamically weighted to obtain the first weighted spectral features and the second weighted spectral features. The first weighted spectral feature and the second weighted spectral feature are fused together to obtain the speech fusion spectrum.

6. The speech synthesis method as described in claim 5, characterized in that, The step of using the first Mel spectrum and the second Mel spectrum as input to a preset fusion network, and extracting spectral features and generating corresponding attention weights through the fusion network, specifically includes: The first Mel spectrum and the second Mel spectrum are subjected to feature normalization processing to obtain standardized first spectral features and second spectral features; The standardized first spectral feature and the second spectral feature are spliced ​​together to form a joint spectral feature; The joint spectral features are input into the feature extraction module in the fusion network to extract high-dimensional spectral representation features; Based on the high-dimensional spectral representation features, the attention layer in the fusion network is used to calculate the corresponding attention weight parameters; The attention weight parameters are normalized to obtain the attention weights.

7. The speech synthesis method as described in claim 6, characterized in that, The step of calculating the corresponding attention weight parameters based on the high-dimensional spectral representation features through the attention layer in the fusion network specifically includes: The high-dimensional spectral representation features are input into the linear mapping region in the attention layer to obtain the intermediate attention feature representation; The intermediate attention feature representation is subjected to nonlinear activation processing; Based on the activated intermediate attention feature representation, attention weight parameters corresponding to the first Mel spectrum and the second Mel spectrum are generated through the weight calculation region.

8. A speech synthesis method, characterized in that, include: The text acquisition module is used to acquire the input text sequence; The text encoding module is used to obtain the text feature vector corresponding to the input text sequence through a preset text encoder; An acoustic coding module is used to acquire input sound conditions and obtain the sound feature vector corresponding to the input sound conditions through a preset acoustic encoder. The audio-text splicing module is used to splice the text feature vector and the audio feature vector to obtain a first speech feature vector; The first spectrum generation module is used to input the first speech feature vector into the pre-trained first speech processing model to obtain the first Mel spectrum; The semantic recognition module is used to perform semantic recognition on the input text sequence using a pre-trained large language model to obtain a semantic token sequence; The second spectrum generation module is used to input the semantic token sequence into a pre-trained second speech processing model to obtain the second Mel spectrum; The spectrum fusion module is used to perform spectrum fusion on the first Mel spectrum and the second Mel spectrum to obtain the speech fusion spectrum; The vocode conversion module is used to input the speech fusion spectrum into a preset vocoder to obtain synthesized speech.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the speech synthesis method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the speech synthesis method as described in any one of claims 1 to 7.