An end-to-end speech synthesis method, apparatus, computer device, and storage medium

CN121075306BActive Publication Date: 2026-09-01PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511066062.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2026-09-01
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

[0006]本申请实施例的目的在于提出一种端到端语音合成方法、装置、计算机设备及存储介质,以解决现有技术不能从带噪低采样率语音准确提取有效语音特征并合成高质量语音的技术问题

Benefits of technology

[0026] This application proposes an end-to-end speech synthesis method. It obtains a high-sampling-rate reference audio set and corresponding input text, downsamples the high-sampling-rate reference audio set to obtain a low-sampling-rate reference audio set and a medium-sampling-rate reference audio set, then injects noise with different signal-to-noise ratios into the low-sampling-rate reference audio set to obtain a noisy low-sampling-rate audio sample set. A multimodal sample set is then constructed, and a pre-trained speech synthesis model is trained using this multimodal sample set to obtain the final end-to-end speech synthesis model. During training, the model parameters are dynamically adjusted through weighted balancing of multiple loss functions, improving the model's ability to extract effective speech features from noisy low-sampling-rate speech. Simultaneously, it effectively converts low-sampling-rate speech into high-sampling-rate speech, improving the naturalness and clarity of the synthesized speech and ensuring high-quality speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075306B_ABST
    Figure CN121075306B_ABST
Patent Text Reader

Abstract

This application belongs to the field of artificial intelligence technology and relates to an end-to-end speech synthesis method. The method includes downsampling a high-sampling-rate reference audio set to obtain a low-sampling-rate reference audio set; injecting noise with different signal-to-noise ratios into the low-sampling-rate reference audio set to obtain a noisy low-sampling-rate audio sample set, and constructing a multimodal sample set with the corresponding input text; using the multimodal sample set to train a speech synthesis model to obtain a predicted high-sampling-rate audio waveform; adjusting the model parameters based on the calculated total loss function, and continuing iterative training to obtain the final end-to-end speech synthesis model; inputting the text to be synthesized and the noisy low-sampling-rate reference speech into the end-to-end speech synthesis model to synthesize the target speech waveform. This application also provides an end-to-end speech synthesis device, computer equipment, and storage medium. This application can be applied to business management system programs in financial technology, healthcare, etc., to improve the naturalness and clarity of synthesized speech and ensure high-quality speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and is applied to online processing business scenarios such as fintech / digital healthcare. In particular, it relates to an end-to-end speech synthesis method, device, computer equipment, and storage medium. Background Technology

[0002] With continuous breakthroughs in artificial intelligence technology, large-scale speech synthesis models have experienced unprecedented development in recent years. As an important human-computer interaction tool, speech synthesis technology has permeated multiple fields such as intelligent customer service, voice assistants, and education and training, greatly improving user experience and efficiency. Particularly in scenarios such as financial and medical services, speech synthesis technology is used in intelligent voice customer service for answering financial questions and promoting medical knowledge. Its natural and fluent voice and personalized expression capabilities bring customers more humanized financial and medical services. Especially in the financial sector, speech synthesis technology is also widely used in telephone banking, intelligent investment advisory, and other scenarios, enhancing customer trust and satisfaction through the synthesis of realistic voices.

[0003] Traditional speech synthesis systems typically consist of multiple independent modules, such as a text analysis front-end, an acoustic model, and an audio synthesis module. These modules have complex dependencies, which not only requires extensive domain expertise during system construction, increasing the difficulty and cost of system development, but also leads to severe error propagation problems when processing noisy, low-sampling-rate speech. For example, during acoustic model training, if the input is noisy, low-sampling-rate speech features, the information loss and distortion caused by noise and the low sampling rate will result in inaccurate learned acoustic features, thus affecting the speech quality generated by the audio synthesis module. In intelligent customer service systems in the financial sector, this error propagation can lead to discontinuous and semantically incoherent synthesized speech, impacting customer experience; in the medical field, it can result in inaccurate voice prompts, posing potential risks to medical work. Furthermore, traditional systems have limited ability to increase the sampling rate, making it difficult to effectively recover lost high-frequency information from low-sampling-rate speech to generate high-quality audio with a high sampling rate. This makes traditional speech synthesis systems inadequate when facing the diverse demands for high-quality speech in the financial and medical fields.

[0004] Most existing speech synthesis models are trained on specific clean speech data and at specific sampling rates. When inputting noisy, low-sampling-rate speech, the models cannot process it effectively and lack adaptability to such complex inputs. The models cannot accurately extract effective speech features from noisy, low-sampling-rate speech, resulting in synthesized speech with problems such as residual noise, distorted timbre, and low naturalness. In the financial sector, this may make the voice of intelligent customer service sound stiff and unnatural, reducing customer satisfaction; in the medical field, it may lead to voice prompts being difficult for medical staff to understand accurately, affecting the normal operation of medical work.

[0005] In summary, existing speech synthesis technologies cannot meet the demand for high-quality speech synthesis in diverse real-world scenarios. Summary of the Invention

[0006] The purpose of this application is to provide an end-to-end speech synthesis method, apparatus, computer device, and storage medium to solve the technical problem that the prior art cannot accurately extract effective speech features from noisy, low-sampling-rate speech and synthesize high-quality speech.

[0007] Firstly, an end-to-end speech synthesis method is provided, which adopts the following technical solution:

[0008] Obtain a high sampling rate reference audio set and the corresponding input text, and downsample the high sampling rate reference audio set to obtain a low sampling rate reference audio set and a medium sampling rate reference audio set;

[0009] Noise with different signal-to-noise ratios is injected into the low-sampling-rate reference audio set to obtain a noisy low-sampling-rate audio sample set. The noisy low-sampling-rate audio sample set and the corresponding input text are then combined to generate a multimodal sample set.

[0010] The multimodal sample set is input into a pre-trained speech synthesis model, and the speech synthesis model performs noise reduction, prosody matching, and super-resolution processing on the multimodal sample set to obtain a predicted high sampling rate audio waveform.

[0011] According to the preset loss function, the weighted loss is calculated based on the predicted high sampling rate audio waveform, the low sampling rate reference audio set, the medium sampling rate reference audio set, and the high sampling rate reference audio set to obtain the total loss function;

[0012] The model parameters of the speech synthesis model are adjusted based on the total loss function, and iterative training continues until the iteration stopping condition is met to obtain the final end-to-end speech synthesis model.

[0013] Obtain the text to be synthesized and the corresponding noisy low-sampling-rate reference speech, input the text to be synthesized and the noisy low-sampling-rate reference speech into the end-to-end speech synthesis model, and synthesize the target speech waveform.

[0014] Secondly, an end-to-end speech synthesis device is provided, which adopts the following technical solution:

[0015] The acquisition module is used to acquire a high sampling rate reference audio set and a corresponding input text, and to downsample the high sampling rate reference audio set to obtain a low sampling rate reference audio set and a medium sampling rate reference audio set.

[0016] The module is used to inject noise with different signal-to-noise ratios into the low sampling rate reference audio set to obtain a noisy low sampling rate audio sample set, and to combine the noisy low sampling rate audio sample set and the corresponding input text to generate a multimodal sample set.

[0017] The training module is used to input the multimodal sample set into a pre-trained speech synthesis model, and to perform noise reduction, prosody matching, and super-resolution processing on the multimodal sample set through the speech synthesis model to obtain a predicted high sampling rate audio waveform.

[0018] The loss calculation module is used to calculate the weighted loss according to the predicted high sampling rate audio waveform, the low sampling rate reference audio set, the medium sampling rate reference audio set, and the high sampling rate reference audio set, and obtain the total loss function according to the preset loss function;

[0019] An iterative module is used to adjust the model parameters of the speech synthesis model based on the total loss function, continue iterative training until the iteration stopping condition is met, and obtain the final end-to-end speech synthesis model.

[0020] The synthesis module is used to acquire the text to be synthesized and the corresponding noisy low-sampling-rate reference speech, input the text to be synthesized and the noisy low-sampling-rate reference speech into the end-to-end speech synthesis model, and synthesize the target speech waveform.

[0021] Thirdly, a computer device is provided that adopts the technical solution described below:

[0022] The computer device includes a memory and a processor, the memory storing computer-readable instructions, and the processor executing the computer-readable instructions to implement the steps of the end-to-end speech synthesis method as described above.

[0023] Fourthly, a computer-readable storage medium is provided, which adopts the technical solution described below:

[0024] The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the end-to-end speech synthesis method as described above.

[0025] Compared with the prior art, this application has the following main advantages:

[0026] This application proposes an end-to-end speech synthesis method. It obtains a high-sampling-rate reference audio set and corresponding input text, downsamples the high-sampling-rate reference audio set to obtain a low-sampling-rate reference audio set and a medium-sampling-rate reference audio set, then injects noise with different signal-to-noise ratios into the low-sampling-rate reference audio set to obtain a noisy low-sampling-rate audio sample set. A multimodal sample set is then constructed, and a pre-trained speech synthesis model is trained using this multimodal sample set to obtain the final end-to-end speech synthesis model. During training, the model parameters are dynamically adjusted through weighted balancing of multiple loss functions, improving the model's ability to extract effective speech features from noisy low-sampling-rate speech. Simultaneously, it effectively converts low-sampling-rate speech into high-sampling-rate speech, improving the naturalness and clarity of the synthesized speech and ensuring high-quality speech synthesis. Attached Figure Description

[0027] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0028] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;

[0029] Figure 2 This is a flowchart of an embodiment of the end-to-end speech synthesis method according to this application;

[0030] Figure 3 yes Figure 2 A flowchart of a specific implementation of step S203;

[0031] Figure 4 This is a schematic diagram of the structure of an embodiment of the end-to-end speech synthesis apparatus according to this application;

[0032] Figure 5 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0034] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0035] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0036] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0037] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0038] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptop computer 1011, tablet computer 1012 or mobile phone 1013, terminal device 101 can also be e-book reader, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer and desktop computer, etc.

[0039] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0040] It should be noted that the end-to-end speech synthesis method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the end-to-end speech synthesis device is generally located in the server / terminal device.

[0041] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0042] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the end-to-end speech synthesis method according to this application, including the following steps:

[0043] Step S201: Obtain the high sampling rate reference audio set and the corresponding input text, and downsample the high sampling rate reference audio set to obtain the low sampling rate reference audio set and the medium sampling rate reference audio set.

[0044] High sampling rate refers to audio with a sampling frequency greater than or equal to 44.1kHz. The high sampling rate reference audio set contains a large number of high sampling rate reference audio samples, which can be acquired through public data platforms or using high-quality professional recording equipment, such as high-fidelity microphones, audio interfaces, and professional recording software. The collected raw audio samples are preprocessed, including noise removal and standardization, to obtain the preprocessed high sampling rate reference audio set. The input text corresponding to the high sampling rate reference audio set is then obtained, and the input text is time-aligned with the corresponding audio. The input text can be generated using a trained speech recognition model.

[0045] In this embodiment, the high sampling rate reference audio set includes high sampling rate audio samples with different accents and dialects, as well as various speaking speeds and intonations.

[0046] For example, in the financial field, a high-sampling-rate reference audio set could be audio data containing financial terminology, financial market data, stock trading information, financial reports, account inquiries, etc., or it could be recordings of conversations between financial customer service representatives and customers. In the medical field, a high-sampling-rate reference audio set could be audio data containing medical diagnoses, treatment plans, drug names, disease names, etc., or it could be audio data containing medical guidance information such as drug usage instructions, health check guidelines, and disease prevention advice, or it could be recordings of conversations between doctors and patients of different ages and genders during diagnosis and treatment.

[0047] Downsampling is performed on the high sampling rate reference audio set to obtain a low sampling rate reference audio set and a medium sampling rate reference audio set. The low sampling rate refers to audio with a sampling frequency between 8kHz and 16kHz; the medium sampling rate refers to audio with a sampling frequency between 16kHz and 44.1kHz.

[0048] Downsampling refers to reducing the sampling frequency of an audio signal, that is, converting high-sampling-rate audio to low-sampling-rate audio. Downsampling methods can be selected from at least one of the following: direct decimation, anti-aliasing filtering, polyphase filtering, or neural network-based downsampling methods.

[0049] Step S202: Inject noise with different signal-to-noise ratios into the low sampling rate reference audio set to obtain a noisy low sampling rate audio sample set, and combine the noisy low sampling rate audio sample set with the corresponding input text to generate a multimodal sample set.

[0050] In this embodiment, noise and reverberation with different signal-to-noise ratios are injected into a low-sampling-rate reference audio set. Noise refers to unwanted, random, or irregular interference components in the audio signal; reverberation refers to the acoustic phenomenon of continuous attenuation formed after sound is reflected multiple times in a closed space, which is a delayed and superimposed version of the target signal.

[0051] The process involves acquiring preset signal-to-noise ratios (SNRs) and extracting corresponding noise data from a noise database. This noise data is then injected into a low-sampling-rate reference audio set. Specifically, based on the noise and reverberation characteristics of the actual application scenario, data with similar noise and reverberation features are selected from open-source datasets such as AudioSet and BIRD (BigImpulse Response Dataset) to construct a scene dataset. A data fusion algorithm is then used to fuse the selected scene dataset and the low-sampling-rate reference audio set, generating a noisy low-sampling-rate audio sample set. This noisy low-sampling-rate audio sample set contains audio data with a realistic scene feel. The data fusion algorithm can employ weighted averaging, signal superposition, or a Gaussian mixture model (GMM), among others.

[0052] In some alternative implementations, data augmentation techniques are used to enrich and expand the noisy low-sampling-rate audio sample set, including changing the speech rate, volume, pitch, etc., to obtain an enhanced noisy low-sampling-rate audio sample set, which is then used to construct a multimodal sample set.

[0053] Among them, the multimodal sample set is labeled with a clean reference audio label. The noisy low sampling rate audio sample set corresponds to the low sampling rate reference audio set, the medium sampling rate reference audio set, and the high sampling rate reference audio set. The clean reference audio label includes the low sampling rate reference audio, the medium sampling rate reference audio, and the high sampling rate reference audio corresponding to each noisy low sampling rate audio sample in the noisy low sampling rate audio sample set.

[0054] By injecting noise and reverberation, the description of special scenarios is expanded, enabling the model to learn the variation patterns and characteristics of speech under different noise backgrounds. As a result, the model can accurately simulate speech in various special scenarios according to the needs when synthesizing speech, which significantly improves the applicability and expressiveness of speech synthesis models in complex real-world scenarios.

[0055] Step S203: Input the multimodal sample set into the pre-trained speech synthesis model, and perform noise reduction, prosody matching and super-resolution processing on the multimodal sample set through the speech synthesis model to obtain the predicted high sampling rate audio waveform.

[0056] In this embodiment, a speech synthesis model is used to convert noisy, low-sampling-rate reference audio and corresponding text into high-sampling-rate, high-quality synthesized speech. The speech synthesis model includes a denoising sub-model, a speech synthesis sub-model, a supramolecular model, and a vocoder. The pre-trained speech synthesis model refers to the speech synthesis model that has already undergone the first stage of training. In the first stage of training, the denoising sub-model, the speech synthesis sub-model, the supramolecular model, and the vocoder are trained separately. Then, the denoising sub-model, the speech synthesis sub-model, the supramolecular model, and the vocoder trained in the first stage are concatenated to obtain the pre-trained speech synthesis model. The pre-trained speech synthesis model is then subjected to end-to-end second-stage training.

[0057] The multimodal sample set is input into the pre-trained speech synthesis model. The multimodal sample set is then subjected to noise reduction, prosody matching, and super-resolution processing through the noise reduction sub-model, speech synthesis sub-model, supramolecular model, and vocoder in sequence. Finally, a predicted high sampling rate audio waveform is synthesized.

[0058] In some alternative implementations, see [link to relevant documentation]. Figure 3 As shown, the steps described above for performing noise reduction, prosody matching, and super-resolution processing on a multimodal sample set using a speech synthesis model to obtain a predicted high-sampling-rate audio waveform include:

[0059] Step S301: The noisy low-sampling-rate audio sample set is processed by the noise reduction sub-model to generate noise-reduced audio features.

[0060] The noise reduction sub-model aims to extract relatively clean speech features from noisy low-sampling-rate audio samples in a noisy low-sampling-rate audio sample set, providing high-quality input for subsequent speech synthesis. Specifically, the noise reduction sub-model analyzes the spectral characteristics of the input noisy low-sampling-rate audio samples, separates the audio and noise components, and generates feature data that can characterize the audio content, i.e., noise-reduced audio features.

[0061] In this embodiment, the noise reduction frequency feature is the Mel spectrum feature, which is a commonly used speech domain feature that can effectively characterize the time-frequency distribution of audio signals.

[0062] In some optional implementations of this embodiment, the steps of processing the noisy low-sampling-rate audio sample set through a noise reduction sub-model to generate noise-reduced audio features include:

[0063] The noisy, low-sampling-rate audio sample set is input into the denoising sub-model, which includes a feature input layer, a feature extraction layer, and a recurrent layer.

[0064] The noisy low-sampling-rate audio sample set is input into the feature input layer for audio conversion to obtain noisy low-sampling-rate spectral features;

[0065] Local audio enhancement features are obtained by extracting local features from the noisy low sampling rate spectral features through a feature extraction layer.

[0066] By analyzing the temporal dynamic characteristics of local audio enhancement features through a recurrent layer, noise reduction frequency features are obtained.

[0067] The noise reduction sub-model is a deep learning-based architecture that performs nonlinear transformations on noisy, low-sampling-rate audio samples to separate audio and noise components. Specifically, the noise reduction sub-model adopts a MossFormer2-based model framework, including a feature input layer, a feature extraction layer, and a recurrent layer. The feature input layer converts the noisy, low-sampling-rate audio samples into a frequency domain representation. Specifically, it uses a Fast Fourier Transform to convert each frame of the noisy, low-sampling-rate audio sample into the frequency domain, obtaining a short-time spectrum, i.e., the noisy, low-sampling-rate spectral features. The feature extraction layer enhances the input noisy, low-sampling-rate spectral features, extracting audio features from high-energy regions to obtain local audio enhancement features. The recurrent layer analyzes the temporal dynamics of the local audio enhancement features, identifies audio and noise components, suppresses the noise components, and outputs the noise-reduced frequency features.

[0068] In one specific example, the feature extraction layer can adopt a Transformer encoder structure, which consists of multiple stacked encoder layers. The multi-head attention mechanism of each encoder layer identifies the spectral differences between noise components and speech components in the noisy low sampling rate spectral features, and extracts audio features in high-energy regions to obtain local audio enhancement features.

[0069] The recurrent layer can employ a multi-layer feedforward sequential memory network (FSMN). This network includes memory modules that construct audio temporal features from local audio enhancement features to predict whether each frame is audio, thus outputting denoising audio features. It should be understood that the memory modules store historical and future information useful for determining the current speech frame. Skip connections are added between memory modules, allowing historical information from lower-level memory modules to directly flow into higher-level memory modules. During backpropagation, gradients from higher-level memory modules also flow directly into lower-level memory modules, helping to overcome the vanishing gradient problem.

[0070] By processing noisy low-sampling-rate audio sample sets through a noise reduction sub-model, audio and noise can be accurately separated, thereby generating high-quality noise-reduced audio features under different noise environments and improving the accuracy of effective speech feature extraction.

[0071] Step S302: Using the speech synthesis sub-model, generate medium-sampling-rate synthesized audio features based on the noise reduction frequency features and the input text.

[0072] The speech synthesis sub-model maps denoised audio features and input text to higher-quality, medium-sampling-rate synthesized audio features. Specifically, by combining the acoustic information of the denoised audio features with the semantic information of the input text, the speech synthesis sub-model generates synthesized audio features that are consistent with the text content and have a natural rhythm.

[0073] In this embodiment, the noise-reduced frequency features are typically noise-reduced frequency Mel-spectrums, with 24 or 40 dimensions, containing spectral information of the audio. The corresponding input text is a character sequence containing pinyin annotations or prosodic tags, used to guide the semantic and prosodic expression of speech synthesis.

[0074] In some optional implementations of this embodiment, the step of generating medium-sampling-rate synthesized audio features based on noise-reduced audio features and input text using a speech synthesis sub-model includes:

[0075] The noise reduction frequency features and input text are input into the speech synthesis sub-model, which includes a text feature layer, a prosodic feature layer, and a speech synthesis layer.

[0076] The text feature layer extracts features from the input text to obtain a semantic feature vector;

[0077] The prosodic features of the noise-reduced frequency features are extracted through the prosodic feature layer to obtain the prosodic feature vector;

[0078] The semantic feature vector and prosodic feature vector are input into the speech synthesis layer. The semantic feature vector and prosodic feature vector are fused based on the stream matching algorithm to generate medium sampling rate synthesized audio features.

[0079] The speech synthesis sub-model uses the F5-TTS main network structure based on the FlowMatching algorithm to process noise reduction audio features and input text, generating synthesized audio features with a medium sampling rate.

[0080] The text feature layer is used to obtain semantic feature vectors from the input text. Specifically, the text feature layer includes a text embedding layer, a Transformer encoder layer, and a pooling layer. The input text is fed into the text feature layer, where it is semantically encoded through the text embedding layer, Transformer encoder layer, and pooling layer to obtain medium-sampling-rate synthesized audio features. The Transformer encoder layer contains multiple stacked encoders, each including a multi-head attention mechanism layer and a feedforward neural network layer.

[0081] Furthermore, the text data in the text dataset is vector-embedded through a text embedding layer to obtain text embedding vectors; the text embedding vectors are input into the Transformer encoding layer, and semantic features are extracted through a multi-layer self-attention mechanism, that is, the semantic information and contextual relationships of the text in the text embedding vectors are captured to obtain semantic encoding features; the semantic encoding features are pooled through a pooling layer to obtain semantic feature vectors.

[0082] The prosodic feature layer is used to extract prosodic features from the noise-reduced frequency features, generating a prosodic feature vector that matches the length of the input text. Specifically, the prosodic feature layer analyzes the pitch, intonation, and rhythm features of the noise-reduced frequency features to extract a vector representation reflecting the prosody of speech, namely the prosodic feature vector. These prosodic features include pitch, duration, intensity, rhythm, intonation, and emotion.

[0083] In this embodiment, the prosodic feature layer can employ an adversarial prosodic encoder to extract and encode the prosodic information from the noise-reduced audio features, thereby obtaining a prosodic feature vector. The adversarial prosodic encoder extracts prosodic information from the noise-reduced audio features through competitive learning and generates a prosodic feature vector. Specifically, the adversarial prosodic encoder includes a generator network and a discriminator network. The generator network extracts prosodic features from the noise-reduced audio features, and the discriminator network evaluates the authenticity of the extracted audio features. The feedback signal from the discriminator network is used to enhance the prosodic feature expressive power of the generator network.

[0084] The speech synthesis layer employs a decoder structure, comprising multiple stacked decoders. Each decoder includes a masked multi-head attention mechanism layer and a fully connected layer. Layer normalization is performed after the masked multi-head attention mechanism layer, and post-normalization is performed after the fully connected layer. Specifically, based on a stream matching algorithm, semantic feature vectors and prosodic feature vectors are aligned and mapped to generate a fused feature vector. This fused feature vector is input to the first decoder. Each head of the masked multi-head attention mechanism layer performs self-attention calculation on the fused feature vector, obtaining the output of each head. The outputs of multiple heads are fused to obtain the attention feature representation of that decoder. The attention feature representation and the fused feature vector are used as input to the next decoder, and the calculation is repeated until all decoders have completed processing, outputting the final attention feature representation. This final attention feature representation is then mapped to a medium-sampling-rate feature space to obtain the medium-sampling-rate synthesized audio features.

[0085] By fusing semantic and prosodic information through a speech synthesis sub-model, the naturalness of synthesized speech can be enhanced, the mechanical feel of the synthesized speech can be reduced, and the user experience can be improved.

[0086] Step S303: Super-resolution calculation is performed on the medium sampling rate synthesized audio features using a supramolecular model to obtain the high sampling rate synthesized audio features.

[0087] By analyzing the spectral characteristics of medium-sampling-rate synthesized audio features using a supramolecular model, high-frequency components are reconstructed to compensate for the lack of high-frequency components in low-sampling-rate signals, thereby generating high-sampling-rate synthesized audio features that can characterize high-fidelity speech and enhance the detail and clarity of the speech.

[0088] In some optional implementations of this embodiment, the step of calculating the oversampling rate of the medium-sampling-rate synthesized audio features using a supramolecular model to obtain the high-sampling-rate synthesized audio features includes:

[0089] The synthesized audio features at a medium sampling rate are input into a supramolecular model, which includes a high-frequency reconstruction layer and a decoding output layer.

[0090] By using a high-frequency reconstruction layer, a frequency band-aware attention mechanism is employed to extract and reconstruct high-frequency features from the synthesized audio features, thereby obtaining enhanced high-frequency features.

[0091] The enhanced high-frequency features are decoded and generated by the decoding output layer to obtain high-sampling-rate synthesized audio features.

[0092] The medium-sampling-rate synthesized audio features are input into the supramolecular model. A frequency band-aware attention mechanism is used in the high-frequency reconstruction layer to analyze the frequency distribution of the medium-sampling-rate synthesized audio features and dynamically adjust the frequency distribution. Specifically, the frequency band distribution information of the medium-sampling-rate synthesized audio features is obtained, where the frequency band distribution information is the energy distribution of the medium-sampling-rate synthesized audio features in different frequency ranges; the attention weights of the high-frequency components in the frequency band distribution information are calculated, where the attention weights characterize the importance of the high-frequency components in the generation of high-sampling-rate synthesized audio features; the high-frequency components of the medium-sampling-rate synthesized audio features are dynamically adjusted according to the attention weights. In one possible implementation, the adjustment process also includes a smoothing operation, which reduces distortion caused by abrupt changes by interpolating the high-frequency components of adjacent frames; and enhanced high-frequency features are generated based on the adjusted high-frequency components.

[0093] The decoding output layer adopts a Transformer decoding structure, which is composed of multiple decoders stacked together. It captures the long-distance dependencies between enhanced high-frequency features through the multi-head attention mechanism and cross-attention mechanism of each encoder, and obtains the predicted high-sampling-rate synthesized audio features.

[0094] The frequency band awareness attention mechanism introduced by the supramolecular module can more accurately identify the frequency band regions that need to be enhanced, improve the generation effect of high-frequency components, ensure the accuracy of high-frequency reconstruction, and thus improve the clarity of synthesized speech.

[0095] Step S304: The high sampling rate synthesized audio features are converted using a vocoder to synthesize a predicted high sampling rate audio waveform.

[0096] The vocoder submodule converts high-sampling-rate synthesized audio features into directly playable high-sampling-rate synthesized audio waveforms. Specifically, the vocoder submodule analyzes the spectral characteristics of the high-sampling-rate synthesized audio features, reconstructs the time-domain waveform of the audio signal, and generates high-fidelity speech output.

[0097] In some alternative implementations, the steps described above, which involve using a vocoder to perform audio conversion on the high-sampling-rate synthesized audio features and predict the synthesized audio waveform, include:

[0098] High-sampling-rate synthesized audio features are input into a vocoder, which includes an input layer, a feature extraction layer, a convolutional layer, a prediction layer, a speech reconstruction layer, and a speech output layer.

[0099] The high-sampling-rate synthesized audio features are preprocessed by the input layer to obtain high-frequency input features;

[0100] High-frequency input features are input into the feature extraction layer for feature extraction to obtain high-frequency spectral features;

[0101] The high-frequency spectral features are fused by channel convolution through a convolutional layer to obtain fused spectral features;

[0102] The fused spectral features are input into the prediction layer for prediction, resulting in an audio temporal spectrogram.

[0103] The audio time-spectrum is restored to an audio signal through the speech restoration layer;

[0104] The audio signal is synthesized by predicting the audio waveform through the output layer, and the predicted audio waveform is output.

[0105] The vocoder employs a neural network-based architecture, capable of mapping high-sampling-rate synthesized audio features to continuous audio waveforms. The vocoder can use the VOCOS model, which does not rely on transposed convolution for upsampling. Instead, it directly predicts the short-time Fourier transform (STFT) spectrum through network stacking, and then uses the inverse short-time Fourier transform (iSTFT) to reconstruct the speech.

[0106] In this embodiment, the input layer preprocesses the high-sampling-rate synthesized audio features, converting them into a data format that the model can process. The feature extraction layer decomposes, extracts, and encodes the high-frequency input features, obtaining high-frequency spectral features including frequency bands, frequency band signal strength, fundamental frequency, and voiced / unvoiced tones. The convolutional layer typically uses the ConvNeXt convolutional module, which consists of depthwise convolution and two pointwise convolutions. STFT spectrum reconstruction is achieved through residual connections, yielding the fused spectral features. The prediction layer predicts the audio time-frequency spectrum, i.e., the STFT spectrum, based on the extracted fused spectral features. The speech reconstruction layer is an iSTFT layer, which uses the predicted STFT spectrum to reconstruct the audio signal. The output layer generates the final audio waveform and outputs it based on the input audio signal's feature parameters (such as spectral envelope and fundamental frequency) and model parameters through inverse transform or synthesis algorithms.

[0107] In one embodiment, the vocoder introduces a phase reconstruction mechanism that optimizes the smoothness and naturalness of the waveform by predicting the phase information of the synthesized audio features at high sampling rates, thereby significantly improving the listening quality of the synthesized audio.

[0108] By converting high-sampling-rate synthesized audio features into audio and outputting the final synthesized audio waveform, clear and natural audio is ensured, effectively improving the quality and realism of the synthesized audio. The entire speech synthesis process combines deep learning and digital signal processing technologies to meet the demands of high-fidelity voice interaction, enhance user experience, and expand application scenarios.

[0109] By leveraging the collaborative efforts of a noise reduction sub-model, a speech synthesis sub-model, a supramolecular model, and a vocoder, the entire process from noise suppression to high-fidelity speech generation is optimized. This reduces complex inter-module interactions and data transfer processes, lowering system complexity and construction costs, and minimizing the impact of inter-module error propagation on the final synthesis effect. This improves audio synthesis efficiency and quality. Specifically, the noise reduction sub-model extracts clear audio features, the speech synthesis sub-model generates prosodic-fitting audio features based on a stream matching algorithm, the supramolecular model reconstructs high-frequency information using a frequency band-aware attention mechanism, and the vocoder sub-model generates high-sampling-rate, high-quality synthesized audio that closely matches the prosodic pattern of the reference audio through inverse Mel-spectrum transformation and waveform reconstruction.

[0110] Step S204: According to the preset loss function, calculate the weighted loss based on the predicted high sampling rate audio waveform, low sampling rate reference audio set, medium sampling rate reference audio set and high sampling rate reference audio set to obtain the total loss function.

[0111] The preset loss function includes the first loss function, Loss. denoise Second loss function Loss sythesis The third loss function Loss super and the fourth loss function Loss vocoder Among them, the first loss function is Loss denoise The second loss function represents the loss between the denoised audio features predicted by the denoising sub-model and the actual low-sampling-rate reference audio features; sythesis The third loss function represents the difference between the mid-sampled rate synthesized audio features predicted by the speech synthesis sub-model and the actual mid-sampled rate reference audio features; super The fourth loss function represents the loss between the high-sampling-rate synthesized audio features predicted by the supramolecular model and the actual high-sampling-rate reference audio features; vocoder This represents the loss between the audio waveform predicted by the vocoder and the actual audio waveform.

[0112] Furthermore, the steps described above, which involve calculating a weighted loss based on the predicted high-sampling-rate audio waveform, low-sampling-rate reference audio set, medium-sampling-rate reference audio set, and high-sampling-rate reference audio set according to a preset loss function, to obtain the total loss function, include:

[0113] Extract low-sampling-rate reference audio features from the low-sampling-rate reference audio set, and calculate the first loss function based on the low-sampling-rate reference audio features;

[0114] Extract the mid-sampling rate reference audio features from the mid-sampling rate reference audio set, and calculate the second loss function based on the mid-sampling rate reference audio features;

[0115] Extract high-sampling-rate reference audio features from the high-sampling-rate reference audio set, and calculate the third loss function based on the high-sampling-rate reference audio features;

[0116] The fourth loss function is calculated based on the predicted high-sampling-rate audio waveform and the corresponding high-sampling-rate reference audio in the high-sampling-rate reference audio set.

[0117] The first loss function, the second loss function, the third loss function, and the fourth loss function are weighted and calculated to obtain the total loss function.

[0118] Specifically, clean audio features are pre-extracted from low-sampling-rate, medium-sampling-rate, and high-sampling-rate reference audio sets as reference standards. Taking the extraction of clean low-sampling-rate reference audio features from the low-sampling-rate reference audio set as an example, the low-sampling-rate reference audio in the low-sampling-rate reference audio set is first segmented into frames to obtain audio frames of a preset length. Then, each audio frame is converted to the frequency domain using a Fast Fourier Transform to obtain a short-time spectrum. Next, a set of Mel filter banks is used to filter the short-time spectrum to generate a clean low-sampling-rate audio Mel spectrum as the low-sampling-rate reference audio feature. The dimension of the low-sampling-rate audio Mel spectrum is consistent with the number of filters in the Mel filter bank. Similarly, clean medium-sampling-rate audio Mel spectra are extracted from the medium-sampling-rate reference audio set as medium-sampling-rate reference audio features, and clean high-sampling-rate audio Mel spectra are extracted from the high-sampling-rate reference audio set as high-sampling-rate reference audio features.

[0119] In one embodiment, the total loss function is calculated as follows:

[0120] Loss = αLoss denoise +βLoss sythesis +γLoss super +δLoss vocoder

[0121] Among them, Loss denoise Loss sythesis Loss super and Loss vocoder Using the mean squared error (MES) approach, specifically, the sum of squared differences between the noise-reduced audio features and the true low-sampling-rate reference audio features in each dimension is calculated to obtain the first loss function, Loss. denoise The second loss function, Loss, is obtained by calculating the sum of squared differences between the synthesized audio features at the medium sampling rate and the true reference audio features at the medium sampling rate in each dimension. sythesis The third loss function, Loss, is obtained by calculating the sum of squared differences between the high-sampling-rate synthesized audio features and the real high-sampling-rate reference audio features in each dimension. superThe fourth loss function, Loss, is obtained by calculating the sum of squared differences between the predicted high-sampling-rate audio waveform and the true audio waveform (i.e., the high-sampling-rate reference audio) in each dimension. vocoder α, β, γ, and δ represent the weights of the first loss function, the second loss function, the third loss function, and the fourth loss function, respectively, and α + β + γ + δ = 1.

[0122] By optimizing the model using multiple loss functions, we can improve the model's prediction accuracy, enhance its robustness, promote its generalization ability, and improve the naturalness of synthesized speech, thereby achieving high-quality speech synthesis.

[0123] Step S205: Adjust the model parameters of the speech synthesis model based on the total loss function, continue iterative training until the iteration stopping condition is met, and obtain the final end-to-end speech synthesis model.

[0124] In this embodiment, the Adam or SGD optimizer is used to adjust the model parameters of the speech synthesis model according to the total loss function. The adjusted model continues to be trained iteratively until the iteration stopping condition is met, that is, the number of iterations reaches the preset number or the loss value of the total loss function does not change significantly, and the final end-to-end speech synthesis model is output.

[0125] Step S206: Obtain the text to be synthesized and the corresponding noisy low-sampling-rate reference speech, input the text to be synthesized and the noisy low-sampling-rate reference speech into the end-to-end speech synthesis model, and synthesize the target speech waveform.

[0126] In this embodiment, the text to be synthesized and the corresponding noisy low-sampling-rate reference speech are obtained. The noisy low-sampling-rate reference speech is used as the imitation object to realize the speech synthesis of the specified text to be synthesized, and the prosody of the synthesized speech is highly consistent with the provided reference speech signal.

[0127] It should be emphasized that, to further ensure the privacy and security of the text to be converted, the text can also be stored in a node of a blockchain.

[0128] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0129] This application trains a pre-trained speech synthesis model using a multimodal sample set to obtain the final end-to-end speech synthesis model. During the training process, the model parameters of the speech synthesis model are dynamically adjusted by weighted balancing of multiple loss functions, which improves the model's ability to extract effective speech features from noisy, low-sampling-rate speech. At the same time, it can effectively convert low-sampling-rate speech into high-sampling-rate speech, ensuring high-quality speech synthesis and improving the naturalness and clarity of the synthesized speech.

[0130] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0131] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0132] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0133] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0134] Further reference Figure 4 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of an end-to-end speech synthesis device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0135] like Figure 4 As shown, the end-to-end speech synthesis device 400 described in this embodiment includes: an acquisition module 401, a construction module 402, a training module 403, a loss calculation module 404, an iteration module 405, and a synthesis module 406. Wherein:

[0136] The acquisition module 401 is used to acquire a high sampling rate reference audio set and a corresponding input text, and to downsample the high sampling rate reference audio set to obtain a low sampling rate reference audio set and a medium sampling rate reference audio set.

[0137] The construction module 402 is used to inject noise with different signal-to-noise ratios into the low sampling rate reference audio set to obtain a noisy low sampling rate audio sample set, and to combine the noisy low sampling rate audio sample set and the corresponding input text to generate a multimodal sample set;

[0138] The training module 403 is used to input the multimodal sample set into the pre-trained speech synthesis model, and to perform noise reduction, prosody matching and super-resolution processing on the multimodal sample set through the speech synthesis model to obtain a predicted high sampling rate audio waveform.

[0139] The loss calculation module 404 is used to calculate the weighted loss according to the predicted high sampling rate audio waveform, the low sampling rate reference audio set, the medium sampling rate reference audio set and the high sampling rate reference audio set according to the preset loss function, so as to obtain the total loss function;

[0140] The iteration module 405 is used to adjust the model parameters of the speech synthesis model based on the total loss function, continue iterative training until the iteration stopping condition is met, and obtain the final end-to-end speech synthesis model.

[0141] The synthesis module 406 is used to acquire the text to be synthesized and the corresponding noisy low-sampling-rate reference speech, input the text to be synthesized and the noisy low-sampling-rate reference speech into the end-to-end speech synthesis model, and synthesize the target speech waveform.

[0142] It should be emphasized that, to further ensure the privacy and security of the text to be synthesized, the text can also be stored in a node of a blockchain.

[0143] Based on the aforementioned end-to-end speech synthesis device 400, the pre-trained speech synthesis model is trained using a multimodal sample set to obtain the final end-to-end speech synthesis model. During the training process, the model parameters of the speech synthesis model are dynamically adjusted through the weighted balancing of multiple loss functions, thereby improving the model's ability to extract effective speech features from noisy, low-sampling-rate speech. At the same time, it can effectively convert low-sampling-rate speech into high-sampling-rate speech, ensuring high-quality speech synthesis and improving the naturalness and clarity of the synthesized speech.

[0144] In some optional implementations, the speech synthesis model includes a noise reduction sub-model, a speech synthesis sub-model, a supramolecular model, and a vocoder, and the training module 403 includes:

[0145] The noise reduction submodule is used to process the noisy low-sampling-rate audio sample set through the noise reduction submodel to generate noise-reduced audio features;

[0146] The matching submodule is used to generate medium-sampling-rate synthesized audio features based on the noise-reduced audio features and the input text through the speech synthesis submodel;

[0147] The supramolecular module is used to perform super-resolution calculations on the medium-sampling-rate synthesized audio features using the supramolecular model to obtain high-sampling-rate synthesized audio features.

[0148] The synthesis submodule is used to perform audio conversion on the high sampling rate synthesized audio features through the vocoder to synthesize a predicted high sampling rate audio waveform.

[0149] Through the collaborative work of multiple modules, the entire process from noise suppression to high-fidelity speech generation is optimized, reducing complex inter-module interactions and data transfer processes, lowering system complexity and construction costs, and minimizing the impact of inter-module error propagation on the final synthesis effect. This improves audio synthesis efficiency and quality.

[0150] In some optional implementations of this embodiment, the noise reduction submodule is further used for:

[0151] The noisy low-sampling-rate audio sample set is input into the noise reduction sub-model, wherein the noise reduction sub-model includes a feature input layer, a feature extraction layer, and a recurrent layer;

[0152] The noisy low-sampling-rate audio sample set is input into the feature input layer for audio conversion to obtain noisy low-sampling-rate spectral features;

[0153] Local audio enhancement features are obtained by extracting local features from the noisy low sampling rate spectral features through the feature extraction layer.

[0154] By analyzing the temporal dynamic characteristics of the local audio enhancement features through the loop layer, the noise reduction frequency features are obtained.

[0155] By processing noisy, low-sampling-rate audio sample sets, audio and noise can be accurately separated, thereby generating high-quality noise-reduced audio features under different noise environments and improving the accuracy of effective speech feature extraction.

[0156] In some optional implementations of this embodiment, the matching submodule is further used for:

[0157] The noise reduction frequency features and the input text are input into the speech synthesis sub-model, wherein the speech synthesis sub-model includes a text feature layer, a prosodic feature layer and a speech synthesis layer;

[0158] The text feature layer extracts features from the input text to obtain a semantic feature vector;

[0159] The prosodic features of the noise-reduced frequency features are extracted through the prosodic feature layer to obtain the prosodic feature vector;

[0160] The semantic feature vector and the prosodic feature vector are input into the speech synthesis layer, and the semantic feature vector and the prosodic feature vector are fused based on the stream matching algorithm to generate medium sampling rate synthesized audio features.

[0161] By fusing semantic and prosodic information, the naturalness of synthesized speech can be enhanced, the mechanical feel of synthesized speech can be reduced, and the user experience can be improved.

[0162] In some optional implementations of this embodiment, the supramolecular module is further used for:

[0163] The medium sampling rate synthesized audio features are input into the supramolecular model, wherein the supramolecular model includes a high-frequency reconstruction layer and a decoding output layer;

[0164] Through the high-frequency reconstruction layer, a frequency band-aware attention mechanism is used to extract and reconstruct the high-frequency features in the synthesized audio features to obtain enhanced high-frequency features;

[0165] The enhanced high-frequency features are decoded and generated through the decoding output layer to obtain high-sampling-rate synthesized audio features.

[0166] By introducing a frequency band awareness attention mechanism, it is possible to more accurately identify the frequency band regions that need enhancement, improve the generation effect of high-frequency components, ensure the accuracy of high-frequency reconstruction, and thus improve the clarity of synthesized speech.

[0167] In some optional implementations of this embodiment, the synthesis submodule is further used for:

[0168] The high-sampling-rate synthesized audio features are input into the vocoder, which includes an input layer, a feature extraction layer, a convolutional layer, a prediction layer, a speech reconstruction layer, and a speech output layer.

[0169] The high-sampling-rate synthesized audio features are preprocessed through the input layer to obtain high-frequency input features;

[0170] The high-frequency input features are input into the feature extraction layer for feature extraction to obtain high-frequency spectral features;

[0171] The high-frequency spectral features are fused by channel convolution through the convolutional layer to obtain fused spectral features;

[0172] The fused spectral features are input into the prediction layer for prediction to obtain the audio temporal spectrum.

[0173] The audio time-spectrum is restored to an audio signal through the speech restoration layer;

[0174] The audio signal is subjected to audio waveform prediction and synthesis through the output layer, and a predicted high sampling rate audio waveform is output.

[0175] By converting high-sampling-rate synthesized audio features into audio and outputting the final synthesized audio waveform, clear and natural audio can be generated, effectively improving the quality and realism of the synthesized audio.

[0176] In some alternative implementations, the loss calculation module 404 is used for:

[0177] Extract low-sampling-rate reference audio features from the low-sampling-rate reference audio set, and calculate a first loss function based on the low-sampling-rate reference audio features;

[0178] Extract the mid-sampling rate reference audio features from the mid-sampling rate reference audio set, and calculate the second loss function based on the mid-sampling rate reference audio features;

[0179] Extract the high sampling rate reference audio features from the high sampling rate reference audio set, and calculate the third loss function based on the high sampling rate reference audio features;

[0180] The fourth loss function is calculated based on the predicted high-sampling-rate audio waveform and the corresponding high-sampling-rate reference audio in the high-sampling-rate reference audio set.

[0181] The first loss function, the second loss function, the third loss function, and the fourth loss function are weighted and calculated to obtain the total loss function.

[0182] By optimizing the model using multiple loss functions, we can improve the model's prediction accuracy, enhance its robustness, promote its generalization ability, and improve the naturalness of synthesized speech, thereby achieving high-quality speech synthesis.

[0183] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference] for details. Figure 5 , Figure 5 This is a basic structural block diagram of the computer device in this embodiment.

[0184] The computer device 5 includes a memory 51, a processor 52, and a network interface 53 that are interconnected via a system bus. It should be noted that only a computer device 5 with a memory 51, a processor 52, and a network interface 53 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0185] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0186] The memory 51 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 51 may be an internal storage unit of the computer device 5, such as the hard disk or memory of the computer device 5. In other embodiments, the memory 51 may also be an external storage device of the computer device 5, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 5. Of course, the memory 51 may include both the internal storage unit and its external storage device of the computer device 5. In this embodiment, the memory 51 is typically used to store the operating system and various application software installed on the computer device 5, such as computer-readable instructions for end-to-end speech synthesis methods. In addition, the memory 51 can also be used to temporarily store various types of data that have been output or will be output.

[0187] In some embodiments, the processor 52 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 52 is typically used to control the overall operation of the computer device 5. In this embodiment, the processor 52 is used to execute computer-readable instructions stored in the memory 51 or to process data, for example, to execute computer-readable instructions of the end-to-end speech synthesis method.

[0188] The network interface 53 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 5 and other electronic devices.

[0189] By training the pre-trained speech synthesis model using a multimodal sample set, the final end-to-end speech synthesis model is obtained. During the training process, the model parameters of the speech synthesis model are dynamically adjusted through the weighted balancing of multiple loss functions, which improves the model's ability to extract effective speech features from noisy low-sampling-rate speech. At the same time, it can effectively convert low-sampling-rate speech into high-sampling-rate speech, ensuring high-quality speech synthesis and improving the naturalness and clarity of the synthesized speech.

[0190] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the end-to-end speech synthesis method as described above.

[0191] By training the pre-trained speech synthesis model using a multimodal sample set, the final end-to-end speech synthesis model is obtained. During the training process, the model parameters of the speech synthesis model are dynamically adjusted through the weighted balancing of multiple loss functions, which improves the model's ability to extract effective speech features from noisy low-sampling-rate speech. At the same time, it can effectively convert low-sampling-rate speech into high-sampling-rate speech, ensuring high-quality speech synthesis and improving the naturalness and clarity of the synthesized speech.

[0192] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0193] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

[0194] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.

Claims

1. An end-to-end speech synthesis method, characterized in that, Includes the following steps: Obtain a high sampling rate reference audio set and the corresponding input text, and downsample the high sampling rate reference audio set to obtain a low sampling rate reference audio set and a medium sampling rate reference audio set; Noise with different signal-to-noise ratios is injected into the low-sampling-rate reference audio set to obtain a noisy low-sampling-rate audio sample set. The noisy low-sampling-rate audio sample set and the corresponding input text are then combined to generate a multimodal sample set. The multimodal sample set is input into a pre-trained speech synthesis model. The speech synthesis model performs noise reduction, prosody matching, and super-resolution processing on the multimodal sample set to obtain a predicted high sampling rate audio waveform. The speech synthesis model includes a noise reduction sub-model, a speech synthesis sub-model, a supramolecular model, and a vocoder. According to the preset loss function, the weighted loss is calculated based on the predicted high sampling rate audio waveform, the low sampling rate reference audio set, the medium sampling rate reference audio set, and the high sampling rate reference audio set to obtain the total loss function; The model parameters of the speech synthesis model are adjusted based on the total loss function, and iterative training continues until the iteration stopping condition is met to obtain the final end-to-end speech synthesis model. Obtain the text to be synthesized and the corresponding noisy low-sampling-rate reference speech, input the text to be synthesized and the noisy low-sampling-rate reference speech into the end-to-end speech synthesis model, and synthesize the target speech waveform; The step of performing noise reduction, prosody matching, and super-resolution processing on the multimodal sample set using the speech synthesis model to obtain a predicted high-sampling-rate audio waveform includes: processing the noisy low-sampling-rate audio sample set using the noise reduction sub-model to generate noise-reduced audio features; generating medium-sampling-rate synthesized audio features using the speech synthesis sub-model based on the noise-reduced audio features and the input text; performing super-resolution calculation on the medium-sampling-rate synthesized audio features using the supramolecular model to obtain high-sampling-rate synthesized audio features; and performing audio conversion on the high-sampling-rate synthesized audio features using the vocoder to synthesize a predicted high-sampling-rate audio waveform. The step of generating medium-sampling-rate synthesized audio features based on the noise-reduced frequency features and the input text through the speech synthesis sub-model includes: inputting the noise-reduced frequency features and the input text into the speech synthesis sub-model, wherein the speech synthesis sub-model includes a text feature layer, a prosodic feature layer, and a speech synthesis layer; extracting features from the input text through the text feature layer to obtain a semantic feature vector; extracting prosodic features from the noise-reduced frequency features through the prosodic feature layer to obtain a prosodic feature vector; inputting the semantic feature vector and the prosodic feature vector into the speech synthesis layer, and fusing the semantic feature vector and the prosodic feature vector based on a stream matching algorithm to generate medium-sampling-rate synthesized audio features.

2. The end-to-end speech synthesis method according to claim 1, characterized in that, The step of processing the noisy low-sampling-rate audio sample set using the noise reduction sub-model to generate noise-reduced audio features includes: The noisy low-sampling-rate audio sample set is input into the noise reduction sub-model, wherein the noise reduction sub-model includes a feature input layer, a feature extraction layer, and a recurrent layer; The noisy low-sampling-rate audio sample set is input into the feature input layer for audio conversion to obtain noisy low-sampling-rate spectral features; Local audio enhancement features are obtained by extracting local features from the noisy low sampling rate spectral features through the feature extraction layer. By analyzing the temporal dynamic characteristics of the local audio enhancement features through the loop layer, the noise reduction frequency features are obtained.

3. The end-to-end speech synthesis method according to claim 1, characterized in that, The step of performing super-resolution calculation on the medium-sampling-rate synthesized audio features using the supramolecular model to obtain high-sampling-rate synthesized audio features includes: The medium sampling rate synthesized audio features are input into the supramolecular model, wherein the supramolecular model includes a high-frequency reconstruction layer and a decoding output layer; Through the high-frequency reconstruction layer, a frequency band-aware attention mechanism is used to extract and reconstruct the high-frequency features in the medium sampling rate synthesized audio features to obtain enhanced high-frequency features; The enhanced high-frequency features are decoded and generated through the decoding output layer to obtain high-sampling-rate synthesized audio features.

4. The end-to-end speech synthesis method according to claim 1, characterized in that, The step of converting the high-sampling-rate synthesized audio features using the vocoder to synthesize a predicted high-sampling-rate audio waveform includes: The high-sampling-rate synthesized audio features are input into the vocoder, which includes an input layer, a feature extraction layer, a convolutional layer, a prediction layer, a speech reconstruction layer, and a speech output layer. The high-sampling-rate synthesized audio features are preprocessed through the input layer to obtain high-frequency input features; The high-frequency input features are input into the feature extraction layer for feature extraction to obtain high-frequency spectral features; The high-frequency spectral features are fused by channel convolution through the convolutional layer to obtain fused spectral features; The fused spectral features are input into the prediction layer for prediction to obtain the audio temporal spectrum. The audio time-spectrum is restored to an audio signal through the speech restoration layer; The audio signal is subjected to audio waveform prediction and synthesis through the output layer, and a predicted high sampling rate audio waveform is output.

5. The end-to-end speech synthesis method according to claim 1, characterized in that, The step of calculating a weighted loss based on the predicted high-sampling-rate audio waveform, the low-sampling-rate reference audio set, the medium-sampling-rate reference audio set, and the high-sampling-rate reference audio set, according to a preset loss function, to obtain the total loss function includes: Extract low-sampling-rate reference audio features from the low-sampling-rate reference audio set, and calculate a first loss function based on the low-sampling-rate reference audio features; Extract the mid-sampling rate reference audio features from the mid-sampling rate reference audio set, and calculate the second loss function based on the mid-sampling rate reference audio features; Extract the high sampling rate reference audio features from the high sampling rate reference audio set, and calculate the third loss function based on the high sampling rate reference audio features; The fourth loss function is calculated based on the predicted high-sampling-rate audio waveform and the corresponding high-sampling-rate reference audio in the high-sampling-rate reference audio set. The first loss function, the second loss function, the third loss function, and the fourth loss function are weighted and calculated to obtain the total loss function.

6. An end-to-end speech synthesis device, characterized in that, include: The acquisition module is used to acquire a high sampling rate reference audio set and a corresponding input text, and to downsample the high sampling rate reference audio set to obtain a low sampling rate reference audio set and a medium sampling rate reference audio set. The module is used to inject noise with different signal-to-noise ratios into the low sampling rate reference audio set to obtain a noisy low sampling rate audio sample set, and to combine the noisy low sampling rate audio sample set and the corresponding input text to generate a multimodal sample set. The training module is used to input the multimodal sample set into the pre-trained speech synthesis model, and to perform noise reduction, prosody matching and super-resolution processing on the multimodal sample set through the speech synthesis model to obtain a predicted high sampling rate audio waveform. The speech synthesis model includes a noise reduction sub-model, a speech synthesis sub-model, a supramolecular model and a vocoder. The loss calculation module is used to calculate the weighted loss according to the predicted high sampling rate audio waveform, the low sampling rate reference audio set, the medium sampling rate reference audio set, and the high sampling rate reference audio set, and obtain the total loss function according to the preset loss function; An iterative module is used to adjust the model parameters of the speech synthesis model based on the total loss function, continue iterative training until the iteration stopping condition is met, and obtain the final end-to-end speech synthesis model. The synthesis module is used to acquire the text to be synthesized and the corresponding noisy low-sampling-rate reference speech, input the text to be synthesized and the noisy low-sampling-rate reference speech into the end-to-end speech synthesis model, and synthesize the target speech waveform; The training module includes: a noise reduction submodule, used to process the noisy low-sampling-rate audio sample set through the noise reduction submodel to generate noise-reduced audio features; a matching submodule, used to generate medium-sampling-rate synthesized audio features based on the noise-reduced audio features and the input text through the speech synthesis submodel; a supramolecular module, used to perform super-resolution calculation on the medium-sampling-rate synthesized audio features through the supramolecular model to obtain high-sampling-rate synthesized audio features; and a synthesis submodule, used to perform audio conversion on the high-sampling-rate synthesized audio features through the vocoder to synthesize a predicted high-sampling-rate audio waveform. The matching submodule is further configured to: input the noise-reduced frequency features and the input text into the speech synthesis submodel, wherein the speech synthesis submodel includes a text feature layer, a prosodic feature layer, and a speech synthesis layer; extract features from the input text through the text feature layer to obtain a semantic feature vector; extract the prosodic features of the noise-reduced frequency features through the prosodic feature layer to obtain a prosodic feature vector; input the semantic feature vector and the prosodic feature vector into the speech synthesis layer, and fuse the semantic feature vector and the prosodic feature vector based on a stream matching algorithm to generate medium sampling rate synthesized audio features.

7. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the end-to-end speech synthesis method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the end-to-end speech synthesis method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Speech synthesis method, model training method, equipment and storage medium

    CN114283783A

  • Text-guided speech synthesis method and device, computer equipment and storage medium

    CN120015011A