Voice synthesis method combining variational inference and rhythm perception and related equipment

By combining variational inference and rhythm-aware speech synthesis methods, the difficulties of traditional TTS systems in terms of intonation and rhythm are solved, achieving a more natural and fluent speech synthesis effect.

CN121938348APending Publication Date: 2026-04-28SHANGHAI JITU SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI JITU SCI & TECH CO LTD
Filing Date
2024-02-20
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional text-to-speech systems have difficulties with intonation and rhythm, resulting in synthesized speech that is not natural or fluent.

Method used

Combining variational inference and rhythm-aware speech synthesis methods, this paper uses a variational inference text-to-speech model to perform pinyin conversion, phoneme extraction, and audio feature synthesis, and utilizes a rhythm-aware speech conversion model to capture prosody and rhythm and incorporate timbre, thereby achieving natural and fluent speech synthesis.

Benefits of technology

It improves the naturalness of intonation and rhythm during text-to-speech conversion, resulting in more natural and fluent speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121938348A_ABST
    Figure CN121938348A_ABST
Patent Text Reader

Abstract

The invention provides a variational inference and rhythm perception combined speech synthesis method and related equipment. The method comprises the steps of receiving a to-be-converted speech text input by a user; inputting the to-be-converted voice text into a variational inference text-to-voice model; performing pinyin conversion, phoneme extraction and audio feature synthesis on the to-be-converted voice text through the variational inference text-to-voice model; inputting the audio features into a rhythm perception voice conversion model; and performing rhythm capture on the audio features through the rhythm perception voice conversion model, and fusing timbres to obtain a target timbre audio. The variational inference speech synthesis technology and the rhythm perception speech synthesis technology are utilized in the process of converting the text into the speech, and the intonation rhythm of the text synthesis speech is more natural and smoother.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech synthesis technology, and in particular to a speech synthesis method and related equipment that combines variational inference and rhythm perception. Background Technology

[0002] Text-to-speech (TTS) is a technology that intelligently converts text into natural speech. A TTS system converts text into synthesized speech that is as close as possible to real human speech, according to the pronunciation rules of a specific language. Traditional TTS systems often face difficulties in intonation, rhythm, and other aspects during the speech synthesis process, resulting in synthesized speech that is not natural or fluent.

[0003] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0004] This invention provides a speech synthesis method and related equipment that combines variational inference and rhythm awareness. The main objective of this invention is to solve the technical problems mentioned in the background section of the prior art.

[0005] The first aspect of this invention provides a speech synthesis method combining variational inference and rhythm awareness, comprising:

[0006] Receive user-inputted voice-to-text;

[0007] The text to be converted into speech is input into the variational inference text-to-speech model;

[0008] The variational inference text-to-speech model is used to perform pinyin conversion, phoneme extraction, and audio feature synthesis on the text to be converted.

[0009] The audio features are input into the rhythm-aware speech conversion model;

[0010] The rhythm-aware speech conversion model captures the rhythm and timbre of the audio features and integrates them to obtain the target timbre audio.

[0011] In an optional embodiment of the first aspect of the present invention, the step of performing pinyin conversion, phoneme extraction, and audio feature synthesis on the text to be converted using the variational inference text-to-speech model includes:

[0012] The speech text to be converted is converted into a Pinyin set using the PyPinyin library;

[0013] The pinyin set is split into initials and finals to obtain a phoneme set;

[0014] The phoneme set is subjected to audio embedding processing to obtain an audio set;

[0015] The audio set is encoded using a variational autoencoder to obtain audio features.

[0016] In an optional embodiment of the first aspect of the present invention, encoding the audio set using a variational autoencoder to obtain audio features includes:

[0017] Extract an acoustic feature set from the audio set;

[0018] The acoustic feature set is normalized to obtain a stable acoustic feature set;

[0019] The stable acoustic feature set is mapped to the latent space to obtain a latent vector dataset;

[0020] The latent vector dataset is input into an adversarial generative network to obtain audio features expressed in array form.

[0021] In an optional embodiment of the first aspect of the present invention, the step of capturing the rhythm and rhythm of the audio features and integrating them with timbre through the rhythm-aware speech conversion model to obtain the target timbre audio includes:

[0022] The audio features are converted into audio waveforms;

[0023] The audio waveform is subjected to prosody and rhythm capture to obtain prosodic feature points and rhythmic feature points;

[0024] Based on the prosodic feature points and the rhythmic feature points, the audio waveform is fused to obtain the target timbre audio.

[0025] In an optional embodiment of the first aspect of the present invention, the step of capturing the prosody and rhythm of the audio waveform to obtain prosodic feature points and rhythmic feature points includes:

[0026] The fundamental frequency and speech rate changes in the audio waveform are analyzed, and temporal features are introduced for reference to obtain the prosodic feature points in the audio waveform.

[0027] The syllable levels in the audio waveform are analyzed, and rhythmic feature points in the audio waveform are obtained based on the degree of the syllable levels.

[0028] A second aspect of the present invention provides a speech synthesis apparatus combining variational inference and rhythm awareness, the speech synthesis apparatus comprising:

[0029] The text receiving module is used to receive user-inputted text to be converted into speech;

[0030] The text-to-speech processing module is used to input the text to be converted into the variational inference text-to-speech model;

[0031] The variational inference processing module is used to perform pinyin conversion, phoneme extraction, and audio feature synthesis on the text to be converted to speech using the variational inference text-to-speech model.

[0032] An audio feature input module is used to input the audio features into a rhythm-aware speech conversion model;

[0033] The rhythm-aware processing module is used to capture the rhythm and timbre of the audio features through the rhythm-aware speech conversion model and integrate them to obtain the target timbre audio.

[0034] In an optional embodiment of the second aspect of the present invention, the variational inference processing module includes:

[0035] The Pinyin conversion unit is used to convert the speech text to be converted into a Pinyin set using the PyPinyin library;

[0036] A phoneme conversion unit is used to split the pinyin set into initials and finals to obtain a phoneme set;

[0037] An audio conversion unit is used to perform audio embedding processing on the phoneme set to obtain an audio set;

[0038] The variational automatic coding unit is used to encode the audio set using a variational automatic encoder to obtain audio features.

[0039] In an optional embodiment of the second aspect of the present invention, the rhythm perception processing module includes:

[0040] A waveform conversion unit is used to convert the audio features into an audio waveform.

[0041] The prosody and rhythm capture unit is used to capture the prosody and rhythm of the audio waveform to obtain prosodic feature points and rhythmic feature points.

[0042] The timbre fusion unit is used to perform timbre fusion on the audio waveform based on the prosodic feature points and the rhythmic feature points to obtain the target timbre audio.

[0043] A third aspect of the present invention provides a speech synthesis device combining variational inference and rhythm awareness, the speech synthesis device comprising: a memory and at least one processor, the memory storing instructions, the memory and the at least one processor being interconnected via a circuit;

[0044] The at least one processor invokes the instructions in the memory to cause the speech synthesis device combining variational inference and rhythm awareness to perform the speech synthesis method combining variational inference and rhythm awareness as described above.

[0045] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the speech synthesis method combining variational inference and rhythm awareness as described in any of the preceding claims.

[0046] Beneficial Effects: This invention provides a speech synthesis method and related equipment combining variational inference and rhythm-aware techniques. The method includes receiving user-inputted text to be converted into speech; inputting the text into a variational inference text-to-speech model; performing phonetic conversion, phoneme extraction, and audio feature synthesis on the text using the variational inference text-to-speech model; inputting the audio features into a rhythm-aware speech conversion model; and capturing the prosody and rhythm of the audio features and integrating timbre to obtain a target timbre audio. This invention utilizes variational inference speech synthesis technology and rhythm-aware speech synthesis technology in the text-to-speech process, resulting in more natural and fluent intonation and rhythm in the synthesized speech. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of an embodiment of a speech synthesis method combining variational inference and rhythm awareness according to the present invention;

[0048] Figure 2 This is a schematic diagram of an embodiment of a speech synthesis device combining variational inference and rhythm awareness according to the present invention;

[0049] Figure 3 This is a schematic diagram of an embodiment of a speech synthesis device that combines variational inference and rhythm awareness according to the present invention. Detailed Implementation

[0050] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0051] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1The first aspect of this invention provides a speech synthesis method combining variational inference and rhythm awareness, comprising:

[0052] S100: Receive the user-inputted voice text to be converted;

[0053] S200. Input the text to be converted into a variational inference text-to-speech model. In this invention, the variational inference text-to-speech model uses the VITS (Variational Inference-based Text-to-Speech) speech conversion model. The VITS model, as a variational inference-based speech synthesis method, can generate high-quality speech by introducing a variational autoencoder and a generative model.

[0054] S300. The variational inference text-to-speech model is used to convert the text to be converted into pinyin, extract phonemes, and synthesize audio features. In this invention, the text to be converted into pinyin is input into the VITS text-to-speech model, and it is first converted into pinyin, then the pinyin is converted into phonemes, and then audio is synthesized based on the phonemes.

[0055] S400. Input the audio features into the rhythm-aware speech conversion model; In this invention, the rhythm-aware speech conversion model uses the URhythmic (Unsupervised Rhythm-based TTS) model. The URhythmic model uses rhythm-aware voice-conversion technology based on unsupervised learning, which can automatically learn and capture the rhythm features of speech from the input speech audio.

[0056] S500: The rhythm-aware speech conversion model is used to capture the prosodic rhythm of the audio features and incorporate timbre to obtain the target timbre audio. In this invention, by utilizing the variational inference technology of the VITS speech conversion model and the rhythm-aware technology of the URhythmic speech conversion model, the audio synthesized from text can be more natural and fluent.

[0057] In an optional embodiment of the first aspect of the present invention, the step of performing pinyin conversion, phoneme extraction, and audio feature synthesis on the text to be converted using the variational inference text-to-speech model includes:

[0058] The text to be converted into speech is then converted into a Pinyin set using the PyPinyin library. PyPinyin is a Python library for converting Chinese characters into Pinyin. It provides various Pinyin styles (such as initial consonant style, style with tone marks, style with separators, etc.), supports polyphonic characters, and allows for the setting of various default processing methods.

[0059] The pinyin set is split into initials and finals to obtain a phoneme set; in this invention, the initials and finals of each pinyin in the pinyin set are separated so that subsequent steps can better obtain the smallest unit of audio pronunciation (i.e., the pronunciation of a phoneme).

[0060] The phoneme set is subjected to audio embedding processing to obtain an audio set; in this step, the pronunciations of each initial consonant and final vowel are then integrated in sequence to obtain the audio set.

[0061] The audio set is encoded using a variational autoencoder to obtain audio features. In this invention, the obtained audio set is further encoded using a variational autoencoder to obtain higher quality audio data.

[0062] In an optional embodiment of the first aspect of the present invention, encoding the audio set using a variational autoencoder to obtain audio features includes:

[0063] Extracting acoustic feature sets from the audio set; this typically involves using Short-Time Fourier Transform (STFT) or Mel filter banks to capture the spectral information of the audio;

[0064] The acoustic feature set is normalized to obtain a stable acoustic feature set. In this step, the acoustic features can have zero mean and unit variance after normalization, which improves the stability of the acoustic features.

[0065] The stable acoustic feature set is mapped to the latent space to obtain the latent vector dataset; in this step, the acoustic features are expressed in a higher-dimensional vector form.

[0066] The latent vector dataset is input into an adversarial generative network to obtain audio features expressed in array form. In this step, the adversarial generative network processes the latent vector dataset using linear transformations, nonlinear transformations, residual connections, and attention mechanisms.

[0067] In an optional embodiment of the first aspect of the present invention, the step of capturing the rhythm and rhythm of the audio features and integrating them with timbre through the rhythm-aware speech conversion model to obtain the target timbre audio includes:

[0068] The audio features are converted into audio waveforms. In this step, the rhythm-aware speech conversion model converts the audio features into waveforms to quickly capture prosody and rhythm in the model graph.

[0069] The audio waveform is subjected to prosody and rhythm capture to obtain prosodic feature points and rhythmic feature points. In this step, the capture of prosody and rhythm is mainly based on the characteristics of the waveform formed by prosody and rhythm, such as fundamental frequency changes, speech rate changes and syllable levels.

[0070] Based on the prosodic feature points and rhythmic feature points, the audio waveform is timbre fused to obtain the target timbre audio. This step mainly involves determining the location of the timbre transformation based on the prosodic and rhythmic feature points. The timbre encoding primarily utilizes the vocoder in the rhythm-aware speech conversion model. By locating the prosody and rhythm, the timbre fusion can be more natural and have a higher matching degree.

[0071] In an optional embodiment of the first aspect of the present invention, the step of capturing the prosody and rhythm of the audio waveform to obtain prosodic feature points and rhythmic feature points includes:

[0072] The fundamental frequency and speech rate changes in the audio waveform are analyzed, and temporal features are introduced for reference to obtain prosodic feature points in the audio waveform. In this step, the fluctuations of the fundamental frequency and speech rate are mainly analyzed. When the fluctuation exceeds a certain threshold, it can be preliminarily identified as a prosodic feature point. Temporal features are mainly used to exclude some non-positive prosodic feature points.

[0073] The syllable levels in the audio waveform are analyzed, and rhythmic feature points are obtained based on the degree of the syllable level. In this invention, the syllable level includes intensity, stress, and syllable duration, etc. When a certain syllable condition is met, it is determined to be a rhythmic feature point.

[0074] In summary, the synthesized speech and voice synthesis system proposed in this patent has the following innovations:

[0075] (1) It combines the advantages of VITS and URhythmic, providing a more natural and fluent speech synthesis and sound simulation effect.

[0076] (2) By utilizing VITS’s pinyin embedding and audio feature synthesis functions, as well as URhythmic’s rhythm perception capabilities, the content and prosodic features of the input text were fully captured.

[0077] (3) The second half of the timbre transfer section adopts an unsupervised learning approach, which reduces the dependence on labeled data and improves the scalability and adaptability of the system.

[0078] See Figure 2 A second aspect of the present invention provides a speech synthesis apparatus combining variational inference and rhythm awareness, the speech synthesis apparatus comprising:

[0079] The text receiving module 10 is used to receive the text to be converted into speech input by the user;

[0080] Text-to-speech processing module 20 is used to input the speech text to be converted into the variational inference text-to-speech model;

[0081] The variational inference processing module 30 is used to perform pinyin conversion, phoneme extraction and audio feature synthesis on the text to be converted to speech using the variational inference text-to-speech model.

[0082] Audio feature input module 40 is used to input the audio features into the rhythm-aware speech conversion model;

[0083] The rhythm perception processing module 50 is used to capture the rhythm and timbre of the audio features through the rhythm perception speech conversion model and integrate them to obtain the target timbre audio.

[0084] In an optional embodiment of the second aspect of the present invention, the variational inference processing module includes:

[0085] The Pinyin conversion unit is used to convert the speech text to be converted into a Pinyin set using the PyPinyin library;

[0086] A phoneme conversion unit is used to split the pinyin set into initials and finals to obtain a phoneme set;

[0087] An audio conversion unit is used to perform audio embedding processing on the phoneme set to obtain an audio set;

[0088] The variational automatic coding unit is used to encode the audio set using a variational automatic encoder to obtain audio features.

[0089] In an optional embodiment of the second aspect of the present invention, the variational automatic coding unit includes:

[0090] An acoustic feature extraction subunit is used to extract an acoustic feature set from the audio set;

[0091] The normalization processing subunit is used to normalize the acoustic feature set to obtain a stable acoustic feature set.

[0092] The feature mapping subunit is used to map the stable acoustic feature set to the latent space to obtain a latent vector dataset.

[0093] An array processing subunit is used to input the latent vector dataset into an adversarial generative network to obtain audio features expressed in array form.

[0094] In an optional embodiment of the second aspect of the present invention, the rhythm perception processing module includes:

[0095] A waveform conversion unit is used to convert the audio features into an audio waveform.

[0096] The prosody and rhythm capture unit is used to capture the prosody and rhythm of the audio waveform to obtain prosodic feature points and rhythmic feature points.

[0097] The timbre fusion unit is used to perform timbre fusion on the audio waveform based on the prosodic feature points and the rhythmic feature points to obtain the target timbre audio.

[0098] In an optional embodiment of a second aspect of the invention, the rhythmic capture unit includes:

[0099] The prosodic feature point capture subunit is used to analyze the fundamental frequency change and speech rate change in the audio waveform diagram, and introduce temporal features for reference to obtain the prosodic feature points in the audio waveform diagram;

[0100] The rhythm feature point capture subunit is used to analyze the syllable level in the audio waveform diagram and obtain the rhythm feature points in the audio waveform diagram based on the degree of the syllable level.

[0101] Figure 3 This is a schematic diagram of a speech synthesis device combining variational inference and rhythm awareness, provided by an embodiment of the present invention. Speech synthesis devices combining variational inference and rhythm awareness can vary considerably due to different configurations or performance characteristics. They may include one or more processors 70 (central processing units, CPUs) (e.g., one or more processors) and memory 80, and one or more storage media 90 (e.g., one or more mass storage devices) for storing application programs or data. The memory and storage media can be temporary or persistent storage. The program stored in the storage media may include one or more modules (not shown in the diagram), each module including a series of instruction operations within the speech synthesis device combining variational inference and rhythm awareness. Furthermore, the processor may be configured to communicate with the storage media and execute the series of instruction operations in the storage media on the speech synthesis device combining variational inference and rhythm awareness.

[0102] The speech synthesis device of this invention, combining variational inference and rhythm awareness, may also include one or more power supplies 100, one or more wired or wireless network interfaces 110, one or more input / output interfaces 120, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 3The illustrated speech synthesis device structure combining variational inference and rhythm awareness does not constitute a specific limitation on the speech synthesis device combining variational inference and rhythm awareness. It may also include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0103] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the speech synthesis method combining variational inference and rhythm awareness.

[0104] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system or system / unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0105] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0106] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A speech synthesis method combining variational inference and rhythm awareness, characterized in that, include: Receive user-inputted voice-to-text; The text to be converted into speech is input into the variational inference text-to-speech model; The variational inference text-to-speech model is used to perform pinyin conversion, phoneme extraction, and audio feature synthesis on the text to be converted. The audio features are input into the rhythm-aware speech conversion model; The rhythm-aware speech conversion model captures the rhythm and timbre of the audio features and integrates them to obtain the target timbre audio.

2. The speech synthesis method combining variational inference and rhythm awareness according to claim 1, characterized in that, The step of performing pinyin conversion, phoneme extraction, and audio feature synthesis on the text to be converted using the variational inference text-to-speech model includes: The speech text to be converted is converted into a Pinyin set using the PyPinyin library; The pinyin set is split into initials and finals to obtain a phoneme set; The phoneme set is subjected to audio embedding processing to obtain an audio set; The audio set is encoded using a variational autoencoder to obtain audio features.

3. The speech synthesis method combining variational inference and rhythm awareness according to claim 2, characterized in that, The process of encoding the audio set using a variational autoencoder to obtain audio features includes: Extract an acoustic feature set from the audio set; The acoustic feature set is normalized to obtain a stable acoustic feature set; The stable acoustic feature set is mapped to the latent space to obtain a latent vector dataset; The latent vector dataset is input into an adversarial generative network to obtain audio features expressed in array form.

4. The speech synthesis method combining variational inference and rhythm awareness according to claim 1, characterized in that, The step of capturing the rhythm and rhythm of the audio features and integrating them with timbre through the rhythm-aware speech conversion model to obtain the target timbre audio includes: The audio features are converted into audio waveforms; The audio waveform is subjected to prosody and rhythm capture to obtain prosodic feature points and rhythmic feature points; Based on the prosodic feature points and the rhythmic feature points, the audio waveform is fused to obtain the target timbre audio.

5. The speech synthesis method combining variational inference and rhythm awareness according to claim 4, characterized in that, The step of capturing prosody and rhythm from the audio waveform to obtain prosodic and rhythmic feature points includes: The fundamental frequency and speech rate changes in the audio waveform are analyzed, and temporal features are introduced for reference to obtain the prosodic feature points in the audio waveform. The syllable levels in the audio waveform are analyzed, and rhythmic feature points in the audio waveform are obtained based on the degree of the syllable levels.

6. A speech synthesis device combining variational inference and rhythm perception, characterized in that, The speech synthesis device combining variational inference and rhythm awareness includes: The text receiving module is used to receive user-inputted text to be converted into speech; The text-to-speech processing module is used to input the text to be converted into the variational inference text-to-speech model; The variational inference processing module is used to perform pinyin conversion, phoneme extraction, and audio feature synthesis on the text to be converted to speech using the variational inference text-to-speech model. An audio feature input module is used to input the audio features into a rhythm-aware speech conversion model; The rhythm-aware processing module is used to capture the rhythm and timbre of the audio features through the rhythm-aware speech conversion model and integrate them to obtain the target timbre audio.

7. The speech synthesis device combining variational inference and rhythm perception according to claim 6, characterized in that, The variational inference processing module includes: The Pinyin conversion unit is used to convert the speech text to be converted into a Pinyin set using the PyPinyin library; A phoneme conversion unit is used to split the pinyin set into initials and finals to obtain a phoneme set; An audio conversion unit is used to perform audio embedding processing on the phoneme set to obtain an audio set; The variational automatic coding unit is used to encode the audio set using a variational automatic encoder to obtain audio features.

8. The speech synthesis apparatus combining variational inference and rhythm perception according to claim 6, characterized in that, The rhythm perception processing module includes: A waveform conversion unit is used to convert the audio features into an audio waveform. The prosody and rhythm capture unit is used to capture the prosody and rhythm of the audio waveform to obtain prosodic feature points and rhythmic feature points. The timbre fusion unit is used to perform timbre fusion on the audio waveform based on the prosodic feature points and the rhythmic feature points to obtain the target timbre audio.

9. A speech synthesis device combining variational inference and rhythm awareness, characterized in that, The speech synthesis device combining variational inference and rhythm awareness includes: a memory and at least one processor, wherein the memory stores instructions and the memory and the at least one processor are interconnected via a line. The at least one processor invokes the instructions in the memory to cause the speech synthesis device combining variational inference and rhythm awareness to perform the speech synthesis method combining variational inference and rhythm awareness as described in any one of claims 1-5.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the speech synthesis method combining variational inference and rhythm awareness as described in any one of claims 1-5.