Duration-Aware Network for Text-to-Speech Conversion Analysis

By introducing the duration model and CBHG module in the Tacotron system, the problem of text skipping and repetition in the end-to-end speech synthesis system is solved, and a more natural and stable speech synthesis effect is achieved.

CN113711305BActive Publication Date: 2025-07-04TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080028696.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-29
Filing Date
2020-03-05
Publication Date
2025-07-04
Estimated Expiration
2040-03-05

AI Technical Summary

Technical Problem

The end-to-end speech synthesis system based on Tacotron is prone to skipping or repeating text input when synthesizing speech, and the attention mechanism is uncontrollable, resulting in unstable synthetic speech.

Method used

The duration-based attention mechanism is adopted to predict the duration of input characters and phonemes through the duration model to ensure the alignment of the input text with the spectral frame, and use the CBHG module to generate more accurate spectral frames and synthesize high-quality speech.

Benefits of technology

Improves the naturalness and stability of speech synthesis, avoids text skipping and repetition, and generates more natural and accurate speech output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113711305B_ABST
    Figure CN113711305B_ABST
Patent Text Reader

Abstract

A method and apparatus, comprising: receiving a text input including a sequence of text components; using a duration model to determine corresponding durations of the text components; generating a first spectrogram set based on the sequence of text components; generating a second spectrogram set based on the first spectrogram set and the corresponding durations of the sequence of text components; generating spectrogram frames based on the second spectrogram set; generating an audio waveform based on the spectrogram frames; and providing the audio waveform as an output.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Application No. 16 / 397,349, filed on Apr. 29, 2019, the disclosure of which is hereby incorporated by reference in its entirety. Background Art

[0003] Recently, Tacotron - based end - to - end speech synthesis systems have shown impressive text - to - speech (TTS) results in terms of the prosody and naturalness of the synthesized speech. However, such systems have significant drawbacks in skipping or repeating some words in the input text during speech synthesis. This problem is caused by the end - to - end nature of such systems, where an uncontrollable attention mechanism is used for speech generation. The present disclosure addresses these problems by replacing the end - to - end attention mechanism inside the Tacotron system with an attention network that notifies durations. The network proposed by the present disclosure achieves comparable or improved synthesis performance and solves the problems within the Tacotron system. Summary of the Invention

[0004] According to some possible implementations, a method includes: receiving, by a device, a text input including a sequence of text components; determining, by the device and using a duration model, corresponding durations of the text components; generating, by the device, a first spectrogram set based on the sequence of text components; generating, by the device, a second spectrogram set based on the first spectrogram set and the corresponding durations of the text components; generating, by the device, spectrogram frames based on the second spectrogram set; generating, by the device, an audio waveform based on the spectrogram frames; and providing, by the device, the audio waveform as an output.

[0005] According to some possible implementations, a device includes: at least one memory configured to store program code; at least one processor configured to read the program code and operate in accordance with instructions of the program code, the program code including: receiving code configured to cause the at least one processor to receive a text input including a sequence of text components; determining code configured to cause the at least one processor to use a duration model to determine corresponding durations of the text components; generating code configured to cause the at least one processor to: generate a first spectrogram set based on the sequence of text components; generate a second spectrogram set based on the first spectrogram set and the corresponding durations of the text components; generate spectrogram frames based on the second spectrogram set; generate an audio waveform based on the spectrogram frames; and providing code configured to cause the at least one processor to provide the audio waveform as an output.

[0006] According to some possible implementations, a non-transitory computer-readable medium stores instructions that include one or more instructions which, when executed by one or more processors of a device, cause the one or more processors to: receive a text input that includes a sequence of text components; use a duration model to determine corresponding durations of the text components; generate a first spectrogram set based on the sequence of text components; generate a second spectrogram set based on the first spectrogram set and the corresponding durations of the text components; generate spectrogram frames based on the second spectrogram set; generate an audio waveform based on the spectrogram frames; and provide the audio waveform as an output. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 is a schematic diagram of an example implementation described herein;

[0008] Figure 2 is a diagram of an example environment in which the systems and / or methods described herein can be implemented;

[0009] Figure 3 is Figure 2 a diagram of example components of one or more devices of

[0010] Figure 4 is a flowchart of an example process for generating an audio waveform using a duration-aware attention network for text-to-speech synthesis. DETAILED DESCRIPTION

[0011] TTS systems have a wide variety of applications. However, most of the commercially adopted systems are mainly based on parametric systems, which have a large gap compared with human natural speech. Tacotron is a TTS synthesis system that is significantly different from traditional parametric-based TTS systems and can generate highly natural speech sentences. The entire system can be trained in an end-to-end manner, and the traditional complex language feature extraction part is replaced by an encoder-convolutional-stack-highway network-bidirectional gated recurrent unit (CBHG) module.

[0012] An end-to-end attention mechanism is used to replace the duration model used in traditional parametric systems, where in the end-to-end attention mechanism, the alignment between the input text (or phoneme sequence) and the speech signal is learned from the attention model, rather than the alignment based on the hidden Markov model (HMM). Another main difference associated with the Tacotron system is that it directly predicts the mel / linear spectrogram, which can be directly used by advanced vocoders (such as Wavenet and WaveRNN) to synthesize high-quality speech.

[0013] The Tacotron-based system is capable of generating more accurate and natural-sounding speech. However, the Tacotron system includes instabilities such as skipping and / or repeating the input text, which are inherent drawbacks when synthesizing the speech waveform.

[0014] Some implementations of this disclosure address the aforementioned input text skipping and repeating issues of the Tacotron-based system while maintaining its excellent synthesis quality. Additionally, some implementations of this disclosure address these instability issues and achieve a significantly improved naturalness in the synthesized speech.

[0015] The instability of Tacotron is mainly caused by its uncontrollable attention mechanism, which cannot guarantee that each input text can be synthesized sequentially without skipping or repeating.

[0016] Some implementations of this disclosure replace this unstable and uncontrollable attention mechanism with a duration-based attention mechanism, in which the input text is guaranteed to be synthesized sequentially without skipping or repeating. The main reason for the attention in the Tacotron-based system is the lack of alignment information between the source text and the target mel-spectrogram.

[0017] Generally, the length of the input text is much shorter than the length of the generated mel-spectrogram. A single character / phoneme from the input text can generate multiple frames of the mel-spectrogram, and this information is needed to model the input / output relationship through any neural network architecture.

[0018] The Tacotron-based system mainly uses an end-to-end mechanism to solve this problem, in which the generation of the mel-spectrogram depends on the known attention to the source input text. However, this attention mechanism is basically unstable because the attention of this attention mechanism is highly uncontrollable. Some implementations of this disclosure replace the end-to-end attention mechanism within the Tacotron system with a duration model that predicts how long each input character and / or phoneme lasts. In other words, the alignment between the output mel-spectrogram and the input text is achieved by replicating each input character and / or phoneme within a predetermined duration. The ground truth duration of the input text learned from our system is achieved through HMM-based forced alignment. Using the predicted duration, each target frame in the mel-spectrogram can be matched with a character / phoneme in the input text. The overall model architecture is depicted in the following figure.

[0019] Figure 1 is a schematic diagram of an embodiment described in this disclosure. As Figure 1As shown, with reference numeral 110, a platform (e.g., a server) may receive a text input that includes a sequence of text components. As shown, the text input may include a phrase, such as "This is a cat". The text input may include a sequence of text components shown as the characters "DH", "IH", "S", "IH", "Z", "AX", "K", "AE", and "AX".

[0020] Further as Figure 1 shown, with reference numeral 120, the platform may use a duration model to determine the corresponding durations of the text components. The duration model may include a model that receives an input text component and determines the duration of the text component. As an example, the phrase "This is a cat" may include a total duration of one second when audibly output. The corresponding text components of the phrase may include different durations that together make up the total duration.

[0021] As an example, the word "This" may include a duration of 400 milliseconds, the word "is" may include a duration of 200 milliseconds, the word "a" may include a duration of 100 milliseconds, and the word "cat" may include a duration of 300 milliseconds. The duration model may determine the corresponding constituent durations of the text components.

[0022] Further as Figure 1 shown, with reference numeral 130, the platform may generate a first spectrogram set based on the sequence of text components. For example, the platform may input the text components into a model that generates an output spectrogram based on the input text components. As shown, the first spectrogram set may include the corresponding spectrograms of each text component (e.g., shown as "1", "2", "3", "4", "5", "6", "7", "8", and "9").

[0023] Further as Figure 1 shown, with reference numeral 140, the platform may generate a second spectrogram set based on the first spectrogram set and the corresponding durations of the sequence of text components. The platform may generate the second spectrogram set by duplicating spectrograms based on the corresponding durations of the spectrograms. As an example, the spectrogram "1" may be duplicated such that the second spectrogram set includes three spectrogram components corresponding to the spectrogram "1", etc. The platform may use the output of the duration model to determine how to generate the second spectrogram set.

[0024] Further as Figure 1 shown, with reference numeral 140, the platform may generate spectrogram frames based on the second spectrogram set. The spectrogram frames may be formed by the corresponding constituent spectrogram components of the second spectrogram set. As Figure 1 shown, the spectrogram frames may be aligned with the prediction frames. In other words, the spectrogram frames generated by the platform may be precisely aligned with the expected audio output of the text input.

[0025] The platform can use various technologies to generate an audio waveform based on spectrogram frames and provide the audio waveform as an output.

[0026] In this way, some implementations of the present disclosure allow for more precise audio output generation associated with text-to-speech synthesis by leveraging a duration model that determines the corresponding durations of the input text components.

[0027] Figure 2 FIG. is an illustration of an example environment 200 that can implement the systems and / or methods described herein. As Figure 2 shown, the environment 200 can include a user device 210, a platform 220, and a network 230. The devices of the environment 200 can be interconnected by a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.

[0028] The user device 210 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with the platform 220. For example, the user device 210 can include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., a pair of smart glasses or a smartwatch), or a similar device. In some implementations, the user device 210 can receive information from and / or send information to the platform 220.

[0029] The platform 220 includes a device capable of generating an audio waveform using an attention network that notifies duration for text-to-speech synthesis, as described elsewhere herein. In some implementations, the platform 220 can include a cloud server or a group of cloud servers. In some implementations, the platform 220 can be designed as a modular platform such that certain software components can be swapped in or out according to specific needs. Thus, the platform 220 can be easily and / or quickly reconfigured for different uses.

[0030] In some implementations, as shown, the platform 220 can be hosted in a cloud computing environment 222. It should be noted that while the implementations described herein depict the platform 220 as being hosted in the cloud computing environment 222, in some implementations, the platform 220 is not cloud-based (i.e., can be implemented outside of a cloud computing environment) or can be partially cloud-based.

[0031] The cloud computing environment 222 includes the environment hosting the platform 220. The cloud computing environment 222 can provide services such as computing, software, data access, storage, etc. without the need for the end user (e.g., the user device 210) to know the physical location and configuration of the systems and / or devices of the hosting platform 220. As shown in the figure, the cloud computing environment 222 can include a set of computing resources 224 (this set of computing resources is collectively referred to as "computing resources 224", and a single computing resource is called "computing resource 224").

[0032] The computing resources 224 include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some implementations, the computing resources 224 can control the platform 220. Cloud resources can include computing instances running in the computing resources 224, storage devices provided in the computing resources 224, data transmission devices provided by the computing resources 224, etc. In some implementations, the computing resources 224 can communicate with other computing resources 224 through a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.

[0033] Further as Figure 2 shown, the computing resources 224 include a set of cloud resources, such as one or more applications ("APP") 224-1, one or more virtual machines ("VM") 224-2, virtualized memory ("VS") 224-3, one or more hypervisors ("HYP") 224-4, etc.

[0034] The application 224-1 includes one or more software applications that can be provided to or accessed by the user device 210 and / or the sensor device 220. The application 224-1 can eliminate the need to install and run software applications on the user device 210. For example, the application 224-1 can include software associated with the platform 220 and / or any other software that can be provided through the cloud computing environment 222. In some implementations, one application 224-1 can send information to / receive information from one or more other applications 224-1 through the virtual machine 224-2.

[0035] The virtual machine 224-2 includes a software implementation of a machine (e.g., a computer) that runs programs, similar to a physical machine. Depending on the correspondence and use of the virtual machine 224-2 to any actual machine, the virtual machine 224-2 can be a system virtual machine or a process virtual machine. The system virtual machine can provide a complete system platform that supports the operation of a complete operating system ("OS"). The process virtual machine can run a single program and can support a single process. In some implementations, the virtual machine 224-2 can run on behalf of a user (e.g., the user device 210) and can manage the infrastructure of the cloud computing environment 222, such as data management, synchronization, or long-term data transmission.

[0036] The virtualized memory 224-3 includes one or more storage systems and / or one or more devices that use virtualization technology within the storage system or device of the computing resource 224. In some implementations, in the context of a storage system, the types of virtualization may include block virtualization and file virtualization. Block virtualization may refer to the abstraction (or separation) of logical storage from physical storage such that the storage system can be accessed without regard to the physical storage or heterogeneous architecture. The separation may allow the administrator of the storage system to have flexibility in how the administrator manages the storage of end users. File virtualization may eliminate the dependency between the data accessed at the file level and the location where the files are physically stored. This may optimize memory usage, server consolidation, and / or non-disruptive file migration performance.

[0037] The hypervisor 224-4 may provide hardware virtualization technology that allows multiple operating systems (e.g., “guest operating systems”) to run simultaneously on a host computer such as the computing resource 224. The hypervisor 224-4 may present a virtual operating platform to the guest operating systems and may manage the operation of the guest operating systems. Multiple instances of each operating system may share the virtualized hardware resources.

[0038] The network 230 includes one or more wired networks and / or wireless networks. For example, the network 230 may include a cellular network (e.g., a fifth-generation (5G) network, a long-term evolution (LTE) network, a third-generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, etc., and / or a combination of these or other types of networks.

[0039] Figure 2 The number and arrangement of the devices and networks shown are provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks arranged differently from those Figure 2 shown. Additionally, Figure 2 two or more of the devices shown may be implemented within a single device, or Figure 2 a single device shown may be implemented as multiple distributed devices. Additionally or alternatively, a set of devices (e.g., one or more devices) of the environment 200 may perform one or more functions described as being performed by another set of devices of the environment 200.

[0040] Figure 3A diagram of an example component of device 300. Device 300 may correspond to user device 210 and / or platform 220. As Figure 3 shown, device 300 may include bus 310, processor 320, memory 330, storage component 340, input component 350, output component 360, and communication interface 370.

[0041] Bus 310 includes components that permit communication between the components of device 300. Processor 320 is implemented in hardware, firmware, or a combination of hardware and software. Processor 320 is a central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or another type of processing component. In some implementations, processor 320 includes one or more processors that can be programmed to perform functions. Memory 330 includes random access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by processor 320.

[0042] Storage component 340 stores information and / or software related to the operation and use of device 300. For example, storage component 340 may include a hard disk (e.g., a magnetic disk, optical disk, magneto-optical disk, and / or solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cassette tape, a magnetic tape, and / or another type of non-transitory computer-readable medium, as well as corresponding drives.

[0043] Input component 350 includes components that permit device 300 to receive information, such as through user input (e.g., a touch screen display, keyboard, keypad, mouse, button, switch, and / or microphone). Additionally or alternatively, input component 350 may include sensors for sensing information (e.g., a global positioning system (GPS) component, accelerometer, gyroscope, and / or actuator). Output component 360 includes components that provide output information from device 300 (e.g., a display, speaker, and / or one or more light-emitting diodes (LEDs)).

[0044] The communication interface 370 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter), which enable the device 300 to communicate with other devices, such as through a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. The communication interface 370 may allow the device 300 to receive information from another device and / or provide information to another device. For example, the communication interface 370 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.

[0045] The device 300 may execute one or more of the processes described herein. The device 300 may execute these processes in response to the processor 320 running software instructions stored by a non-transitory computer-readable medium such as the memory 330 and / or the storage component 340. In this document, a computer-readable medium is defined as a non-transitory memory device. A memory device includes a memory space within a single physical storage device or a memory space distributed across multiple physical storage devices.

[0046] The software instructions may be read into the memory 330 and / or the storage component 340 from another computer-readable medium through the communication interface 370, or from another device into the memory 330 and / or the storage component 340. When run, the software instructions stored in the memory 330 and / or the storage component 340 may cause the processor 320 to execute one or more of the processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with the software instructions to execute one or more of the processes described herein. Accordingly, the implementations described herein are not limited to any particular combination of hardware circuitry and software.

[0047] Figure 3 The number and arrangement of the components shown are provided as an example. In practice, the device 300 may include additional components, fewer components, different components, or components arranged differently from Figure 3 the components shown. Additionally or alternatively, a set of components (e.g., one or more components) of the device 300 may perform one or more functions described as being performed by another set of components of the device 300.

[0048] Figure 4 is a flowchart of an example process 400 for generating an audio waveform using an attention network that notifies a duration for text-to-speech synthesis. In some implementations, one or more process blocks may be performed by the platform 220 Figure 4 In some implementations, one or more process blocks may be performed by another device or set of devices (e.g., the user device 210) separate from or including the platform 220 Figure 4 of one or more process blocks.

[0049] As Figure 4 shown, process 400 may include receiving, by a device, a text input that includes a sequence of text components (block 410).

[0050] For example, platform 220 may receive a text input that is to be converted into an audio output. The text components may include characters, phonemes, n-grams, words, letters, etc. The sequence of text components may form a sentence, phrase, etc.

[0051] Further as Figure 4 shown, process 400 may include determining, by the device and using a duration model, the respective durations of the text components (block 420).

[0052] The duration model may include a model that receives the input text components and determines the durations of the input text components. Platform 220 may train the duration model. For example, platform 220 may use machine learning techniques to analyze data (e.g., training data, such as historical data, etc.) and create a duration model. Machine learning techniques may include, for example, supervised techniques and / or unsupervised techniques, such as artificial networks, Bayesian statistics, learning automata, hidden Markov modeling, linear classifiers, quadratic classifiers, decision trees, association rule learning, etc.

[0053] Platform 220 may train the duration model by aligning spectrogram frames of known durations with the sequence of text components. For example, platform 220 may use HMM-based forced alignment to determine the ground truth durations of the input text sequence of the text components. Platform 220 may train the duration model by using predicted or target spectrogram frames of known durations and a known input text sequence that includes text components.

[0054] Platform 220 may input the text components into the duration model and determine, based on the output of the model, the respective durations of the identified text components or information associated with the respective durations of the text components. The identified respective durations or information associated with the respective durations may be used to generate a second spectrogram set, as described below.

[0055] Further as Figure 4 shown, process 400 may include determining whether the duration model has been used to determine the respective durations of each text component (block 430).

[0056] For example, platform 220 may iteratively or simultaneously determine the respective durations of the text components. Platform 220 may determine whether the duration has been determined for each text component of the input text sequence.

[0057] Further as Figure 4As shown, if the duration model is not used to determine the corresponding duration of each text component (box 430 - No), then process 400 may include returning to box 420.

[0058] For example, platform 220 may input text components for which the duration has not been determined into the duration model until the duration of each text component is determined.

[0059] Further as Figure 4 shown, if the duration model is used to determine the corresponding duration of each text component (box 430 - Yes), then process 400 may include generating a first spectrogram set by the device based on the sequence of text components (box 440).

[0060] For example, platform 220 may generate an output spectrogram of text components corresponding to the sequence of input text components. Platform 220 may utilize a CBHG module to generate the output spectrogram. The CBHG module may include a stack of one-dimensional (1-D) convolutional filters, a set of highway networks, bidirectional gated recurrent units (GRUs), recurrent neural networks (RNNs), and / or other components.

[0061] In some implementations, the output spectrogram may be a Mel-frequency cepstral (MFC) spectrogram. The output spectrogram may include any type of spectrogram for generating spectrogram frames.

[0062] Further as Figure 4 shown, process 400 may include generating a second spectrogram set by the device based on the first spectrogram set and the corresponding durations of the sequence of text components (box 450).

[0063] For example, platform 220 may use the first spectrogram set and the information identifying or associated with the corresponding durations of the text components to generate the second spectrogram set.

[0064] As an example, platform 220 may copy each spectrogram of the first spectrogram set based on the corresponding duration of the underlying text component corresponding to the spectrogram. In some cases, platform 220 may copy the spectrogram based on a replication factor, a time factor, etc. In other words, the output of the duration model may be used to determine factors for copying a specific spectrogram, generating additional spectrograms, etc.

[0065] Further as Figure 4 shown, process 400 may include generating spectrogram frames by the device based on the second spectrogram set (box 460).

[0066] For example, platform 220 may generate spectrogram frames based on a second spectrogram set. The second spectrogram set jointly forms the spectrogram frames. As described elsewhere in this document, spectrogram frames generated using a duration model may more precisely resemble a target frame or a predicted frame. In this way, some implementations herein improve the accuracy of TTS synthesis, improve the naturalness of the generated speech, improve the prosody of the generated speech, and so on.

[0067] Further as Figure 4 shown, process 400 may include generating an audio waveform based on the spectrogram frames by a device (block 470), and providing the audio waveform as an output by the device (block 480).

[0068] For example, platform 220 may generate an audio waveform based on the spectrogram frames and provide the audio waveform for output. As an example, platform 220 may provide the audio waveform to an output component (such as a speaker, etc.), may provide the audio waveform to another device (such as user device 210), may transmit the audio waveform to a server or another terminal, etc.

[0069] Although Figure 4 illustrative blocks of process 400 are shown, in some implementations, process 400 may include additional blocks, fewer blocks, different blocks, or blocks arranged differently from Figure 4 the blocks depicted. Additionally or alternatively, two or more blocks of process 400 may be executed in parallel.

[0070] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Modifications and variations may be made in light of the above disclosure, or may be obtained from the practice of the implementations.

[0071] As used herein, the term "component" is intended to be broadly construed as hardware, firmware, or a combination of hardware and software.

[0072] Obviously, the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specific control hardware or software code for implementing these systems and / or methods does not limit the implementations. Thus, the operation and behavior of the systems and / or methods are not described herein with reference to specific software code - it should be understood that the software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0073] Even if specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may only directly depend on one claim, the disclosure of possible implementations includes the combination of each dependent claim with every other claim in the claim set.

[0074] Elements, acts, or instructions used herein should not be construed as critical or essential unless explicitly described as such. Additionally, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Further, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." Where only one item is intended, the term "a" or similar language is used. Additionally, as used herein, the terms "having," "have," "containing," or similar terms are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise explicitly stated.

Claims

1. A method for text-to-speech conversion analysis, comprising: Receiving, by a device, a text input including a sequence of text components; The text component being at least one of a phoneme and a character; Determining, by the device and using a duration model, corresponding durations of the text components; The duration model being a model obtained by training by aligning spectrogram frames of known durations with a sequence of text components; Generating, by the device based on the sequence of text components, a first spectrogram set through a CBHG module; the CBHG module including a stack of one-dimensional convolutional filters, a set of highway networks, bidirectional gated recurrent units, recurrent neural networks, and / or other components; Generating, by the device based on the first spectrogram set and the corresponding durations of the text components, a second spectrogram set by copying spectrograms in the first spectrogram set; Generating, by the device, spectrogram frames based on corresponding constituent spectrogram components of the second spectrogram set; The spectrogram frames being aligned with an expected audio output of the text input; The length of the text input being less than the length of the spectrogram frames; a single character or phoneme from the text input being used to generate multiple frames in the spectrogram frames; Generating, by the device, an audio waveform based on the spectrogram frames; And Providing, by the device, the audio waveform as an output.

2. The method according to claim 1, wherein The second spectrogram set includes Mel-frequency cepstrum spectrograms.

3. The method according to claim 1, the method further comprising: Training the duration model using a set of prediction frames and training text components.

4. The method according to claim 1, the method further comprising: Training the duration model using hidden Markov model forced alignment techniques.

5. A device for text-to-speech conversion analysis, comprising: At least one memory configured to store program code; At least one processor configured to read the program code and operate according to instructions of the program code, the program code including: A receiving code configured to cause the at least one processor to receive a text input including a sequence of text components; the text component being at least one of a phoneme and a character; A determining code configured to cause the at least one processor to use a duration model to determine corresponding durations of the text components; the duration model being a model obtained by training by aligning spectrogram frames of known durations with a sequence of text components; A generating code configured to cause the at least one processor to: Generate a first spectrogram set based on the sequence of text components through a CBHG module; the CBHG module including a stack of one-dimensional convolutional filters, a set of highway networks, bidirectional gated recurrent units, recurrent neural networks, and / or other components; Generate a second spectrogram set based on the first spectrogram set and the corresponding durations of the text components by copying spectrograms in the first spectrogram set; Generate spectrogram frames based on the corresponding constituent spectrogram components of the second spectrogram set; the spectrogram frames are aligned with the expected audio output of the text input; the length of the text input is less than the length of the spectrogram frames; a single character or phoneme from the text input is used to generate multiple frames in the spectrogram frames; Generate an audio waveform based on the spectrogram frames; and Provide code configured to cause the at least one processor to provide the audio waveform as an output.

6. The device according to claim 5, wherein The second spectrogram set includes Mel-frequency cepstral spectrograms.

7. The apparatus according to claim 5, the apparatus further comprising: Training code configured to cause the at least one processor to train the duration model using a set of prediction frames and training text components.

8. The apparatus according to claim 5, the apparatus further comprising: Training code configured to cause the at least one processor to train the duration model using Hidden Markov Model forced alignment techniques.

9. A non-transitory computer-readable medium storing instructions that include one or more instructions that, when executed by one or more processors of a device, cause the one or more processors to: Receive a text input including a sequence of text components; the text components are at least one of phonemes and characters; Use a duration model to determine the corresponding duration of the text component; The duration model is a model obtained by training by aligning spectrogram frames with known durations and a sequence of text components; Generate a first spectrogram set based on the sequence of text components by a CBHG module; the CBHG module includes a stack of one-dimensional convolutional filters, a set of highway networks, bidirectional gated recurrent units, recurrent neural networks, and / or other components; Generate a second spectrogram set by replicating spectrograms in the first spectrogram set based on the first spectrogram set and the corresponding durations of the text components; Generate spectrogram frames based on the corresponding constituent spectrogram components of the second spectrogram set; The spectrogram frames are aligned with the expected audio output of the text input; The length of the text input is less than the length of the spectrogram frames; a single character or phoneme from the text input is used to generate multiple frames in the spectrogram frames; Generate an audio waveform based on the spectrogram frames; And Provide the audio waveform as an output.

10. The non-transitory computer-readable medium according to claim 9, wherein, The second spectrogram set includes Mel-frequency cepstral spectrograms.

11. The non-transitory computer-readable medium according to claim 9, wherein, The second spectrogram set includes a different number of spectrograms compared to the first spectrogram set.

12. A computer device, comprising: A processor; And A memory for storing a computer-readable program; Wherein, the processor is configured to execute the computer-readable program to perform the method according to any one of claims 1 to 4.

13. A device for text-to-speech conversion analysis, comprising: A receiving module configured to receive a text input including a sequence of text components; The text components are at least one of phonemes and characters; A determining module configured to use a duration model to determine the corresponding durations of the text components; The duration model is a model obtained by training by aligning spectrogram frames and text components with known durations; A generation module, configured to: Generate a first spectrogram set through a CBHG module based on the sequence of the text components; the CBHG module includes a stack of one-dimensional convolutional filters, a set of highway networks, bidirectional gated recurrent units, recurrent neural networks, and / or other components; Generate a second spectrogram set by replicating spectrograms in the first spectrogram set based on the first spectrogram set and the corresponding durations of the text components; Generate spectrogram frames based on the corresponding constituent spectrogram components of the second spectrogram set; The spectrogram frames are aligned with the expected audio output of the text input; The length of the text input is less than the length of the spectrogram frames; a single character or phoneme from the text input is used to generate multiple frames in the spectrogram frames; Generate an audio waveform based on the spectrogram frames; And A providing module, configured to provide the audio waveform as an output.

Citation Information

Patent Citations

  • Text to speech synthesis using deep neural network with constant unit length spectrogram

    US10186252B1