Generating audio data using artificial intelligence techniques

By integrating AI techniques to process phonetic and prosodic data, the method addresses delays and accuracy issues in conventional speech synthesis, enhancing the capture of style and emotion in voice data generation.

JP2026528702APending Publication Date: 2026-08-25INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026503586
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-15
Filing Date
2024-07-01
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Conventional speech synthesis methods face challenges such as large delays, low accuracy, and an inability to adequately capture styles and emotions in voice data.

Method used

Implementing artificial intelligence techniques to generate voice data by combining language models, text-to-speech frontend models, and text-to-speech prosody models, using a language model-text-to-speech algorithm adapter to convert output into phonetic and prosodic data, and incorporating phonetic and prosodic information into the synthesis process.

Benefits of technology

This approach reduces latency and improves accuracy in generating speech data, effectively capturing style and emotion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026528702000001_ABST
    Figure 2026528702000001_ABST
Patent Text Reader

Abstract

This specification provides methods, systems, and computer program products for generating speech data using artificial intelligence techniques. A computer implementation method includes the steps of: implementing one or more artificial intelligence techniques in relation to one or more speech synthesis tasks; generating at least one data sequence, including one or more phonetic and prosodic data, in a plurality of consecutive parts by processing at least one previously generated data sequence using the one or more artificial intelligence techniques; and generating speech data corresponding to at least a portion of the data sequence by processing the at least portion of the data sequence using at least one artificial intelligence-based speech synthesis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to information technology, and more specifically, to language and speech processing.

Background Art

[0002] More specifically, there are cases where it is desired to predict the text output of words or tokens based on previous text, and when such output should be used in a conversation-related context, it is necessary to convert the text into voice data. However, conventional speech synthesis methods have limitations such as, for example, the problem of large delays, the problem of accuracy, and the inability to sufficiently capture styles and emotions in voice data.

Summary of the Invention

[0003] In at least one embodiment, a method for generating voice data using artificial intelligence techniques is provided. An exemplary computer-implemented method includes implementing one or more artificial intelligence techniques in relation to one or more speech synthesis tasks; and generating in a plurality of consecutive parts by processing at least one data sequence including one or more of acoustic data and prosodic data using the one or more artificial intelligence techniques on at least one previously generated data sequence. Additionally, this method also includes generating voice data corresponding to at least a part of the data sequence by processing the at least a part of the data sequence using at least one artificial intelligence-based speech synthesis model.

[0004] At least one embodiment may include combining at least one language model (LM), at least one text-to-speech frontend (TTS-FE) model, and at least one text-to-speech prosody (TTS-P) model. Additionally or alternatively, one or more embodiments may include using a language model-text-to-speech algorithm adapter to convert at least a portion of the output from at least one LM into phonetic and prosodic data to be used as input to one or more of the at least one TTS-FE model and at least one TTS-P model.

[0005] Another embodiment or element of the present invention may be implemented in the form of a computer program product that tangibly embodies computer-readable instructions, which, when implemented, cause a computer to perform a plurality of method steps described herein. Furthermore, another embodiment or element of the present invention may be implemented in the form of a system comprising memory and at least one processor coupled to the memory, and may be configured to perform the method steps described herein. Still further, another embodiment or element of the present invention may be implemented in the form of means for performing the method steps or elements described herein; these means may include a hardware module or a combination of hardware and software modules, wherein the software module is stored in a tangible computer-readable storage medium (or a plurality of such mediums).

[0006] The exemplary embodiments can offer significant advantages over conventional speech synthesis methods. For example, issues associated with latency, accuracy, and the inability to adequately capture style and emotion in speech data are overcome in one or more embodiments described above by generating speech data using artificial intelligence techniques in relation to phonetic and / or prosodic data.

[0007] These and other objects, features and advantages of the present invention will become apparent from the following detailed description of its exemplary embodiments, which will be read in conjunction with the accompanying drawings. [Brief explanation of the drawing]

[0008] [Figure 1] This figure shows training of a LM using a phonetic transcript, according to an exemplary embodiment of the present invention.

[0009] [Figure 2] This figure shows an exemplary embodiment of the present invention: a text-to-speech (TTS) system and the generation of speech data using a LM trained with phonetic information.

[0010] [Figure 3] This figure shows training of a LM using phonetic transcripts and prosodic features according to an exemplary embodiment of the present invention.

[0011] [Figure 4] This figure shows a TTS system according to an exemplary embodiment of the present invention, and the generation of speech data using a LM trained with phonetic information and prosodic features.

[0012] [Figure 5] This figure shows the fine-tuning of the LM using audio data according to an exemplary embodiment of the present invention.

[0013] [Figure 6] This figure shows a language model-TTS (LM-TTS) adapter architecture according to an exemplary embodiment of the present invention.

[0014] [Figure 7] This figure shows an exemplary LM-TTS adapter output according to an exemplary embodiment of the present invention.

[0015] [Figure 8] This is a flowchart illustrating a method according to an exemplary embodiment of the present invention.

[0016] [Figure 9] This figure shows a computing environment in which at least one embodiment of the present invention may be implemented. [Modes for carrying out the invention]

[0017] As described herein, at least one embodiment includes generating speech data using artificial intelligence techniques. In one or more embodiments, one or more LMs are trained to output the next text token in a sequence, given some text input. Additionally or alternatively, one or more LMs may be trained to output the next text token based on one or more previously generated tokens, and further optionally based on queries. Such output text is converted into speech data, for example, using at least one speech synthesis system, when used in the context and / or implementation of a voice conversation.

[0018] Furthermore, as used herein, a generative language model (G-LM) refers to an autoregressive language model or autoregressive text prediction model that sequentially generates words (e.g., a fixed number of words at a time) given a previously synthesized sequence of words (optionally preceded by a text query), and can also output a sequence of related internal states. Additionally, as used herein, a generative text-to-speech frontend (G-TTS-FE) model refers to a neural model that sequentially generates TTS-FE symbol sequences (e.g., symbols suitable for a speech synthesis procedure (e.g., a fixed number of symbols at a time)) from plain text and / or annotated text (e.g., a word sequence), where the symbol sequence suitable for speech synthesis includes at least one phonetic sequence (e.g., a phoneme, lexicographical stress, etc.), optionally includes phrase type information, and optionally includes one or more annotations such as part of speech, word emphasis, etc.

[0019] Furthermore, as used herein, a generative text-to-speech prosody (G-TTS-P) model refers to a neural model that sequentially generates TTS-P sequences from plain text and / or annotated text, each containing prosodic feature vectors (e.g., a certain number of prosodic feature vectors at a time) that guide at least one subsequent speech synthesis system. Such prosodic feature vectors may include, for example, overall and / or local evaluations of phoneme utterance rate, pitch, and volume. Additionally, such prosodic feature vectors may be derived, for example, from one or more statistical measurements taken over one or more hierarchical time intervals and / or normalized to become one or more speaker-agnostic features. Furthermore, as used herein, a generative speech synthesis (G-SS) model refers to a neural model that sequentially generates at least one output speech waveform from at least one TTS-FE symbol sequence and at least one TTS-P feature vector sequence.

[0020] Thus, at least one embodiment can include generating and / or implementing a combined model of G-LM, G-TTS-FE, and G-TTS-P that sequentially generates the outputs of TTS-FE and TTS-P and optionally also sequentially generates one or more related text outputs, where the combined model is followed by a G-SS model. Additionally or alternatively, one or more embodiments can include generating and / or implementing a combined model of G-TTS-FE and G-TTS-P that sequentially generates the outputs of TTS-FE and TTS-P, is preceded by a G-LM model, uses its internal state sequence in addition to the text sequence as an input, and is followed by a G-SS model. Further, at least one embodiment can include generating and / or implementing a combined model of G-LM, G-TTS-FE, G-TTS-P, and G-SS that performs the entire end-to-end speech synthesis task.

[0021] Also, as described in more detail herein, at least one embodiment includes generating audio data using one or more TTS systems in conjunction with the output of at least one LM (e.g., at least one G-LM). Such embodiments include training at least one LM to produce one or more types of outputs, where such outputs can include phonetic transcription data, one or more phrase types and part-of-speech tagging, word and / or syllable stress information, pause information and / or phrase break information, prosodic information (e.g., prosodic markup that may be useful for real-time speech synthesis), and the like. In one or more embodiments, such outputs can be used as a direct input to a TTS synthesis model, whereby the TTS synthesis model can start synthesis with minimal latency and produce a more acoustically natural output that correctly conveys the meaning of the input.

[0022] As will be described in more detail herein, at least one embodiment includes modifying the LM by training the LM, for example, to output phonemes and linguistic information in parallel with the output words. In such an embodiment, such LM training uses the original LM training text and can include extracting one or more phonemes and linguistic information (e.g., clause type information) from the text using, for example, at least one TTS linguistic front end, grapheme-phoneme conversion, or other similar linguistic analysis tools.

[0023] Additionally, one or more embodiments include modifying the LM by training the LM, for example, to produce prosodic information such as pitch, phoneme duration, clause break information, etc. In such an embodiment, such LM training includes supplying the original LM training text to a TTS system and / or a TTS prosody prediction system (e.g., the G-TTS-P model and / or the G-TTS-FE model as a stand-alone or internal TTS component) and extracting prosodic information (e.g., pitch and energy curves, and phoneme duration) therefrom. One or more TTS systems can include, for example, the G-TTS-P model as an internal component of the TTS system, while other TTS systems (e.g., end-to-end TTS systems) do not include such a model. In such cases, voice data can be synthesized and the desired prosodic features can be extracted from the waveform.

[0024] Furthermore, at least one embodiment includes modifying at least one TTS system to use information generated by at least one LM (e.g., G-LM) as an input instead of just a text input, according to at least a portion of the techniques detailed herein.

[0025] As described above, one or more embodiments include training a LM to produce augmented phonetic transcription information and / or mixed output, where, for example, a query is given as plain text, but the response is produced at least partially as phonetic output. To properly train the LM, at least one embodiment may include using at least one existing LM training text corpus and converting at least a portion of such text into augmented phonetic script information using at least one TTS system linguistic front-end. Such resulting script information can then be used to train the LM.

[0026] One or more embodiments may also include incorporating additional information into such script information by processing at least a portion of the script information using a TTS model. Such additional information, in connection with training at least one LM, can produce improved intonation, phoneme duration, pitch information, pause information, etc., which may be beneficial to the associated TTS system. As just one example, such additional information may include prosodic information that can convey speaker-independent speech synthesis features, such as hierarchical prosody control (HPC) features. Other types of information may include, for example, word emphasis, emotion (e.g., joy, apology, etc.), and speech style (e.g., conversation, announcement, reading, etc.). Also, in connection with incorporating additional information into script information, in such action, there is no need to use and / or acquire additional data, since the additional information can be generated from the original training text corpus.

[0027] In one or more embodiments, the TTS system may use such extended phonetic script information and additional and / or prosodic information to directly input the output of the LM, thereby reducing latency as the TTS system does not need to perform long look-aheads to understand the linguistic context and represent speech in a natural prosody.

[0028] Furthermore, at least one embodiment may include using speech data to fine-tune and / or improve the LM. Such an embodiment may include extracting phoneme and intonation information from the speech data and training the LM using at least a portion of such extracted information. Such speech data may include, for example, a large speech corpus from many speakers that can facilitate general intonation improvements, and / or speech data from a single speaker to fit a specific voice or style (e.g., a conversational voice style). Additionally or alternatively, for multi-speaker data, one or more speaker-independent features (e.g., HPC features) may be used and / or required.

[0029] As detailed herein, at least one embodiment may provide beneficial effects such as reduced latency and improved accuracy in the implementation of LM and TTS systems. In another embodiment, a pre-trained LM is utilized and unmodified. In such an embodiment, an LM-TTS adapter is created and / or implemented, which takes a pre-trained LM output and / or its internal state parameters as input and outputs phonetic and prosodic information required by at least one corresponding TTS. In one exemplary embodiment, the LM-TTS adapter uses an encoder-decoder architecture. In such an embodiment, the encoder may generate an encoded vector for each input word and have an adjustable word lookup function. The decoder takes the encoder output and one or more previous outputs and produces phonetic and prosodic information for the current word. Alternatively, one or more embodiments may include implementing separate phonetic and prosodic decoders.

[0030] Figure 1 shows the training of a LM using phonetic transcripts according to an exemplary embodiment of the present invention. As an example, Figure 1 shows the conversion of a text LM 102 to a phonetic LM 104 using a training corpus 105 of phonetic data, where such a trained phonetic LM can generate phonetic transcript data (e.g., phonemes, stress information, phrase type information, part-of-speech information, segmentation information, etc.) in place of and / or in addition to the text data. In one or more embodiments, the text LM 102 may include a pre-trained text LM, and generating the training corpus 105 of phonetic data may include converting at least a portion of the training corpus 103 of text data to phonetic data using a linguistic front-end 106 (e.g., one or more text processing programs) from a TTS system.

[0031] In at least one embodiment, training speech data (e.g., spoken text 109) derived using an automated speech recognition (ASR) method 108 can also be incorporated into the training corpus 105 along with the text data provided by the linguistic front-end 106, as optionally shown in Figure 1. Thus, the trained phonetic LM 104 can process text data input (e.g., a question) and generate a phoneme response (e.g., an answer to the question). One or more embodiments may also include incorporating one or more external tags, such as sentiment or word emphasis.

[0032] Figure 2 shows a TTS system and the generation of speech data using an LM trained with phonetic information, according to an exemplary embodiment of the present invention. As an example, Figure 2 shows a phonetic LM204 (e.g., a trained phonetic LM as shown in Figure 1) processing a text input (e.g., a query) and generating a phonetic output (e.g., a response to the query). The phonetic output (e.g., a phonetic script) is then provided as input to a TTS system 210 (e.g., a reduced TTS system without a linguistic frontend, and / or a speech synthesis system trained to be controlled by a selected set of prosodic features predicted by an LM such as the phonetic LM304 shown in Figure 3), and the TTS system 210 processes the input and generates corresponding speech data. According to one or more embodiments, the phonetic LM204 has a very large context for producing the necessary annotations of the phonetic output, and the phonetic LM204 can be trained to correctly handle homonym and / or text normalization tasks.

[0033] Figure 3 shows training of an LM using phonetic transcripts and prosodic features according to an exemplary embodiment of the present invention. As an example, Figure 3 shows modifying a text LM 302 into a phonetic LM 304 using a training corpus 307 of phonetic and prosodic data, where such a trained phonetic LM 304 can generate improved phonetic transcript data having associated prosodic information instead of and / or in addition to the text data. Generating the training corpus 307 may include converting at least a portion of the training corpus 303 of text data into phonetic data using the TTS linguistic frontend 306. As also shown in Figure 3, one or more embodiments include improving the training corpus 307 by adding prosodic and / or intonation information derived from a prosodic model 312. The prosodic information incorporated into the training corpus 307 may include, for example, one or more prosodic hint tags (e.g., hints for longer syllables, shorter syllables, higher pitches, lower pitches, etc.), one or more hierarchical tags for speech parts (e.g., sentences, words, and phoneme modifiers), a total pitch curve, and phoneme durations. Such prosodic information may be generated, in addition to using the prosodic model 312, for example, by applying a TTS system to the training corpus 303 of text data (e.g., offline) and / or by using actual speech data (e.g., using the prosodic model 312).

[0034] Figure 4 shows a TTS system and the generation of speech data using an LM trained with phonetic information and prosodic features, according to an exemplary embodiment of the present invention. As an example, Figure 4 shows that a phonetic and prosodic LM 404 (e.g., a trained phonetic and prosodic LM as shown in Figure 3) processes a text input (e.g., a query) and generates a phonetic and prosodic output (e.g., a response to the query). The output is then provided as input to a TTS system 410 (e.g., a low-latency TTS system), which processes the input and generates corresponding speech data. In at least one embodiment, the TTS system 410 is modified to use and / or process a phonetic script and associated prosodic information (generated by the phonetic and prosodic LM 404) as input, thereby reducing the look-ahead function of the TTS system 410 (for example, reducing the look-ahead function from several words to several phonemes), and providing a system that can generate speech data at approximately the same speed as the LM can generate output data supplied to the TTS system.

[0035] Figure 5 shows a fine-tuning of an LM using speech data according to an exemplary embodiment of the present invention. As an example, Figure 5 shows modifying a text LM 502 into a phonetic LM 504 using a training corpus 505 of phonetic data, where such a trained phonetic LM 504 can generate improved phonetic transcription data instead of and / or in addition to the text data. As also shown in Figure 5, one or more embodiments include improving the training corpus 505 using training speech data 509 (e.g., actual speech data) processed using an ASR 508 and a prosodic feature extractor 516, and then fine-tuning the phonetic LM 504. The ASR 508 is used to extract information such as phonemes and phoneme durations, and the prosodic feature extractor 516 can extract information such as pitch curves and energy. By fine-tuning the phonetic LM504 using speech data, the quality and naturalness of the phonetic LM504's output can be improved, making it easier to adapt the phonetic LM504 to specific aspects such as particular speakers, specific speech styles (e.g., conversational voice styles), emotions, and pronunciation.

[0036] Additionally or alternatively, at least one embodiment includes connecting and / or using at least one LM and at least one TTS system in conjunction with at least one LM-TTS adapter that receives LM output and its internal state and converts such output into one or more phonemes and one or more prosodic controls that can be used as input to a TTS system. In such an embodiment, the LM-TTS adapter model can leverage the hidden state of the language model to improve accuracy and support conversion in parallel with LM text generation.

[0037] Figure 6 shows an LM-TTS adapter architecture according to an exemplary embodiment of the present invention. As an example, Figure 6 shows an exemplary LM-TTS adapter model representing an enhanced version of the transformer encoder-decoder architecture. Specifically, Figure 6 shows that LM602 processes a text query to produce one or more internal state vectors, word embedding vectors, and language model tokens 620 (e.g., one or more text word fragments). Figure 6 also shows an LM-TTS adapter 660, which includes an encoder 662 and a decoder 664.

[0038] Encoder 662 processes an input that includes at least a portion of language model tokens 620, combined with internal state and contextual embedding vectors from LM602. Encoder 662 outputs an embedding vector for each word. In one or more embodiments, the at least portion of language model tokens 620 and the semantic information associated with the embeddings may be received from one or more internal layers of LM602 (e.g., one or more deep layers and / or one or more shallow layers) and supplied to encoder 662. As just one example, in at least one embodiment, encoder 662 includes two transformer layers with an embedding dimension of 512 and eight attention heads.

[0039] Additionally, in one or more embodiments, the encoder 662 does not pay attention to future words, and this limitation ensures that the phoneme prediction of a word is invariant to future words. In such embodiments, this limitation can also be relaxed, and a fixed lookahead function can be added and / or incorporated, thereby facilitating the trade-off between delay and contextual increase.

[0040] As shown in Figure 6, the decoder 664 processes at least a portion of the output generated by the encoder 662 to produce and / or output one or more phonemes and prosodic information (e.g., one or more prosodic features). In one or more embodiments, each phoneme output by the decoder 664 is selected from at least one phonetic vocabulary, while the prosodic information may include one or more normalized prosodic observations. Each prosodic observation may include, for example, a normalized linear combination of statistical measures that evaluate a particular prosodic metric (e.g., pitch [Hz], rhythm [sound duration], volume [dB]) over a given period of time. In such embodiments, implementing the decoder 664 may include unfolding a set of hierarchical aggregates (e.g., sentence-level aggregates, word-level aggregates, etc.).

[0041] In at least one embodiment, the decoder 664 outputs one or more phoneme and prosodic information by autoregressive prediction. In such an embodiment, the decoder 664 takes the encoder output (e.g., word embedding vector) and the previously generated phoneme and prosodic output as input and predicts the next phoneme and prosodic output. An example is shown in Figure 7.

[0042] Figure 7 shows an example 700 of a modified LM or LM-TTS adapter output according to an exemplary embodiment of the present invention. As an example, Figure 7 shows an example 700 of a modified LM or LM-TTS adapter output in the form of a table containing information on the decoding step, LM text output, phonetic output, and word-level and sentence-level HPC parameters.

[0043] Referring again to Figure 6, in one or more exemplary embodiments, the decoder 664 may include four transformer layers with an embedding dimension of 512 and eight attention heads. While the LM 602 is generating text, the LM-TTS adapter 660 runs in parallel. As just one example, in at least one embodiment, each time the LM 602 completes the generation of a word, the encoder 662 runs on the LM output, which is then processed multiple times by the decoder 664. The decoder 664 autoregressively generates an output sequence of associated phonemes and prosody until a word-ending token is predicted by the decoder 664. After the decoder 664 stops, its output is sent to the TTS system 610 to be synthesized as speech data.

[0044] Figure 8 is a flowchart illustrating a method according to an embodiment of the present invention. Step 802 includes implementing one or more artificial intelligence methods (e.g., at least one artificial neural network (ANN) module) in relation to one or more speech synthesis tasks. In at least one embodiment, implementing one or more artificial intelligence methods includes combining at least one LM, at least one TTS-FE model, and at least one TTS-P model. Additionally or alternatively, implementing one or more artificial intelligence methods may include using a language model-text-speech algorithm adapter to convert at least a portion of the output from at least one LM into phonetic and prosodic data to be used as input to one or more of the at least one TTS-FE model and at least one TTS-P model. In such embodiments, implementing one or more artificial intelligence techniques may include using a language model-text-speech algorithm adapter, in conjunction with generating one or more HPC features derived from one or more statistical measurements taken over one or more hierarchical time intervals and normalized to represent one or more speaker-independent features, wherein at least a portion of one or more HPC features includes one or more global and one or more local evaluations with respect to at least one of phoneme utterance rate, pitch, and / or volume.

[0045] Furthermore, in one or more embodiments, implementing one or more artificial intelligence techniques includes modifying at least a portion of at least one LM using one or more phonetic data items. In such embodiments, implementing one or more artificial intelligence techniques may include extracting one or more phonetic data items from at least one training text dataset using at least one TTS-FE model.

[0046] Additionally or alternatively, implementing one or more artificial intelligence techniques may include modifying at least a portion of at least one LM using one or more prosodic data items (e.g., information related to pitch, information related to phoneme duration, and / or information related to phrase breaks). In such embodiments, implementing one or more artificial intelligence techniques may include extracting one or more prosodic data items from at least one training text dataset using at least one TTS-P model.

[0047] Furthermore, in at least one embodiment, implementing one or more artificial intelligence techniques includes modifying at least a portion of at least one LM using one or more speech data items in relation to at least one automatic speech recognition technique and one or more prosodic feature extraction techniques.

[0048] Step 804 includes generating at least one data sequence, which includes one or more phonetic and prosodic data, in a plurality of consecutive parts by processing at least one previously generated data sequence (and optionally, related text queries) using one or more artificial intelligence techniques. In one or more embodiments, generating at least one data sequence includes generating one or more phonemes and generating one or more prosodic information items in conjunction with one or more output words related to at least one previously generated data sequence. In such embodiments, generating one or more prosodic information items may involve using a language model-text-speech algorithm adapter in conjunction with generating one or more hierarchical prosodic control (HPC) features derived from one or more statistical measurements taken over one or more hierarchical time intervals and normalized to represent one or more speaker-independent features, wherein at least a portion of the one or more HPC features includes at least one of one or more overall evaluations and one or more local evaluations relating to at least one of phoneme utterance rate, pitch, and volume.

[0049] Additionally or alternatively, generating at least one data sequence containing one or more phonetic and prosodic data in multiple consecutive parts may include generating one or more phonetic vectors and one or more prosodic vectors at once to generate the at least one data sequence.

[0050] Step 806 includes generating speech data corresponding to at least a portion of the data sequence by processing at least a portion of the data sequence using at least one artificial intelligence-based speech synthesis model. In at least one embodiment, generating speech data includes generating one or more speech waveforms corresponding to at least a portion of the data sequence.

[0051] The method shown in Figure 8 may also include automatically training at least one or more artificial intelligence methods using at least a portion of the generated speech data. Additionally, in one or more embodiments, software implementing the method shown in Figure 8 may be provided as a service in a cloud environment.

[0052] The exemplary embodiments described above offer significant advantages over conventional approaches. For example, some embodiments are configured to generate speech data using artificial intelligence techniques in relation to phonetic and / or prosodic data. These and other embodiments can effectively overcome the challenges associated with latency, accuracy, and the inability to fully capture style and emotion in speech data.

[0053] It should be understood that some embodiments described herein utilize one or more artificial intelligence models. As used herein, the term “model” is intended to be interpreted broadly and should be understood to include, for example, a set of executable instructions for generating recommendations and / or predictions implemented by a computer. For example, one or more of the models described herein may be trained to generate recommendations and / or predictions based on input text data, phonetic data, and / or prosodic information, and such recommendations and / or predictions may be used to initiate one or more automated actions (e.g., automatically generating speech data in relation to one or more TTS systems, automatically training one or more artificial intelligence techniques (e.g., one or more LMs)).

[0054] The method shown in Figure 8 may also include providing a system, as described herein, which includes separate software modules, each of which is embodied on a tangible computer-readable recordable storage medium. All modules (or any subset thereof) may reside on the same medium, or, for example, each may reside on a different medium. The modules may include any or all of the components shown in the figure and / or described herein. In embodiments of the present invention, the modules may run, for example, on a hardware processor. The method step may then be carried out using separate software modules of the system running on the hardware processor, as described above. Furthermore, the computer program product may include a tangible computer-readable recordable storage medium having code adapted to be executed to carry out at least one method step described herein, which includes providing a system having separate software modules.

[0055] Additionally, the method shown in Figure 8 may be implemented via a computer program product which may include computer-usable program code stored on a computer-readable storage medium in a data processing system, and the computer-usable program code is downloaded from a remote data processing system via a network. Furthermore, in embodiments of the present invention, the computer program product may include computer-usable program code stored on a computer-readable storage medium in a server data processing system, and the computer-usable program code is downloaded to a remote data processing system via a network for use on the computer-readable storage medium using a remote system.

[0056] Embodiments of the present invention or elements thereof may be implemented in the form of a device comprising a memory and at least one processor coupled to the memory, and may be configured to perform exemplary method steps.

[0057] Various aspects of this disclosure are illustrated by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of mechanical logic included in embodiments of computer program products (CPPs). With respect to any flowchart, operations may be performed in a different order than those shown in a given flowchart, depending on the technology involved. For example, again, depending on the technology involved, two operations shown in consecutive blocks of a flowchart may be performed in reverse order, as a single integrated step, simultaneously, or with at least partial time overlap.

[0058] Embodiments of a computer program product ("CPP Embodiment" or "CPP") are terms used in this disclosure to describe any set of one or more storage media (also called "Multiple Media") that collectively comprise a set of one or more storage devices containing machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A "Storage Device" is any tangible device capable of holding and storing instructions for use by a computer processor. Computer-readable storage media may be, but are not limited to, electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, mechanical storage media, or any preferred combination thereof. Some known types of storage devices that include these media include: diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on the main surface of a disk), or any preferred combination of the foregoing. When the term "computer-readable storage medium" is used in this disclosure, it shall not be construed as storage in the form of transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses passing through optical fiber cables, electrical signals communicated through wires, and / or other transmission media. As those skilled in the art will understand, data is typically moved at several intermittent points in the normal operation of a storage device, such as during access, defragmentation, or garbage collection; however, data is not transient while it is stored, and therefore a storage device is not transient.

[0059] The computing environment 900 includes an example of an environment for executing at least a portion of the computer code involved in performing the method of the present invention, such as the improved voice data generation code 926. In addition to the code 926, the computing environment 900 includes, for example, a computer 901, a wide area network (WAN) 902, an end user device (EUD) 903, a remote server 904, a public cloud 905, and a private cloud 906. In this embodiment, the computer 901 includes a processor set 910 (including a processing circuit configuration 920 and a cache 921), a communication fabric 911, volatile memory 912, persistent storage 913 (including the operating system 922 and code 926 identified above), a peripheral device set 914, a user interface (UI), a device set 923, storage 924, an Internet of Things (IoT) sensor set 925, and a network module 915. The remote server 904 includes a remote database 930. Public Cloud 905 includes Gateway 940, Cloud Orchestration Module 941, Host Physical Machine Set 942, Virtual Machine Set 943, and Container Set 944.

[0060] Computer 901 may take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device currently known or to be developed in the future, capable of executing programs, accessing networks, or querying databases such as the remote database 930. As is well understood in the field of computer technology, and depending on the technology, the execution of a computer implementation method may be distributed among multiple computers and / or multiple locations. On the other hand, in this presentation concerning the computing environment 900, in order to make the presentation as concise as possible, the detailed discussion focuses on a single computer, specifically computer 901. Computer 901 may be located in the cloud, although it is not shown in the cloud in Figure 9. On the other hand, computer 901 is not required to be located in the cloud, except to any extent that can be definitively shown.

[0061] The processor set 910 includes one or more computer processors of any type currently known or to be developed in the future. The processing circuit configuration 920 may be distributed across multiple packages, for example, multiple coordinated integrated circuit chips. The processing circuit configuration 920 may implement multiple processor threads and / or multiple processor cores. The cache 921 is memory located within the processor chip package and is typically used for data or code that should be available for high-speed access by threads or cores running on the processor set 910. The cache memory is typically organized into multiple levels depending on its relative proximity to the processing circuit configuration. Alternatively, some or all of the cache for the processor set may be located "off-chip". In some computing environments, the processor set 910 may operate using qubits and be designed to perform quantum computing.

[0062] Computer-readable program instructions are typically loaded into computer 901, causing the processor set 910 of computer 901 to execute a series of operational steps, thereby enabling the computer implementation method, and as a result, the instructions thus executed instantiate the method specified in the flowcharts and / or narrative descriptions of the computer implementation method contained herein (collectively referred to as the "Methods of the Invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as the cache 921 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 910 to control and direct the execution of the Methods of the Invention. In the computing environment 900, at least some of the instructions for executing the Methods of the Invention may be stored in code 926 in persistent storage 913.

[0063] The communication fabric 911 is a signal conduction path that enables various components of the computer 901 to communicate with one another. Typically, this fabric is made up of switches and conductive paths, such as buses, bridges, physical input / output ports, and similar components. Other types of signal communication paths, such as optical fiber communication paths and / or wireless communication paths, may be used.

[0064] Volatile memory 912 is any type of volatile memory currently known or to be developed in the future. Examples include dynamic RAM or static RAM. Typically, volatile memory 912 is characterized by random access, but this is not required unless explicitly stated. In computer 901, volatile memory 912 is located in a single package and resides inside computer 901, but alternatively or additionally, volatile memory may be distributed across multiple packages and / or located externally to computer 901.

[0065] The persistent storage 913 is any form of non-volatile storage for a computer, currently known or to be developed in the future. The non-volatility of this storage means that the stored data is maintained whether or not power is supplied to the computer 901 and / or directly to the persistent storage 913. The persistent storage 913 may be ROM, but typically at least a portion of the persistent storage allows for writing, erasing, and rewriting of data. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 922 may take several forms, such as various known proprietary operating systems or open-source portable operating system interface (CSI) type operating systems employing a kernel. The code included in code 926 typically includes at least a portion of computer code involved in performing the method of the present invention.

[0066] The peripheral device set 914 includes a set of peripheral devices for the computer 901. Data communication connections between the computer 901's peripheral devices and other components may be implemented in various ways, such as Bluetooth connections, Near-Field Communication (NFC) connections, connections made by cables (such as universal serial bus (USB) type cables), insert-type connections (e.g., secure digital (SD) cards), connections made via local area communication networks, and even connections made via wide area networks such as the internet. In various embodiments, the UI device set 923 may include components such as a display screen, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage 924 is external storage such as an external hard drive or insertable storage such as an SD card. Storage 924 may be persistent and / or volatile. In some embodiments, storage 924 may take the form of a quantum computing memory device for storing data in the form of qubits. In embodiments where computer 901 needs to have a large amount of storage (for example, computer 901 locally stores and manages a large database), this storage may be provided by peripheral storage devices designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 925 consists of sensors that may be used in Internet of Things applications. For example, one sensor may be a thermometer and another may be a motion detector.

[0067] The network module 915 is a collection of computer software, hardware, and firmware that enables computer 901 to communicate with other computers via the WAN 902. The network module 915 may include hardware such as a modem or Wi-Fi® signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of the network module 915 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of the network module 915 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the method of the present invention can typically be downloaded from an external computer or external storage device to computer 901 via a network adapter card or network interface included in the network module 915.

[0068] WAN902 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any currently known or future-developed technology for transmitting computer data. In some embodiments, WAN902 may be replaced and / or supplemented by a local area network (LAN), such as a Wi-Fi network, designed to transmit data between devices located in a local area. WANs and / or LANs typically include computer hardware such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0069] The end-user device 903 is any computer system used and controlled by an end-user (e.g., a customer of the company operating computer 901), and may take any of the forms discussed above in relation to computer 901. The EUD903 typically receives useful and valuable data from the operation of computer 901. For example, in a hypothetical case where computer 901 is designed to provide recommendations to the end-user, these recommendations would typically be communicated from the computer 901's network module 915 to the EUD903 via WAN902. In this way, the EUD903 may display or otherwise present the recommendations to the end-user. In some embodiments, the EUD903 may be a client device such as a thin client, heavy client, mainframe computer, desktop computer, and the like.

[0070] The remote server 904 is any computer system that provides at least some data and / or functions to computer 901. The remote server 904 may be controlled and used by the same entity that operates computer 901. The remote server 904 represents a machine that collects and stores useful and beneficial data for use by other computers, such as computer 901. For example, in a hypothetical case where computer 901 is designed and programmed to provide recommendations based on historical data, this historical data may be provided to computer 901 from the remote database 930 of the remote server 904.

[0071] Public Cloud 905 is any computer system available for use by multiple entities, providing on-demand availability of computer system resources and / or other computer functions, particularly data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct and active management of the computing resources of Public Cloud 905 is performed by the computer hardware and / or software of the Cloud Orchestration Module 941. The computing resources provided by Public Cloud 905 are typically implemented by virtual computing environments running on various computers that make up the computers in Public Cloud 905 and / or the host physical machine set 942, which is a collection of physical computers on which it is available. Virtual computing environments (VCEs) typically take the form of virtual machines from the virtual machine set 943 and / or containers from the container set 944. These VCEs may be stored as images and are understood to be transportable either as images or after instantiation of VCEs, among and between hosts of various physical machines. The cloud orchestration module 941 manages image transfer and storage, deploys new VCE instances, and manages the active instantiation of VCE deployments. The gateway 940 is a collection of computer software, hardware, and firmware that enables the public cloud 905 to communicate over the WAN 902.

[0072] Some further explanations of VCE are provided below. A VCE can be stored as an “image.” A new active instance of a VCE can be instantiated from an image. Two well-known types of VCE are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature in which the kernel allows for the existence of multiple isolated user-space instances called containers. These isolated user-space instances typically behave like actual computers in terms of the programs running within them. Computer programs running on a normal operating system can utilize all the resources of that computer, including connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running inside a container can only use the contents of the container and the devices allocated to that container; this feature is known as containerization.

[0073] Private Cloud 906 is similar to Public Cloud 905, except that its computing resources are available for use by a single enterprise only. While Private Cloud 906 is illustrated as communicating with WAN 902, in other embodiments, the private cloud may be completely isolated from the internet and accessible only through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types), often implemented by different vendors. Each of the multiple clouds remains a separate, discrete entity, but the larger hybrid cloud architecture is coupled together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple configuration clouds. In this embodiment, both Public Cloud 905 and Private Cloud 906 are part of a larger hybrid cloud.

[0074] In computing environment 900, computer 901 is shown as being connected to the Internet (see WAN 902). However, in many embodiments of the present invention, computer 901 is isolated from communications over a communication network and is not connected to the Internet, thereby operating as a standalone computer. In these embodiments, the network module 915 of computer 901 may not be necessary, or even undesirable, to ensure isolation and prevent external communications from entering computer 901. Standalone computer embodiments may be advantageous because they are typically more secure in at least some applications of the present invention. In other embodiments, computer 901 is connected to a secure WAN or secure LAN instead of WAN 902 and / or the Internet. In these network-connected (i.e., non-standalone) embodiments, system designers may wish to implement appropriate security measures, currently known or to be developed in the future, to reduce the risk of security breaches occurring due to incoming network communications.

[0075] The terminology used herein is solely for the purpose of describing specific embodiments and is not intended to limit the invention. Where used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," where used herein, specify the presence of the described features, steps, operations, elements, and / or components, but do not preclude the presence or addition of other features, steps, operations, elements, components, and / or groups thereof.

[0076] The descriptions of various embodiments of the present invention are presented for illustrative purposes but are not intended to be exhaustive or to limit oneself to the disclosed embodiments. Many modifications and variations will become apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best describe the principles of the embodiments, their practical applications, or technical improvements to the art found in the market, or to enable other those skilled in the art to understand the embodiments disclosed herein.

Claims

1. Memory configured to store program instructions; and The memory is operablely coupled to the aforementioned memory and executes the program instructions, Implement one or more artificial intelligence techniques in relation to one or more speech synthesis tasks; At least one data sequence, including one or more phonetic data and prosodic data, is generated in multiple consecutive parts by processing at least one previously generated data sequence using the one or more artificial intelligence methods; and Audio data corresponding to at least a portion of the data sequence is generated by processing at least a portion of the data sequence using at least one artificial intelligence-based speech synthesis model. Processor, A system equipped with these features.

2. The system according to claim 1, wherein generating audio data includes generating one or more audio waveforms corresponding to at least a portion of the data sequence.

3. The system according to claim 1, wherein implementing one or more artificial intelligence techniques includes combining at least one language model (LM), at least one text-to-speech front-end (TTS-FE) model, and at least one text-to-speech prosody (TTS-P) model.

4. The system according to claim 1, wherein implementing one or more artificial intelligence techniques includes using a language model-text-speech conversion algorithm adapter to convert at least a portion of the output from at least one LM into phonetic and prosodic data to be used as input to one or more of at least one TTS-FE model and at least one TTS-P model.

5. The system according to claim 4, wherein the implementation of one or more artificial intelligence techniques includes using a language model-text-speech algorithm adapter, in conjunction with generating one or more hierarchical prosodic control (HPC) features derived from one or more statistical measurements taken over one or more hierarchical time intervals and normalized to represent one or more speaker-independent features, wherein at least a portion of the one or more HPC features includes at least one of one or more overall evaluations and one or more local evaluations relating to at least one of phoneme utterance rate, pitch, and volume.

6. Implementing one or more artificial intelligence methods is Modifying at least a portion of at least one LM using one or more phonetic data items; and Extracting one or more phonetic data items from at least one training text dataset using at least one TTS-FE model, The system according to claim 1, including the following:

7. Implementing one or more artificial intelligence methods is Modifying at least a portion of at least one LM using one or more prosodic data items; and Extracting one or more prosodic data items from at least one training text dataset using at least one TTS-P model, The system according to claim 1, including the following:

8. The system according to claim 1, wherein implementing one or more artificial intelligence techniques includes modifying at least a portion of at least one LM using one or more speech data items in relation to at least one automatic speech recognition technique and one or more prosodic feature extraction techniques.

9. The system according to claim 1, wherein generating at least one data sequence includes generating one or more phonemes and generating one or more prosodic information items in conjunction with one or more output words related to the at least one previously generated data sequence.

10. The system according to claim 9, wherein generating one or more prosodic information items comprises generating one or more HPC features derived from one or more statistical measurement results obtained over one or more hierarchical time intervals and normalized to represent one or more speaker-independent features, wherein at least a portion of the one or more HPC features includes at least one of one or more overall evaluations and one or more local evaluations relating to at least one of phoneme utterance rate, pitch, and volume.

11. The system according to claim 1, wherein generating at least one data sequence containing one or more phonetic data and prosodic data in a plurality of consecutive portions includes generating one or more phonetic vectors and one or more prosodic vectors at once to generate the at least one data sequence.

12. The processor is further operablely coupled with the memory and executes the program instructions. At least a portion of the generated audio data is used to automatically train at least a portion of the one or more artificial intelligence methods. The system according to claim 1.

13. A computer program product comprising a computer-readable storage medium in which program instructions are embodied, wherein the program instructions are executable by a computing device, and the computing device, Implementing one or more artificial intelligence techniques in relation to one or more speech synthesis tasks; At least one data sequence, including one or more phonetic and prosodic data, is generated in multiple consecutive parts by processing at least one previously generated data sequence using the one or more artificial intelligence methods; and Audio data corresponding to at least a portion of the data sequence is generated by processing at least a portion of the data sequence using at least one artificial intelligence-based speech synthesis model. Computer program products.

14. The computer program product according to claim 13, wherein implementing one or more artificial intelligence techniques includes combining at least one LM, at least one TTS-FE model, and at least one TTS-P model.

15. The computer program product according to claim 13, wherein the implementation of one or more artificial intelligence techniques includes using a language model-text-speech conversion algorithm adapter to convert at least a portion of the output from at least one LM into phonetic and prosodic data to be used as input to one or more of at least one TTS-FE model and at least one TTS-P model.

16. The computer program product according to claim 15, wherein the implementation of one or more artificial intelligence techniques includes using a language model-text-speech algorithm adapter, in conjunction with generating one or more HPC features derived from one or more statistical measurement results obtained over one or more hierarchical time intervals and normalized to represent one or more speaker-independent features, wherein at least a portion of the one or more HPC features includes at least one of one or more overall evaluations and one or more local evaluations relating to at least one of speech utterance rate, pitch, and volume.

17. The stage of implementing one or more artificial intelligence techniques in relation to one or more speech synthesis tasks; A step of generating at least one data sequence, including one or more phonetic data and prosodic data, in multiple consecutive parts by processing at least one previously generated data sequence using the one or more artificial intelligence methods; and A step of generating audio data corresponding to at least a portion of the data sequence by processing the at least portion of the data sequence using at least one artificial intelligence-based speech synthesis model. A computer implementation method comprising, The above method is performed by at least one computing device. Computer implementation method.

18. The computer implementation method according to claim 17, wherein implementing one or more artificial intelligence techniques includes combining at least one LM, at least one TTS-FE model, and at least one TTS-P model.

19. The computer implementation method according to claim 17, wherein implementing one or more artificial intelligence techniques includes using a language model-text-speech conversion algorithm adapter to convert at least a portion of the output from at least one LM into phonetic and prosodic data to be used as input to one or more of at least one TTS-FE model and at least one TTS-P model.

20. The computer implementation method according to claim 19, comprising using a language model-text-speech algorithm adapter, in conjunction with generating one or more HPC features derived from one or more statistical measurement results obtained over one or more hierarchical time intervals and normalized to represent one or more speaker-independent features, wherein at least a portion of the one or more HPC features includes at least one of one or more overall evaluations and one or more local evaluations relating to at least one of phoneme utterance rate, pitch, and volume.