Streaming CNN Text-to-Speech with Tensor Reuse for Real-Time Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Intelligent Personal Assistant (IPA) systems are limited in their responsiveness, requiring non-satisfactory time for providing machine-generated utterances in response to user spoken utterances, especially when real-time interaction is desired.

Innovation Solution

Employing a Convolutional Neural Network (CNN) configured for real-time text-to-speech processing, utilizing a streaming buffer to reuse tensor data across iterations and optimize memory usage, allowing for real-time generation and provision of machine-generated utterance portions without compromising quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If conventional IPA systems generate complete machine-generated utterances before providing them to users, then the quality of the utterance is maintained, but the responsiveness time is non-satisfactory and delayed

Engineering Contradiction:
Improveresponsiveness timeVSAvoidutterance quality
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent segments the machine-generated utterance into multiple portions and transmits them sequentially to the user device. The TTS server generates and sends first portions of the utterance before completing the entire utterance generation, allowing the user device to start playing back portions immediately while subsequent portions are still being generated. This segmentation enables real-time streaming without compromising the overall quality of the complete utterance.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary actions by pre-processing the input text into TTS input data and preparing the neural network model before actual utterance generation. The system performs text normalization, tokenization, and feature extraction in advance, so that when the TTS generation starts, the processing pipeline is already optimized and ready to generate portions quickly and continuously.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If real-time generation of waveform segments is implemented, then responsiveness is improved, but computational operations and memory requirements increase

Engineering Contradiction:
Improvereal-time generation speedVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts and reuses tensor data from previous iterations of the neural network that can be applied to subsequent waveform segment generation. Instead of performing complete forward propagation from scratch for each segment, the system identifies and reuses intermediate tensor computations, significantly reducing the computational operations required for real-time generation while maintaining output quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates and maintains copies of relevant tensor data in memory that can be quickly accessed and reused across iterations. By storing intermediate computation results and making them available for subsequent segments, the system avoids redundant calculations and reduces the computational complexity of real-time TTS generation.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12603081B2Method and server for a text-to-speech processing
Publication Date: 2026.04.14 Y E HUB ARMENIA LLC
  • US12603081B2 patent drawing
  • US12603081B2 patent drawing
  • US12603081B2 patent drawing

AI summary

Methods and servers for processing a textual input for generating an audio output are disclosed. The audio output is a sequence of waveform segments generated in real-time by a trained Convolutional Neural Network. The method includes, at a given iteration, generating a given waveform segment which includes storing first tensor data computed by a first hidden layer during the given iteration, and where the first tensor data has tensor-chunk data. The tensor-chunk data is used during the given iteration for generating the given waveform segment and is to be used during a next iteration for generating a next waveform segment. The method includes, at the next iteration, generating the next waveform segment, which comprises storing second tensor data computed by the first hidden layer during the next iteration. The second tensor data excludes redundant tensor-chunk data that is identical to the tensor-chunk data from the first tensor data.