Text-to-Speech Audio Mixing for Network Streaming

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current network streaming methods require pre-recording voice segments, limiting sound richness and requiring multiple tools for adjustment, making it difficult to manage large volumes of voice recordings.

Innovation Solution

A signal transmission method for network video streaming that generates a text voiced audio signal from a database and mixes it with a play signal, allowing for real-time transmission and playback, enabling users to create diverse mixed content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If pre-recording voice segments are used for network streaming, then the mixing function can be enabled to increase content richness, but the sound effect is limited by the streamer's own audio quality and requires multiple tools to adjust

Engineering Contradiction:
Improvecontent richnessVSAvoidnumber of tools required
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent combines multiple audio processing functions (voice recording, text-to-speech conversion, audio mixing, and effect processing) into a single integrated system. The server端 integrates the text voice database, TTS generation capability, and audio mixing function, allowing the streamer to input text and automatically receive professionally processed audio without needing separate tools for each function.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a server as an intermediary between the streamer and the audio processing functions. The server hosts the text voice database, performs TTS conversion, and provides processed audio signals to the streamer's client. This intermediary handles the complex audio processing tasks centrally, freeing the streamer from managing multiple local tools while still accessing rich content creation capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If pre-recording voice segments are used, then mixing can be achieved, but it is difficult to manage a large volume of voice recordings

Engineering Contradiction:
Improvemixing capabilityVSAvoidmanagement of voice recordings
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent replaces physical voice recording files with text-based content stored in a text voice database. Instead of managing large volumes of audio files, the system stores and manages compact text representations. The TTS system generates audio signals on-demand from these text records, eliminating the need to store, organize, and manage extensive audio file libraries while maintaining the ability to create diverse mixed content.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If multiple tools and software are used to adjust sound output, then audio quality can be improved, but the process becomes time-consuming and complex

Engineering Contradiction:
Improveaudio qualityVSAvoidadjustment time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs audio processing adjustments in advance during the content creation phase. The server processes audio signals with appropriate effects, mixing, and quality optimization before transmitting them to the streamer's client. This preliminary processing ensures high audio quality is achieved automatically without requiring the streamer to spend time making manual adjustments during the streaming process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides self-service audio processing where the server automatically performs TTS conversion, audio mixing, and effect application based on text input. The streamer simply needs to input text and select desired effects from predefined options, while the server handles all the complex audio processing tasks automatically, eliminating the need for manual adjustment of multiple audio parameters.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12279016B2Audio mixing and signal transmission method for network streaming
Publication Date: 2025.04.15 AVERMEDIA TECH
  • US12279016B2 patent drawing

AI summary

An audio mixing method for network streaming includes the steps of establishing a network streaming connection between the first end and the second end, generating a text voiced audio signal based on a trigger signal by the first end, mixing a play signal and the text voiced audio signal into a play signal with the text voiced audio signal, in the state of the network streaming connection, transmitting the play signal with text voiced audio signal to the second end, and the play signal with text voiced audio signal is played by the second end.