Text-to-Speech Audio Mixing for Network Streaming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current network streaming methods require pre-recording voice segments, limiting sound richness and requiring multiple tools for adjustment, making it difficult to manage large volumes of voice recordings.
Innovation Solution
A signal transmission method for network video streaming that generates a text voiced audio signal from a database and mixes it with a play signal, allowing for real-time transmission and playback, enabling users to create diverse mixed content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If pre-recording voice segments are used for network streaming, then the mixing function can be enabled to increase content richness, but the sound effect is limited by the streamer's own audio quality and requires multiple tools to adjust
Solution Approach 1:
The patent combines multiple audio processing functions (voice recording, text-to-speech conversion, audio mixing, and effect processing) into a single integrated system. The server端 integrates the text voice database, TTS generation capability, and audio mixing function, allowing the streamer to input text and automatically receive professionally processed audio without needing separate tools for each function.
Solution Approach 2:
The patent introduces a server as an intermediary between the streamer and the audio processing functions. The server hosts the text voice database, performs TTS conversion, and provides processed audio signals to the streamer's client. This intermediary handles the complex audio processing tasks centrally, freeing the streamer from managing multiple local tools while still accessing rich content creation capabilities.
2Adaptability or versatility
If pre-recording voice segments are used, then mixing can be achieved, but it is difficult to manage a large volume of voice recordings
Solution Approach 1:
The patent replaces physical voice recording files with text-based content stored in a text voice database. Instead of managing large volumes of audio files, the system stores and manages compact text representations. The TTS system generates audio signals on-demand from these text records, eliminating the need to store, organize, and manage extensive audio file libraries while maintaining the ability to create diverse mixed content.
3Manufacturing precision
If multiple tools and software are used to adjust sound output, then audio quality can be improved, but the process becomes time-consuming and complex
Solution Approach 1:
The patent performs audio processing adjustments in advance during the content creation phase. The server processes audio signals with appropriate effects, mixing, and quality optimization before transmitting them to the streamer's client. This preliminary processing ensures high audio quality is achieved automatically without requiring the streamer to spend time making manual adjustments during the streaming process.
Solution Approach 2:
The system provides self-service audio processing where the server automatically performs TTS conversion, audio mixing, and effect application based on text input. The streamer simply needs to input text and select desired effects from predefined options, while the server handles all the complex audio processing tasks automatically, eliminating the need for manual adjustment of multiple audio parameters.
Data Source
AI summary
An audio mixing method for network streaming includes the steps of establishing a network streaming connection between the first end and the second end, generating a text voiced audio signal based on a trigger signal by the first end, mixing a play signal and the text voiced audio signal into a play signal with the text voiced audio signal, in the state of the network streaming connection, transmitting the play signal with text voiced audio signal to the second end, and the play signal with text voiced audio signal is played by the second end.
