Multi-Microphone Voice Synthesis for Real-Time Speech-to-Text
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice processing systems face challenges in converting speech voices from multiple users into text in real-time with high accuracy due to increased processing load, leading to impaired real-time text display during conversations.
Innovation Solution
A voice processing system that synthesizes input voices from multiple audio devices into a single voice and individually converts each voice into text using a conversion processing unit, enabling real-time text conversion and improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If individual speech conversion to text is performed for each audio device, then text conversion accuracy is improved, but processing load increases and real-time display is impaired
Solution Approach 1:
The patent combines multiple individual speech inputs into a single synthesized speech stream for text conversion. The synthesis processing unit merges the plurality of input speeches into one unified speech signal, which is then converted to text as a whole. This approach maintains real-time display capability while preserving the accuracy benefits of individual speech processing through the use of speech separation technology in the conversion unit.
2Measurement precision
If multiple individual speeches are converted to text separately, then text conversion accuracy is improved, but processing time increases
Solution Approach 1:
The patent performs speech synthesis as a preliminary action before text conversion. By pre-combining multiple speech inputs into a single synthesized speech stream, the system prepares the data in an optimized format that reduces the computational burden during the actual text conversion process. This preliminary processing step enables faster real-time conversion while maintaining accuracy through subsequent speech separation in the conversion unit.
Data Source
AI summary
A voice processing apparatus includes an acquisition processing unit that acquires a plurality of input voices input to microphones individually included in a plurality of audio devices, a synthesis processing unit that synthesizes the plurality of input voices acquired by the acquisition processing unit into a single synthesized voice, and an output processing unit that outputs the plurality of input voices and the synthesized voice to a conference server that converts the synthesized voice synthesized by the synthesis processing unit into text and individually converts each of the plurality of input voices into a piece of text among a plurality of pieces of text.


