Speech Synthesis for Whispered Audio in Open Offices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Open office environments face challenges in reducing noise interference, particularly speech noise, which diminishes conversation quality and productivity, as traditional microphones are ineffective in noisy settings and quieting voices can make speech less intelligible.
Innovation Solution
A system utilizing proximity-sensitive microphones and speech synthesis technology to convert low-quality, barely audible speech into clear, synthesized speech, transmitted over a telephony network, ensuring high intelligibility and minimizing ambient noise interference.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional microphones are used in noisy open office environments, then speech can be captured, but the speech quality is poor and intelligibility is reduced due to ambient noise interference
Solution Approach 1:
The system extracts the desired speech signal from the noisy environment by using proximity-sensitive microphones to capture only the speech from the immediate vicinity, effectively separating the useful speech from the harmful ambient noise in the open office environment
Solution Approach 2:
The system introduces speech synthesis technology as an intermediary that converts the captured low-quality speech into high-quality synthesized speech, mediating between the noisy capture and the need for clear intelligibility in the telephony network transmission
2Object-affected harmful factors
If employees lower their voice volume to reduce noise distractions to coworkers, then the work environment becomes quieter, but the speech becomes less audible and intelligible
Solution Approach 1:
The system changes the parameters of the speech signal by using speech synthesis to convert low-volume whispered speech into high-volume clear speech, maintaining the quiet workspace while ensuring the synthesized speech is fully intelligible to the intended recipient
Solution Approach 2:
Speech synthesis acts as an intermediary that bridges the gap between the low-volume quiet speech and the need for clear intelligibility, allowing employees to whisper quietly while the system generates clear audible speech for communication
3Measurement precision
If speech synthesis is used to convert low-quality speech to high-quality synthesized speech, then speech intelligibility is improved, but the system complexity increases
Solution Approach 1:
The system uses a multi-functional integrated approach where proximity-sensitive microphones, speech recognition, and speech synthesis work together as a unified telephony solution, allowing the same system to handle both noise reduction and speech quality enhancement across multiple communication scenarios
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system effectively reduces noise distractions in shared workspaces by converting whispered or murmured speech into clear, audible signals, enhancing conversation quality and maintaining a productive work environment without disrupting coworkers.
Implementation Method 1
receiving a first audio signal representing a first utterance of a human end-user
Implementation Method 2
converting the first data file to a second audio signal via implementation of a speech synthesizer
Data Source
AI summary
A method and system of reducing noise associated with telephony-based activities occurring in shared workspaces is provided. An end-user may lower their own voice to a whisper or other less audible or intelligible utterances and submit such low-quality audio signals to an automated speech recognition system via a microphone. The words identified by the automated speech recognition system are provided to a speech synthesizer, and a synthesized audio signal is created artificially that carries the content of the original human-produced utterances. The synthesized audio signal is significantly more audible and intelligible than the original audio signal. The method allows customer support agents to speak at barely audible levels yet be heard clearly by their customers.


