Local Captioning with Audio Priority Handling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional captioning systems for communication sessions require interaction with remote servers, raising privacy concerns and complicating transcriptions due to crosstalk and increased network traffic.
Innovation Solution
The system provides textual representations for a communication session by receiving audio inputs with priority levels, determining the highest priority audio input, and generating a textual representation for display on the device, thereby reducing reliance on remote servers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional captioning systems use remote servers for transcription, then transcription capability is provided, but user privacy is compromised and network traffic increases
Solution Approach 1:
The patent extracts the transcription processing function from remote servers and relocates it to local electronic devices. Each device performs speech-to-text conversion independently using local processing capabilities, eliminating the need to transmit audio data over the network and thereby protecting user privacy while maintaining transcription functionality.
Solution Approach 2:
The system enables each electronic device to transcribe its own audio inputs locally without relying on external server services. By implementing on-device speech recognition and text generation capabilities, the system makes each device self-sufficient for captioning tasks, reducing network dependency and privacy risks.
2Loss of information
If conventional captioning systems process all audio inputs, then complete transcriptions are generated, but computational requirements and network traffic increase due to crosstalk
Solution Approach 1:
The patent implements priority-based processing where different audio inputs are assigned different priority levels based on their relevance to the current speaker or context. High-priority audio streams (e.g., active speaker) receive full processing resources, while low-priority streams (e.g., background noise or non-speaking participants) receive reduced or no processing, thereby reducing computational load while maintaining essential transcription quality.
Solution Approach 2:
Instead of processing all audio inputs equally, the system applies partial processing only to the necessary portions of audio streams. By identifying and processing only the highest priority audio input at any given time, the system achieves sufficient transcription coverage without the excessive computational overhead of processing all concurrent audio sources.
3Loss of information
If multiple audio inputs are processed simultaneously, then all user contributions are captured, but transcription accuracy decreases due to crosstalk
Solution Approach 1:
The system performs preliminary classification of audio inputs into different priority levels before transcription processing. By pre-identifying which audio streams are most relevant (e.g., detecting active speakers or important participants) and assigning them higher priorities, the system prepares the processing sequence in advance, ensuring that accurate transcriptions are generated for critical inputs while managing the complexity of multi-user environments.
Data Source
AI summary
Systems and processes for providing textual representations for a communication session are provided. For example, at least one audio input is received at an electronic device, wherein each audio input of the at least one audio input is associated with a respective priority level. A priority level of an audio input detected at a microphone of the electronic device is determined, wherein a highest priority level among the determined priority level and each received priority level corresponding to the at least one audio input is identified. A textual representation of a respective audio input corresponding to the identified highest priority level is obtained, wherein the obtained textual representation is displayed on a display of the electronic device.


