Local Captioning with Audio Priority Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional captioning systems for communication sessions require interaction with remote servers, raising privacy concerns and complicating transcriptions due to crosstalk and increased network traffic.

Innovation Solution

The system provides textual representations for a communication session by receiving audio inputs with priority levels, determining the highest priority audio input, and generating a textual representation for display on the device, thereby reducing reliance on remote servers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional captioning systems use remote servers for transcription, then transcription capability is provided, but user privacy is compromised and network traffic increases

Engineering Contradiction:
Improvetranscription capabilityVSAvoidprivacy concern
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts the transcription processing function from remote servers and relocates it to local electronic devices. Each device performs speech-to-text conversion independently using local processing capabilities, eliminating the need to transmit audio data over the network and thereby protecting user privacy while maintaining transcription functionality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system enables each electronic device to transcribe its own audio inputs locally without relying on external server services. By implementing on-device speech recognition and text generation capabilities, the system makes each device self-sufficient for captioning tasks, reducing network dependency and privacy risks.

Inventive Principle:
Principle #25Self-service

2Loss of information

If conventional captioning systems process all audio inputs, then complete transcriptions are generated, but computational requirements and network traffic increase due to crosstalk

Engineering Contradiction:
Improvetranscription completenessVSAvoidcomputational requirements
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent implements priority-based processing where different audio inputs are assigned different priority levels based on their relevance to the current speaker or context. High-priority audio streams (e.g., active speaker) receive full processing resources, while low-priority streams (e.g., background noise or non-speaking participants) receive reduced or no processing, thereby reducing computational load while maintaining essential transcription quality.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Instead of processing all audio inputs equally, the system applies partial processing only to the necessary portions of audio streams. By identifying and processing only the highest priority audio input at any given time, the system achieves sufficient transcription coverage without the excessive computational overhead of processing all concurrent audio sources.

Inventive Principle:
Principle #16Partial or excessive action

3Loss of information

If multiple audio inputs are processed simultaneously, then all user contributions are captured, but transcription accuracy decreases due to crosstalk

Engineering Contradiction:
Improvemulti-user coverageVSAvoidtranscription accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The system performs preliminary classification of audio inputs into different priority levels before transcription processing. By pre-identifying which audio streams are most relevant (e.g., detecting active speakers or important participants) and assigning them higher priorities, the system prepares the processing sequence in advance, ensuring that accurate transcriptions are generated for critical inputs while managing the complexity of multi-user environments.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250124927A1Providing textual representations for a communication session
Publication Date: 2025.04.17 APPLE INC
  • US20250124927A1 patent drawing
  • US20250124927A1 patent drawing
  • US20250124927A1 patent drawing

AI summary

Systems and processes for providing textual representations for a communication session are provided. For example, at least one audio input is received at an electronic device, wherein each audio input of the at least one audio input is associated with a respective priority level. A priority level of an audio input detected at a microphone of the electronic device is determined, wherein a highest priority level among the determined priority level and each received priority level corresponding to the at least one audio input is identified. A textual representation of a respective audio input corresponding to the identified highest priority level is obtained, wherein the obtained textual representation is displayed on a display of the electronic device.