Real Time Text Context Encoding for Speaker Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Real Time Text (RTT) technologies for mobile devices lack the ability to effectively convey contextual information, such as speaker identity, volume, and tone, which are essential for accurately representing the dynamic qualities of speech in multi-speaker audio signals, limiting the richness of text-based communication experiences.

Innovation Solution

The implementation of a Real Time Text system that utilizes voice recognition algorithms to identify speakers and generate contextual information, which is then used to determine graphical rendition selections for text display, allowing for dynamic text output and improved contextual representation through the integration of ITU T.140 and UTF-8 encoding, enabling the transmission and display of text with additional context over networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If basic RTT text transmission is used, then real-time text communication is achieved, but contextual information such as speaker identity and speech qualities is lost

Engineering Contradiction:
Improvecontextual informationVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the audio signal processing into distinct functional modules: voice activity detection module, speaker identification module, and speech quality analysis module. Each module processes specific aspects of the audio signal and generates separate contextual attributes that are then integrated into the RTT transmission stream, allowing comprehensive context capture without overwhelming system complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer between audio input and text output that extracts and tags contextual information. This intermediary layer analyzes audio characteristics and embeds metadata (speaker IDs, volume levels, tone indicators) into the text stream, preserving contextual information without requiring complete system redesign

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If voice recognition algorithms are implemented to identify speakers, then speaker identification capability is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary voice activity detection and speaker diarization before full text transcription. By pre-identifying speakers and their turn-taking patterns in the audio stream, the system prepares contextual metadata in advance, reducing real-time processing delays while maintaining accurate speaker identification

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial processing by focusing voice recognition algorithms only on active speech segments identified by voice activity detection, rather than processing the entire audio stream continuously. This selective approach reduces computational load and processing time while maintaining speaker identification accuracy during actual speech periods

Inventive Principle:
Principle #16Partial or excessive action

3Loss of information

If graphical rendition selections are added to convey speech qualities, then text display richness is improved, but data transmission volume increases

Engineering Contradiction:
Improvespeech quality informationVSAvoiddata transmission volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent applies graphical rendition selections locally to specific text segments or words that correspond to particular speech events, rather than applying uniform formatting to entire messages. This targeted approach conveys speech quality information (volume, tone, emphasis) only where relevant, enriching text display while minimizing additional data transmission

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent encodes speech quality attributes as compact parameter tags within the text stream rather than transmitting separate detailed audio descriptions. By converting analog speech qualities into discrete categorical parameters (e.g., volume levels, tone types), the system preserves speech quality information efficiently with minimal increase in data transmission volume

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2584745B1Determining and conveying contextual information for real time text
Publication Date: 2018.03.07 BLACKBERRY LTD
  • EP2584745B1 patent drawingFigure 1
  • EP2584745B1 patent drawingFigure 2
  • EP2584745B1 patent drawingFigure 3~4

AI summary

Aspects relate to machine recognition of human voices in live or recorded audio content, and delivering text derived from such live or recorded content as real time text, with contextual information derived from characteristics of the audio. For example, volume information can be encoded as larger and smaller font sizes. Speaker changes can be detected and indicated through text additions, or color changes to the font. A variety of other context information can be detected and encoded in graphical rendition commands available through RTT, or by extending the information provided with RTT packets, and processing that extended information accordingly for modifying the display of the RTT text content.