Live Conversation Broadcasting With Speaker-Segmented Transcription

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for capturing and recording conversations are inefficient and prone to human error, leading to inaccurate information extraction and distraction from the conversation.

Innovation Solution

A system and method for capturing, processing, and rendering context-aware moment-associating elements, such as conversations, by segmenting audio-form conversations, assigning speaker labels, and generating synchronized transcriptions for live broadcasting.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional note-taking is performed during a conversation, then information can be recorded, but the note-taker is distracted from the conversation and accuracy decreases due to human error

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoidnote-taker distraction
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent replaces the mechanical system of manual note-taking with an automated speech recognition and processing system. The system captures conversation audio, transcribes it to text, segments it by speaker, and generates context-aware information automatically, eliminating the need for human note-takers and their associated distractions and errors.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The conversation recording system performs its own information extraction and processing functions without requiring external human intervention. The automated system transcribes, segments, and analyzes the conversation content independently, making the process self-sufficient and removing the burden from human participants.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual transcription is performed, then conversation content can be captured, but human error reduces accuracy and real-time processing is difficult

Engineering Contradiction:
Improvetranscription accuracyVSAvoidreal-time processing capability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual transcription with automated speech recognition technology. The system processes audio input through ASR engines that convert speech to text automatically, eliminating human transcription errors and enabling real-time processing of conversation content.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary processing of the audio signal by pre-segmenting it into speaker-specific segments before final transcription. This preliminary organization of audio data by speaker improves the accuracy and efficiency of the subsequent transcription process, enabling real-time accurate capture.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated speech processing is implemented, then transcription accuracy and real-time capability improve, but system complexity increases

Engineering Contradiction:
Improvereal-time transcription capabilityVSAvoidprocessing system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the conversation audio into distinct speaker-specific segments before processing. This segmentation approach simplifies the overall processing task by breaking down the complex multi-speaker audio stream into manageable, speaker-specific segments that can be processed independently and efficiently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces intermediate processing stages including audio pre-processing, speaker segmentation, and context-aware processing as mediators between the raw audio input and final transcription output. These intermediary steps organize and prepare the data in ways that simplify subsequent processing and improve overall system efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12406672B2Systems and methods for live broadcasting of context-aware transcription and/or other elements related to conversations and/or speeches
Publication Date: 2025.09.02 OTTER AI INC
  • US12406672B2 patent drawing
  • US12406672B2 patent drawing
  • US12406672B2 patent drawing

AI summary

Computer-implemented method and system for processing and broadcasting one or more moment-associating elements. For example, the computer-implemented method includes granting subscription permission to one or more subscribers; receiving the one or more moment-associating elements; transforming the one or more moment-associating elements into one or more pieces of moment-associating information; and transmitting at least one piece of the one or more pieces of moment-associating information to the one or more subscribers. In certain examples, transforming the one or more moment-associating elements includes: segmenting the one or more moment-associating elements into a plurality of moment-associating segments; assigning a segment speaker for each segment of the plurality of moment-associating segments; transcribing the plurality of moment-associating segments into a plurality of transcribed segments; and generating the one or more pieces of moment-associating information based the plurality of transcribed segments and the segment speaker assigned for each segment of the plurality of moment-associating segments.