Transcript And Caption Integration for Continuous Chat Conversations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing communication systems fail to seamlessly integrate audio/video conversations with chat sessions, leading to fragmented discussions, inefficiencies, and excessive resource usage due to manual switching between modalities.

Innovation Solution

A system that allows users to quote and reply directly from transcriptions or captions to chat messages, using a multi-dimensional data structure for context tracking and automated integration, enabling efficient and context-aware messaging.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users manually switch between transcription/caption mode and chat mode to share information, then information sharing between modes is enabled, but user productivity decreases and conversation continuity is lost

Engineering Contradiction:
Improveease of information sharingVSAvoiduser productivity
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent merges transcription/caption functionality with chat messaging functionality into a unified interface. Users can directly quote and reply to transcript/caption content within the chat interface without switching modes, combining what were previously separate operations into a single integrated system.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system pre-processes and structures transcription and caption data with embedded metadata (timestamps, speaker information, contextual markers) that enables direct quoting and replying. This preliminary structuring allows users to immediately interact with transcript content in chat without manual copying or switching operations.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If users interrupt current collaboration activities to pursue discussions in chat session, then chat discussions can be pursued, but conversation continuity is fragmented and productivity is reduced

Engineering Contradiction:
Improvecommunication flexibilityVSAvoidconversation continuity
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent combines audio/video conversation streams with chat message streams into a single unified conversation thread. Transcript excerpts from audio/video are directly embeddable in chat messages, and chat messages are synchronized with the conversation timeline, maintaining a continuous view of all communication modalities.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces a unified conversation context as an intermediary layer that mediates between audio/video transmissions and chat messages. This context layer correlates timestamps, speakers, and content across modalities, allowing users to reference and respond to any communication type without breaking the overall conversation flow.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If content is re-sent when participants miss salient points or cues, then information can be reinforced, but network and computing resources are used inefficiently

Engineering Contradiction:
Improveinformation deliveryVSAvoidcomputing resource efficiency
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system provides real-time feedback mechanisms including transcript generation during audio/video calls, automated captioning, and contextual suggestions in chat. Users receive immediate feedback about what is being said and can reference specific transcript portions, reducing the need to re-send content and optimizing resource utilization.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP4356567B1Transition to messaging from transcription and captioning
Publication Date: 2025.09.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4356567B1 patent drawingFigure 1
  • EP4356567B1 patent drawingFigure 2A
  • EP4356567B1 patent drawingFigure 2B

AI summary

The techniques disclosed herein improve existing systems by controlling a data processing system for generating messages associated with a communication session. A first UI is rendered on a user device that includes a text-based transcription or caption of dialogue being communicated between users of the communication session. in response to receiving a selection of a portion of the transcription or caption for corresponding via a messaging function of the communication session, rendering a second UI including the selected portion and current messages exchanged between the users of the communication session. The selected portion is rendered along with subsequent messages exchanged between the users of the communication session.