Voice Recognition Transcript Segmentation for Meeting Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional transcription services struggle to accurately represent conversations involving multiple speakers, as they produce single flow text that mixes words spoken by different individuals, making it difficult to understand the context and identify who spoke specific parts.

Innovation Solution

The system uses voice recognition to attribute spoken text to individual users, creating a graphical user interface that separates text segments by speaker, combining words spoken within a predefined period or as part of a linguistic unit, even if interrupted by others, and allows filtering by user or keyword.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional transcription services use single flow text to transcribe conversations, then the transcription process is simple and fast, but the context understanding and speaker identification become difficult

Engineering Contradiction:
Improvetranscription speedVSAvoidcontext information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent segments the continuous speech transcript into discrete speaker-specific segments. Each segment is attributed to a specific speaker based on voice recognition, allowing the system to maintain transcription speed while organizing information to preserve context and speaker identity. This segmentation transforms the monolithic flow text into structured, attributable segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces voice recognition technology as an intermediary between the audio signal and the text transcript. This intermediary automatically identifies and attributes speech segments to specific speakers, enabling context preservation without manual intervention and maintaining efficient automated transcription processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If the system attributes text to individual speakers using voice recognition, then speaker identification and context understanding improve, but the system complexity increases

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The voice recognition system performs self-service by automatically identifying speakers and attributing segments without requiring manual configuration or complex setup. The system autonomously processes audio, recognizes voices, and structures transcripts, reducing the need for manual intervention while managing system complexity through automated intelligence.

Inventive Principle:
Principle #25Self-service

3Ease of manufacture

If the system combines words from different speakers in a single flow text, then the transcription is simple to generate, but the difficulty of following the conversation increases

Engineering Contradiction:
Improvetranscription generation easeVSAvoidconversation followability
Core Design Contradiction:
Ease of manufactureVSEase of operation

Solution Approach 1:

The patent applies segmentation to divide the mixed speaker transcript into distinct, attributable segments. Each segment is clearly marked with speaker identification, making the conversation easy to follow while maintaining automated generation simplicity. This segmentation preserves the ease of manufacturing through automation while dramatically improving ease of operation for readers.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3797413B1Use of voice recognition to generate a transcript of conversation(s)
Publication Date: 2024.07.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3797413B1 patent drawingFigure 1
  • EP3797413B1 patent drawingFigure 2
  • EP3797413B1 patent drawingFigure 3

AI summary

Examples described herein improve the way in which a transcript is generated and displayed so that the context of a conversation taking place during a meeting or another type of collaboration event can be understood by a person that reviews the transcript, e.g., reads or browses through the transcript. The techniques described herein use voice recognition to identify a user that is speaking during the meeting. Accordingly, when the speech of the user is converted to text for the transcript, the text can be attributed to the identified user. The techniques described herein further configure a graphical user interface layout, in which the transcript can be displayed. The graphical user interface layout enables users to better understand the context of a conversation that takes place during a meeting.