Structured Transcription for Speech File Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for consuming speech audio files lack efficient topic-specific filtering and navigation, as they are linear and do not provide structured transcription or convenient interfaces for users to quickly access key portions, with abstractive speech summarization remaining unaddressed and user navigation through highlighted content being inconvenient.

Innovation Solution

A structured transcription system that converts speech files into navigable structured transcriptions using Automatic Speech Recognition (ASR) and Text-to-Text (STT) techniques, generating document trees with both extractive and abstractive summaries, allowing users to interactively navigate and highlight key sections through a user interface.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech files are consumed in linear format, then the content can be delivered continuously, but users cannot efficiently navigate to key portions or perform topic-specific filtering

Engineering Contradiction:
Improvenavigation efficiencyVSAvoidtranscription structure
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent segments the continuous speech file into discrete sections based on topic changes detected through transcription analysis. Each section represents a distinct topic or theme, allowing users to navigate directly to specific topics of interest rather than listening linearly through the entire file. This segmentation enables efficient topic-specific filtering and navigation while maintaining the original speech content integrity.

Inventive Principle:
Principle #1Segmentation

2Productivity

If playback speed is doubled to reduce consumption time, then users can consume content faster, but users still cannot selectively navigate to important portions

Engineering Contradiction:
Improveconsumption speedVSAvoidtime to locate key content
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of the speech file to generate a structured transcription with identified key sections and topics before the user consumes the content. This advance structuring allows users to quickly locate and jump to important portions without having to listen through content at accelerated speeds, thereby reducing both consumption time and the time needed to locate key information.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If automated summarization is applied to speech files, then key points can be identified, but abstractive speech summarization has not been meaningfully addressed

Engineering Contradiction:
Improvekey point retentionVSAvoidsummarization system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts key information from speech files by generating both extractive summaries (selecting important original segments) and abstractive summaries (generating new synthesized content that captures essential meaning). This dual approach identifies and highlights key points while managing system complexity through a structured pipeline that processes transcription data through multiple analysis stages.

Inventive Principle:
Principle #2Taking out (Extraction)

4Productivity

If users prefer textual content over speech content, then reading speed may be faster, but users need structured transcriptions with highlighted key points

Engineering Contradiction:
Improvecontent consumption rateVSAvoidtranscription usability
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent applies local quality enhancement by providing different representations of the same content tailored to user preferences. For textual consumers, the system provides structured transcriptions with highlighted key sections, bolded important phrases, and organized topic headers. This allows users to read at their own pace while easily identifying key information without having to process entire paragraphs uniformly.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10783314B2Emphasizing key points in a speech file and structuring an associated transcription
Publication Date: 2020.09.22 ADOBE INC
  • US10783314B2 patent drawing
  • US10783314B2 patent drawing
  • US10783314B2 patent drawing

AI summary

Techniques are disclosed for generating a structured transcription from a speech file. In an example embodiment, a structured transcription system receives a speech file comprising speech from one or more people and generates a navigable structured transcription object. The navigable structured transcription object may comprise one or more data structures representing multimedia content with which a user may navigate and interact via a user interface. Text and/or speech relating to the speech file can be selectively presented to the user (e.g., the text can be presented via a display, and the speech can be aurally presented via a speaker).