Audio Video Summarization with Visual and Semantic Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language processing systems fail to generate coherent summaries by accounting for visual characteristics and formatting structure of input text documents, resulting in disjointed or incomplete summaries.

Innovation Solution

A computer-based system that analyzes both semantic context and visual formatting of input text documents to segment and generate summary snippets, incorporating visual cues and formatting features for improved summary generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If prior systems generate document summaries using only semantic analysis, then the summarization process is simple and fast, but the summaries become disjointed or incomplete when visual characteristics or formatting structure are present

Engineering Contradiction:
Improvesummary coherenceVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines semantic analysis with visual characteristic analysis in a unified summarization system. The system integrates multiple analysis components (semantic understanding, formatting detection, visual feature recognition) to generate coherent summaries that account for both text meaning and document structure, resolving the contradiction between summary reliability and system complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The summarization system uses a composite approach by integrating multiple types of analysis (semantic, visual, formatting) into a unified framework. This composite methodology allows the system to leverage strengths of different analysis types while mitigating their individual weaknesses, producing more reliable summaries without excessive complexity increase.

Inventive Principle:
Principle #40Composite materials

2Reliability

If prior systems ignore visual characteristics and formatting structure, then the processing speed is high, but the summary quality deteriorates in the presence of such characteristics

Engineering Contradiction:
Improvesummary completenessVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system segments the document analysis process into distinct components: semantic analysis, visual characteristic detection, and formatting structure recognition. This segmentation allows parallel processing of different document features, improving processing efficiency while maintaining comprehensive analysis quality that produces complete summaries.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis of visual characteristics and formatting structures before generating the final summary. By pre-identifying important structural elements and visual features, the system prepares the data in advance, reducing the time required for the actual summarization process while ensuring no important information is missed.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If the system analyzes both semantic context and visual formatting, then summary coherence is improved, but the computational resources required increase

Engineering Contradiction:
Improvesummary qualityVSAvoidcomputational energy
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system applies partial analysis by focusing computational resources on the most relevant visual and semantic features rather than analyzing every aspect of the document equally. This selective approach maintains high summary quality by capturing essential information while reducing unnecessary computational energy consumption on less critical document elements.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12353465B2Audio video summarizer
Publication Date: 2025.07.08 AL21 LABS
  • US12353465B2 patent drawing
  • US12353465B2 patent drawing
  • US12353465B2 patent drawing

AI summary

The presently disclosed embodiments may include a computer readable medium including instructions that when executed by one or more processing devices cause the one or more processing devices to: receive an identification of at least one source audio or video with audio file, generate a textual transcript based on an audio component associated with the at least one source audio or video with audio file, edit the textual transcript to provide a formatted textual transcript, segment the formatted textual transcript into two or more segments, generate at least one summary snippet associated with the two or more segments, and cause the at least one summary snippet to be shown on a display together with a representation of the at least one source audio or video with audio file.