Cognitive Slide Extraction from Audiovisual Streams

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods fail to effectively extract and enrich slide presentations from audiovisual content, instead summarizing or extracting semantic information, lacking mechanisms to process and isolate slide presentations within multimodal content.

Innovation Solution

A system and method utilizing cognitive computing, multimodal content processing, and knowledge engineering to automatically extract slides, process synchronized audio, and allow object substitution and user curation, enhancing semantic understanding and presentation enrichment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If current methods are used to process audiovisual content, then semantic information can be extracted, but slide presentations cannot be effectively isolated and extracted

Engineering Contradiction:
Improveslide presentation extractionVSAvoidprocessing mechanism
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments audiovisual content into distinct components by detecting slide boundaries and separating slide content from non-slide content. This segmentation enables the extraction of individual slide presentations from the continuous audiovisual stream, solving the problem of isolating discrete presentation elements from multimodal content.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that includes cognitive computing modules, search engines, and knowledge bases. This intermediary layer bridges the gap between raw audiovisual content and extracted slide presentations, enabling semantic understanding and enrichment while maintaining the ability to isolate and extract presentation content.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If slide presentations are extracted from multimodal content, then manipulation ease is improved, but the extraction mechanism becomes more complex

Engineering Contradiction:
Improvepresentation manipulationVSAvoidextraction system
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically detecting slide boundaries, extracting content, and structuring presentations without requiring manual intervention. The cognitive computing modules autonomously analyze audiovisual content, identify presentation elements, and generate editable slide decks, enabling easy manipulation of extracted content while minimizing the apparent complexity for the end user.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The extraction system performs preliminary actions by pre-processing audiovisual content to detect and mark slide boundaries before actual extraction occurs. This preliminary segmentation simplifies the subsequent extraction and manipulation processes, as the system has already organized the content into discrete, manageable units that can be easily edited and manipulated.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If cognitive computing and knowledge base enrichment are applied, then semantic understanding is enhanced, but processing time increases

Engineering Contradiction:
Improvesemantic understandingVSAvoidprocessing duration
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system applies partial enrichment by selectively applying cognitive computing and knowledge base queries only to relevant content elements that require semantic enhancement. Rather than processing entire presentations uniformly, the system identifies and enriches only the portions that benefit from semantic understanding, reducing overall processing time while maintaining comprehensive semantic coverage where needed.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent replaces traditional mechanical text processing with cognitive computing systems that use natural language understanding, semantic analysis, and knowledge base integration. This substitution enables deeper semantic understanding without proportionally increasing processing time, as cognitive systems can parallelize processing and leverage pre-computed knowledge bases to accelerate semantic enrichment.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11321667B2System and method to extract and enrich slide presentations from multimodal content through cognitive computing
Publication Date: 2022.05.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11321667B2 patent drawing
  • US11321667B2 patent drawing
  • US11321667B2 patent drawing

AI summary

A system, product, and method including automatically performing extraction of slides from multimodal content, performing object extraction from each of the slides, allowing object substitution through semantics and concepts of the objects extracted, processing audio synchronized with the slides enriched with cognitive computing, search engine, and knowledge base, to provide annotations of the slides, processing the audio synchronized with the object being presented in each slide to enhance semantics and understanding, and curating for each step with human-machine interaction to provide a learning process by the system.