Metadata Extraction for OCR and Speech Recognition Summaries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional OCR and speech recognition technologies fail to provide metadata such as text size, location, volume, and background noise levels, limiting the usefulness of extracted text, which is often stored in raw form for manual review without external exposure or utilization.

Innovation Solution

A method that generates electronic summaries by producing metadata from recognized source content, such as text size, location, and audio volume levels, to select and summarize content, enabling effective indexing, searching, and automatic labeling of online meeting data, without requiring extensive natural language processing or semantic understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional OCR extracts text from images, then text recognition is achieved, but no size or location metadata is provided

Engineering Contradiction:
Improvemetadata informationVSAvoidrecognition system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the text extraction process into two independent components: (1) text recognition that extracts text from images, and (2) metadata extraction that captures size and location information. This segmentation allows each component to specialize, with the metadata extraction module specifically targeting structural information without complicating the core recognition system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a universal metadata extraction module that simultaneously captures multiple types of information (text size, location, formatting) from recognized text elements. This multi-functional approach consolidates what could be separate complex processes into a single unified system that handles all metadata extraction needs.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of information

If conventional speech recognition extracts text from audio, then speech-to-text conversion is achieved, but no volume or background noise metadata is provided

Engineering Contradiction:
Improveaudio metadataVSAvoidspeech recognition system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The speech recognition system is divided into independent modules: audio-to-text conversion and audio metadata extraction. The metadata module separately captures volume levels and background noise characteristics without interfering with the core speech recognition process, preventing complexity propagation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary metadata extraction layer that sits between the audio input and text output. This intermediary component captures audio characteristics (volume, noise levels) without requiring the core speech recognition system to be modified, acting as a buffer that adds functionality without complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of time

If extracted text is stored in raw form for manual review, then complete text is preserved, but manual searching and review is time-consuming

Engineering Contradiction:
Improvemanual review timeVSAvoidcontent accessibility
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The patent applies preliminary action by automatically generating summaries and extracting key metadata from text content before manual review occurs. This pre-processing creates an organized structure with identified important elements, so when manual review is needed, users start with pre-organized content rather than raw unprocessed text.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs self-service by automatically organizing, tagging, and summarizing extracted text content without requiring manual intervention. The metadata extraction and text organization happen autonomously, creating ready-to-search structured data that improves accessibility without adding manual operational steps.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9569428B2Providing an electronic summary of source content
Publication Date: 2017.02.14 GOTO GRP INC
  • US9569428B2 patent drawing
  • US9569428B2 patent drawing
  • US9569428B2 patent drawing

AI summary

A technique provides an electronic summary of source content. The technique involves performing, on the source content, a content recognition operation to electronically generate text output from the source content. The technique further involves electronically evaluating text portions of the text output based on predefined usability criteria to produce a respective set of usability properties for each text portion of the text output. The technique further involves providing, as the electronic summary of the source content, summarization output which summarizes the source content. The summarization output includes a particular text portion of the text output which is selected from the text portions of the text output based on the respective set of usability properties for each text portion of the text output.