Metadata Extraction for OCR and Speech Recognition Summaries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional OCR and speech recognition technologies fail to provide metadata such as text size, location, volume, and background noise levels, limiting the usefulness of extracted text, which is often stored in raw form for manual review without external exposure or utilization.
Innovation Solution
A method that generates electronic summaries by producing metadata from recognized source content, such as text size, location, and audio volume levels, to select and summarize content, enabling effective indexing, searching, and automatic labeling of online meeting data, without requiring extensive natural language processing or semantic understanding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional OCR extracts text from images, then text recognition is achieved, but no size or location metadata is provided
Solution Approach 1:
The patent segments the text extraction process into two independent components: (1) text recognition that extracts text from images, and (2) metadata extraction that captures size and location information. This segmentation allows each component to specialize, with the metadata extraction module specifically targeting structural information without complicating the core recognition system.
Solution Approach 2:
The patent implements a universal metadata extraction module that simultaneously captures multiple types of information (text size, location, formatting) from recognized text elements. This multi-functional approach consolidates what could be separate complex processes into a single unified system that handles all metadata extraction needs.
2Loss of information
If conventional speech recognition extracts text from audio, then speech-to-text conversion is achieved, but no volume or background noise metadata is provided
Solution Approach 1:
The speech recognition system is divided into independent modules: audio-to-text conversion and audio metadata extraction. The metadata module separately captures volume levels and background noise characteristics without interfering with the core speech recognition process, preventing complexity propagation.
Solution Approach 2:
The patent introduces an intermediary metadata extraction layer that sits between the audio input and text output. This intermediary component captures audio characteristics (volume, noise levels) without requiring the core speech recognition system to be modified, acting as a buffer that adds functionality without complexity.
3Loss of time
If extracted text is stored in raw form for manual review, then complete text is preserved, but manual searching and review is time-consuming
Solution Approach 1:
The patent applies preliminary action by automatically generating summaries and extracting key metadata from text content before manual review occurs. This pre-processing creates an organized structure with identified important elements, so when manual review is needed, users start with pre-organized content rather than raw unprocessed text.
Solution Approach 2:
The system performs self-service by automatically organizing, tagging, and summarizing extracted text content without requiring manual intervention. The metadata extraction and text organization happen autonomously, creating ready-to-search structured data that improves accessibility without adding manual operational steps.
Data Source
AI summary
A technique provides an electronic summary of source content. The technique involves performing, on the source content, a content recognition operation to electronically generate text output from the source content. The technique further involves electronically evaluating text portions of the text output based on predefined usability criteria to produce a respective set of usability properties for each text portion of the text output. The technique further involves providing, as the electronic summary of the source content, summarization output which summarizes the source content. The summarization output includes a particular text portion of the text output which is selected from the text portions of the text output based on the respective set of usability properties for each text portion of the text output.


