OCR-Based Media Search Text Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing search engine techniques are inadequate for performing searches based on media such as images, video, or audio, and cannot handle search queries derived from non-search engine content, particularly from brand or service-specific applications.

Innovation Solution

The system and method generate search result data by using machine-encoded text data from optical character recognition (OCR) techniques, allowing users to select and search text within images or videos, which is then processed to generate search queries that can be authenticated and searched across various media and text data sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional text-based search engine techniques are used, then search functionality is simple and reliable, but the system cannot handle media-based queries or extract text from images and videos

Engineering Contradiction:
Improvecapability to handle media-based searchesVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent combines traditional text-based search engine functionality with optical character recognition (OCR) capabilities into a unified system. The search engine now processes both direct text queries and media-based queries by integrating OCR modules that extract text from images and videos, allowing the same search infrastructure to handle multiple query types without requiring separate systems

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary OCR processing layer between the user query and the search index. When a user submits a media-based query, the system uses OCR technology as an intermediary to extract text from the media content, convert it to machine-encoded text, and then process it through the existing search pipeline. This intermediary layer enables media-based searching while preserving the原有 search engine architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If OCR techniques are integrated into the search system, then the system can recognize and process text within images and videos, but processing time and computational resources increase

Engineering Contradiction:
Improvetext recognition capabilityVSAvoidsearch processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by performing OCR text extraction and conversion to machine-encoded text in advance, before the actual search query is executed. The system pre-processes media content to extract and encode text, storing it in a format ready for immediate search processing. This eliminates the need to perform OCR computation at query time, significantly reducing search latency

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the search process into distinct phases: media ingestion and OCR text extraction, text encoding and indexing, and query processing. By separating the computationally intensive OCR operations from the query execution phase, the system can optimize each segment independently - performing OCR during content ingestion rather than during user searches, thereby reducing perceived processing time

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If the system processes and stores machine-encoded text from OCR, then search accuracy improves, but data storage requirements increase

Engineering Contradiction:
Improvesearch result accuracyVSAvoiddata storage volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent uses copying by creating machine-encoded text representations (copies) of the text content extracted from media images and videos. Instead of storing only the original media files, the system generates and stores textual copies that can be efficiently indexed and searched. These text copies are stored in a structured format that optimizes storage efficiency while enabling precise search matching

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies local quality by storing machine-encoded text data in optimized formats and locations within the search infrastructure. The extracted text is stored with appropriate metadata and indexing structures that enhance search accuracy without proportionally increasing overall storage requirements. The system selectively processes and stores text based on relevance and searchability criteria

Inventive Principle:
Principle #3Local quality

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables users to conduct media-based searches by recognizing and selecting text within images or videos, providing relevant information and search results, thereby enhancing the capability of search engines to handle diverse media types and user queries beyond traditional text-based searches.

Implementation Method 1

machine-encoded text data generated by optical character recognition techniques performed on media

Methodology Applied
Scientific EffectOptical character recognition:

Data Source

PatentUS11893815B2Systems and methods for generating search results based on optical character recognition techniques and machine-encoded text
Publication Date: 2024.02.06 YAHOO ASSETS LLC
  • US11893815B2 patent drawing
  • US11893815B2 patent drawing
  • US11893815B2 patent drawing

AI summary

Disclosed are systems and methods for generating search result data based on machine-encoded text generated by computer vision optical character recognition machine learning techniques performed on digital media. The disclosed systems and methods provide a novel framework for performing machine learning visual search or machine learning text extraction techniques on digital media in order to extract and analyze the data therein and further conduct search queries based on the extracted and analyzed data. The disclosed framework may leverage the aforementioned computer vision machine learning techniques in order to provide a user with relevant search results regarding objects and text detect in digital media captured on a user device.