Browser OCR Text Extraction for Searchable Media
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods fail to effectively extract and utilize text from images and videos within web documents for enhancing advertising relevance, limiting user interaction and ad targeting.
Innovation Solution
The method involves extracting text from images and videos using optical character recognition (OCR) techniques, updating the document's markup language to make the extracted text selectable and searchable, and associating tags with objects for relevant ad retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional methods are used to display images and videos in web documents, then the visual content is preserved, but text within the media cannot be extracted or searched
Solution Approach 1:
The patent extracts text information from images and videos by integrating OCR technology into the web browser. The OCR engine processes media objects directly in the browser, extracting text content without requiring external applications. This extracted text is then made searchable and selectable, resolving the contradiction by separating text extraction from the original media format while maintaining system integration.
Solution Approach 2:
The patent introduces an intermediary OCR processing layer between the media display and user interaction. This intermediary component enables text extraction and search functionality without directly modifying the original media files. The OCR engine acts as a mediator that converts visual text content into searchable text data, allowing users to interact with text within images and videos while preserving the original media integrity.
2Reliability
If text is extracted from images and videos using OCR, then ad relevance can be improved, but processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary text extraction from media objects when the page is initially loaded or when media elements become visible. By extracting text in advance rather than on-demand during user interaction, the system prepares searchable text content beforehand, reducing latency when users perform searches. This preliminary processing improves ad relevance while managing processing time through proactive text extraction.
Solution Approach 2:
The patent implements partial text extraction by focusing OCR processing on media objects that are likely to contain relevant text based on contextual analysis. Rather than processing every media object uniformly, the system selectively applies OCR to images and videos with higher probability of containing searchable text, thereby improving ad relevance while reducing overall processing time and computational resource consumption.
3Adaptability or versatility
If OCR processing is applied to all media objects, then text searchability is maximized, but system performance and user experience deteriorate
Solution Approach 1:
The patent applies local quality by enabling text extraction and search functionality selectively for specific media objects rather than uniformly across all media. The system determines which images and videos are most likely to contain searchable text based on contextual cues, file characteristics, and user behavior patterns. This localized approach maximizes text search capability for relevant media while maintaining fast page loading speeds by avoiding unnecessary OCR processing on media unlikely to contain text.
Solution Approach 2:
The patent implements partial text extraction by applying OCR processing only to a subset of media objects that meet specific criteria for text content probability. The system uses heuristics and machine learning models to identify media objects with high likelihood of containing searchable text, applying OCR selectively to these cases. This partial action approach maintains text search capability for relevant content while preserving overall system productivity and user experience by avoiding processing of all media objects.
Data Source
AI summary
A computing device can obtain data describing at least one document, the at least one document referencing at least one media object, wherein a portion of the at least one media object includes one or more characters. The computing device can obtain data describing the one or more characters in the at least one media object in the at least one document. The computing device can generate an updated copy of the at least one document that includes the data describing the one or more characters in the at least one media object. The computing device can present, on a display screen of the computing device and through an interface, the updated copy of the at least one document, wherein the one or more characters in the at least one media object are able to be selected or searched.


