Web Detection System Markup Filtering and OCR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Optical character recognition (OCR) technologies face challenges in efficiently processing large volumes of webpages due to computational expenses and difficulties in detecting obfuscated or fabricated information, especially when malicious entities attempt to deceive automated systems by altering graphical content.

Innovation Solution

A web detection system processes webpage markup language to identify target characteristics, generates images representing the actual user view, and uses optical character recognition to detect graphical features, thereby identifying deviations and associating webpages with target entities, including detecting malicious activity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If optical character recognition is used to process all webpages, then text content can be extracted, but computational costs become excessively high

Engineering Contradiction:
Improvetext extraction accuracyVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent segments the webpage processing into multiple stages: first filtering webpages using markup language analysis to identify those containing target characteristics, then applying OCR only to the filtered subset. This segmentation reduces the number of webpages requiring computationally expensive OCR processing while maintaining text extraction accuracy for relevant content.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary filtering of webpages using markup language analysis before applying OCR. By pre-identifying and filtering webpages that contain target characteristics in their markup structure, the system prepares the data in advance to minimize the scope of subsequent OCR operations, thereby reducing overall computational cost.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If markup language is analyzed to filter webpages, then processing efficiency improves, but information loss occurs between markup and graphical representation

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidinformation discrepancy between markup and graphical content
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent employs feedback by comparing information extracted from markup language with information detected from graphical content through OCR. The system identifies deviations between these two information sources and uses this feedback to detect potential obfuscation or fabrication, thereby recovering information that might otherwise be lost in the filtering process.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent combines multiple information sources (markup language analysis and graphical content analysis) to create a composite verification system. By analyzing both the structured markup data and the visual rendering, the system cross-validates information and detects discrepancies, ensuring that filtering based on markup does not result in information loss.

Inventive Principle:
Principle #40Composite materials

3Measurement precision

If graphical content is processed to detect obfuscation, then detection accuracy improves, but processing time increases

Engineering Contradiction:
Improveobfuscation detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments graphical content processing to focus only on webpages that have already been filtered by markup language analysis. By applying OCR and graphical analysis only to the subset of webpages containing target characteristics, the system maintains high obfuscation detection accuracy while significantly reducing the total processing time compared to analyzing all webpages graphically.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11164052B2Image processing of webpages
Publication Date: 2021.11.02 MERCK SHARP & DOHME LLC
  • US11164052B2 patent drawing
  • US11164052B2 patent drawing
  • US11164052B2 patent drawing

AI summary

A web detection system processes webpage information and performs automated feature extraction of webpages including machine processable information. In an embodiment, the web detection system determines a subset of webpages having a target characteristic by processing markup language. For a webpage of the subset, the web detection system determines that a first image overlaps at least a portion of a second image in the webpage. The web detection system generates an image of the webpage such that the portion of the second image is obscured by the first image. The web detection system determines a graphical feature of the webpage by processing the image, e.g., using optical character recognition. Responsive to determining that the graphical feature corresponds to graphical features of images of a different set of webpages associated with a target entity, the web detection system determines that the webpage is also associated with the target entity.