HTML Email Content Extraction via DOM Link Positioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques for identifying key elements in machine-generated electronic mail require large pre-processing and training, leading to resource wastage and inefficiency in processing and network throughput.

Innovation Solution

A novel framework for partitioning HTML content in email messages based on the relative positions of links within the DOM hierarchy, allowing for real-time, scalable, and pre-processing-free identification of meaningful entities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional clustering mechanisms are used to identify key elements in email messages, then measurement precision is improved, but loss of time and productivity deteriorate due to extensive pre-processing requirements

Engineering Contradiction:
Improveidentification accuracyVSAvoidpre-processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts and analyzes only the hyperlink elements from email messages, rather than processing the entire message content. By focusing specifically on hyperlink attributes (URL, anchor text, position), the system achieves accurate content identification without the need for extensive pre-processing of full message bodies, thus resolving the contradiction between identification accuracy and pre-processing time

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates simplified representations of email content by copying only the essential hyperlink data into a structured format for analysis. This copying approach allows the system to work with a reduced data subset that retains the key identifying features needed for accurate content classification, eliminating the need to pre-process and store entire message contents

Inventive Principle:
Principle #26Copying

2Measurement precision

If conventional clustering mechanisms are used to identify key elements in email messages, then measurement precision is improved, but productivity deteriorates due to large system resource requirements

Engineering Contradiction:
Improveidentification accuracyVSAvoidprocessing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts and analyzes only the hyperlink elements from email messages, rather than processing the entire message content. By focusing specifically on hyperlink attributes (URL, anchor text, position), the system achieves accurate content identification without the need for extensive pre-processing of full message bodies, thus resolving the contradiction between identification accuracy and pre-processing time

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates simplified representations of email content by copying only the essential hyperlink data into a structured format for analysis. This copying approach allows the system to work with a reduced data subset that retains the key identifying features needed for accurate content classification, eliminating the need to pre-process and store entire message contents

Inventive Principle:
Principle #26Copying

3Measurement precision

If complex clustering mechanisms are used for message partitioning, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvecontent identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and analyzes only the hyperlink elements from email messages, rather than processing the entire message content. By focusing specifically on hyperlink attributes (URL, anchor text, position), the system achieves accurate content identification without the need for extensive pre-processing of full message bodies, thus resolving the contradiction between identification accuracy and pre-processing time

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates simplified representations of email content by copying only the essential hyperlink data into a structured format for analysis. This copying approach allows the system to work with a reduced data subset that retains the key identifying features needed for accurate content classification, eliminating the need to pre-process and store entire message contents

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20200120054A1Computerized system and method for digital content extraction and propagation in HTML messages
Publication Date: 2020.04.16 YAHOO ASSETS LLC
  • US20200120054A1 patent drawing
  • US20200120054A1 patent drawing
  • US20200120054A1 patent drawing

AI summary

Disclosed are systems and methods for improving interactions with and between computers in content providing, searching and/or hosting systems supported by or configured with devices, servers and/or platforms. The disclosed systems and methods provide a novel framework for partitioning HTML content in electronic messages based on the relative positions of the content's links within the DOM hierarchy of the messages, and basing the propagation (e.g., display or communication) of such content therefrom. The disclosed message partitioning and extraction framework can be applied online, in real-time, at scale, without any pre-processing or pre-learning/training.