Automatic Data Extraction Rules for Electronic Messages

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for extracting data from electronic messages require manual generation of data extraction rules, which is costly, time-consuming, and not feasible due to the high volume and variability of machine-generated emails.

Innovation Solution

The system automatically generates data extraction rules by analyzing a corpus of electronic messages, grouping them into clusters based on structure and domain, and using XML Path Language (XPATH) expressions to identify and extract variable data items.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual generation of data extraction rules is used, then extraction accuracy can be maintained, but cost and time consumption increase significantly

Engineering Contradiction:
Improveextraction accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses machine learning algorithms to automatically generate data extraction rules from electronic messages without requiring manual intervention. The model analyzes message structures and autonomously creates extraction rules, enabling the system to serve itself rather than requiring human editors to manually create rules for each message type and sender.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual rule generation with an automated machine learning-based system. The machine learning model processes electronic messages and generates extraction rules automatically, substituting the human editorial process with an automated computational mechanism that can handle large volumes of messages efficiently.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual generation of data extraction rules is used, then rule accuracy can be ensured, but processing cost becomes prohibitive

Engineering Contradiction:
Improverule accuracyVSAvoidprocessing cost
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The machine learning system automatically generates extraction rules without requiring human editors, enabling the system to process messages independently. This self-service capability eliminates the need for paid human resources and makes the processing cost scalable rather than linear with volume.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent substitutes the costly manual editorial process with an automated machine learning system that can process messages at scale. The machine learning model infers extraction rules from message patterns, providing accurate rules without the prohibitive costs associated with manual generation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If manual generation of data extraction rules is used, then extraction reliability can be maintained, but latency increases

Engineering Contradiction:
Improveextraction reliabilityVSAvoidlatency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The machine learning model is pre-trained on a corpus of electronic messages to automatically generate extraction rules in advance. This preliminary action allows the system to have ready-made extraction rules available immediately when new messages arrive, eliminating the latency associated with manual rule generation and ensuring timely processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the slow manual rule generation process with rapid automated machine learning inference. The machine learning system can quickly analyze message structures and generate extraction rules in real-time, dramatically reducing latency compared to manual processes while maintaining reliable extraction through learned patterns.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Productivity

If automated script standardization is implemented, then data extraction efficiency improves, but adaptability to different senders and message types decreases

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoidadaptability to different senders
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The machine learning model dynamically adapts its extraction rules based on the specific characteristics of each sender and message type. Rather than using a fixed standardized script, the system learns and adjusts its extraction patterns according to the unique structure of each message corpus, maintaining high efficiency while preserving adaptability to different senders and types.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12222973B2Automatic electronic message content extraction method and apparatus
Publication Date: 2025.02.11 YAHOO ASSETS LLC
  • US12222973B2 patent drawing
  • US12222973B2 patent drawing
  • US12222973B2 patent drawing

AI summary

Disclosed are systems and methods for improving interactions with and between computers in electronic messaging, and other, systems supported by or configured with personal computing devices, servers and/or platforms. The systems interact to identify and retrieve data within or across platforms, which can be used to improve the quality of data used in processing interactions between or among processors in such systems. The disclosed systems and methods provide systems and methods for automatically generating data extraction rules, which can then be used to automatically extract data from electronic messages.