Document Data Classification Using Noise-to-Content Ratio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for converting web pages into formats suitable for e-readers and other electronic devices are inefficient in removing noise, such as advertisements and navigation panels, which are not interesting to users, leading to increased operational costs and reduced availability of electronic documents for mobile devices.

Innovation Solution

A document data classification subsystem that calculates a noise-to-content ratio for electronic documents, classifying noise and substantive content using metrics like word count, image presence, and visibility during rendering, and converts only substantive content into a format understandable by e-readers and other mobile devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If existing methods are used to convert web pages into formats for e-readers, then conversion can be performed, but noise such as advertisements and navigation panels cannot be effectively removed, leading to increased operational costs and reduced document availability

Engineering Contradiction:
Improvedocument conversion efficiencyVSAvoidnoise content in converted documents
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The patent segments the web page content into different components by calculating a noise-to-content ratio for each portion. Different metrics are applied to different segments (text portions vs. other portions) to classify them as noise or substantive content, enabling targeted removal of noise while preserving valuable content

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and removes noise portions from the web page by identifying them through noise-to-content ratio calculation. The system separates noise elements (advertisements, navigation panels) from substantive content and converts only the cleaned content to e-reader formats

Inventive Principle:
Principle #2Taking out (Extraction)

2Object-generated harmful factors

If manual noise removal methods are used, then some noise can be removed, but operational costs increase and the number of available electronic documents decreases

Engineering Contradiction:
Improvenoise removal effectivenessVSAvoidmanual processing time
Core Design Contradiction:
Object-generated harmful factorsVSLoss of time

Solution Approach 1:

The patent implements an automated system that performs noise removal without manual intervention. The noise-to-content ratio calculation and classification processes occur automatically, allowing the system to service itself by identifying and removing noise efficiently at scale

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the parameter of noise identification from manual visual inspection to automated metric-based classification. By using calculable parameters like word count, link density, and image ratios, the system transforms noise removal into an automated parameter-driven process

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10275523B1Document data classification using a noise-to-content ratio
Publication Date: 2019.04.30 AMAZON TECH INC
  • US10275523B1 patent drawing
  • US10275523B1 patent drawing
  • US10275523B1 patent drawing

AI summary

A method and system for classifying document data is described. The method may include classifying a first portion of an electronic document as substantive content or noise, classifying a second portion of the electronic document as substantive content or noise, determining a first feature of the first portion of the electronic document indicative of substantive content using a machine learning algorithm, and determining a second feature of the second portion of the electronic document indicative of noise using the machine learning algorithm.