Document Data Classification Using Noise-to-Content Ratio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for converting web pages into formats suitable for e-readers and other electronic devices are inefficient in removing noise, such as advertisements and navigation panels, which are not interesting to users, leading to increased operational costs and reduced availability of electronic documents for mobile devices.
Innovation Solution
A document data classification subsystem that calculates a noise-to-content ratio for electronic documents, classifying noise and substantive content using metrics like word count, image presence, and visibility during rendering, and converts only substantive content into a format understandable by e-readers and other mobile devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing methods are used to convert web pages into formats for e-readers, then conversion can be performed, but noise such as advertisements and navigation panels cannot be effectively removed, leading to increased operational costs and reduced document availability
Solution Approach 1:
The patent segments the web page content into different components by calculating a noise-to-content ratio for each portion. Different metrics are applied to different segments (text portions vs. other portions) to classify them as noise or substantive content, enabling targeted removal of noise while preserving valuable content
Solution Approach 2:
The patent extracts and removes noise portions from the web page by identifying them through noise-to-content ratio calculation. The system separates noise elements (advertisements, navigation panels) from substantive content and converts only the cleaned content to e-reader formats
2Object-generated harmful factors
If manual noise removal methods are used, then some noise can be removed, but operational costs increase and the number of available electronic documents decreases
Solution Approach 1:
The patent implements an automated system that performs noise removal without manual intervention. The noise-to-content ratio calculation and classification processes occur automatically, allowing the system to service itself by identifying and removing noise efficiently at scale
Solution Approach 2:
The patent changes the parameter of noise identification from manual visual inspection to automated metric-based classification. By using calculable parameters like word count, link density, and image ratios, the system transforms noise removal into an automated parameter-driven process
Data Source
AI summary
A method and system for classifying document data is described. The method may include classifying a first portion of an electronic document as substantive content or noise, classifying a second portion of the electronic document as substantive content or noise, determining a first feature of the first portion of the electronic document indicative of substantive content using a machine learning algorithm, and determining a second feature of the second portion of the electronic document indicative of noise using the machine learning algorithm.


