Web Content Extraction for Readable Layout Preservation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current browser 'reader mode' functionalities often inaccurately identify main content, omit critical elements, and disrupt the original design and layout of web pages, leading to a suboptimal user experience due to the indiscriminate zooming and masking of non-essential elements.
Innovation Solution
A system that dynamically optimizes the display of web content by identifying and merging key elements such as title, content, and author areas, then fitting and repositioning them within a user-selected or predefined area using CSS scaling and translation functions, effectively hiding distracting elements while preserving the original layout and design.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-generated harmful factors
If reader mode creates a new webpage or overlays elements to mask original content, then non-essential elements are removed, but main content identification accuracy deteriorates and critical elements may be omitted
Solution Approach 1:
The patent extracts only the essential content elements (article content, title, author, publication date) from the original webpage while preserving their contextual relationships. This selective extraction approach ensures accurate main content identification while removing non-essential elements like advertisements and navigation controls.
Solution Approach 2:
The patent creates a simplified copy of the essential content elements that maintains their original semantic meaning and relationships. This copying approach preserves the integrity of main content identification while presenting a streamlined version free from distracting elements.
2Ease of operation
If reader mode generates a new webpage to display content, then readability is improved, but original design elements and layout are lost
Solution Approach 1:
The patent segments the webpage into essential content elements and non-essential elements. By separating these components, the system can preserve the design integrity of essential elements while removing non-essential ones, thus maintaining readability without sacrificing original design quality.
Solution Approach 2:
The patent applies different processing qualities to different parts of the webpage. Essential content elements retain their original design properties and layout characteristics, while non-essential elements are removed or simplified. This local quality approach ensures that design integrity is preserved where it matters most for readability.
3Ease of operation
If traditional zooming is used to enhance content visibility, then text becomes larger and easier to read, but layout is disrupted and non-essential elements are indiscriminately enlarged
Solution Approach 1:
The patent extracts only essential content elements for display, eliminating the need for indiscriminate zooming of all page elements. By presenting only the necessary content in a simplified layout, the system achieves enhanced visibility without disrupting the original layout structure or enlarging non-essential elements like navigation bars.
4Extent of automation
If reader mode removes all non-essential elements, then focus on main content is improved, but design continuity and aesthetic appeal are compromised
Solution Approach 1:
The patent applies selective quality preservation to essential content elements while removing non-essential elements. This local quality approach maintains design continuity and aesthetic appeal for the main content area, as these elements retain their original design properties, while still achieving effective content filtering.
Data Source
AI summary
This disclosure presents a system and method for optimizing the display of major information on web pages. It identifies key elements such as the title, main content, author, and publication date, dynamically merging these into an “original area.” This area is then fitted into a user-selected or predefined “target area” using advanced zooming and repositioning techniques. The method significantly enhances readability by removing non-essential elements like ads, navigation menus, and other distractions. Various determination methods are employed, including user configurations, known selector matching, external services, and machine learning models. The system is designed to be flexible and can be implemented as either a browser extension or a built-in browser feature. It offers both manual and automatic operation modes, allowing users to customize their viewing experience and maintain focus on the most relevant content, thereby improving overall user engagement and satisfaction.


