Web Page Text Extraction via Dynamic Tag Injection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing web page browsing technologies, such as Safari and Evernote plug-ins, face limitations in extracting text content from web pages, particularly on mobile devices, where some web pages are not suitable for display and the extraction process is complex and not user-friendly.
Innovation Solution
A method and apparatus for a mobile terminal that determines the presence of text content tags in web page source code, extracting text content within these tags if present, or identifying and adding tags to facilitate extraction when tags are absent, to improve text content recognition and display.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing technologies (Safari browser, Evernote plug-in) are used for web page browsing, then text content extraction can be achieved, but the number of supported web pages is limited and browsing experience is affected
Solution Approach 1:
The patent segments the text extraction process into two distinct paths: one for web pages with structured text content tags and another for web pages without such tags. This segmentation allows the system to handle different web page types appropriately, improving both compatibility and reliability of text extraction across diverse web pages.
Solution Approach 2:
The patent performs preliminary detection of text content tags in web page source code before attempting text extraction. This preliminary action allows the system to identify suitable web pages for extraction and prepare appropriate processing methods in advance, thereby improving overall system reliability and adaptability.
2Ease of operation
If Evernote plug-in is used for web page browsing, then text content can be extracted, but the operation is complicated and cannot smartly recognize or inform the user
Solution Approach 1:
The patent implements a self-service mechanism where the system automatically detects the presence of text content tags in web page source code and determines whether text extraction is appropriate without requiring user intervention. The system smartly recognizes suitable web pages and can inform users through UI indicators, thereby improving ease of operation while maintaining high automation.
Solution Approach 2:
The patent introduces feedback mechanisms where the system detects text content tags and provides visual feedback to users through UI elements (such as displaying extraction indicators). This feedback loop enables smart recognition capability while keeping the operation simple for users, as they can see which pages are suitable for extraction without needing to understand the technical process.
3Manufacturing precision
If text content tags are not present in web page source code, then extraction cannot be performed, but adding tags manually increases complexity
Solution Approach 1:
The patent performs preliminary detection of text content tags in web page source code before attempting text extraction. This preliminary action allows the system to identify web pages that already have proper tagging, ensuring high extraction accuracy without requiring manual tag addition.
Solution Approach 2:
The patent introduces an intermediary processing layer that automatically analyzes web page source code to detect text content tags. This intermediary mechanism bridges the gap between web pages with and without tags, maintaining high extraction accuracy by identifying structured content automatically rather than requiring manual tag addition, thus avoiding increased complexity.
Data Source
AI summary
Methods and apparatus for extracting web page content are provided herein. An exemplary method can be implemented by a mobile terminal. A request command to open a first web page can be received. Whether a source code contains text content tags can be determined. When the source code corresponding to the first web page contains the text content tags, text content of the first web page enclosed within the text content tags can be extracted by a reader. When the source code does not contain the text content tags, a start position and an end position to indicate the text content of the first web page can be identified in the source code. The text content tags can be respectively added after the start position and before the end position. The text content of the first web page enclosed within the text content tags can then be extracted.


