Web Content Analysis Using CSS Xpath Facet Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Content mining accumulates substantial amounts of unwanted data, making it difficult to distinguish between desired and undesired text content, which negatively impacts the accuracy of correlation values and facet value calculations in search results.
Innovation Solution
A method and system for analyzing webpage content that extracts facet values using CSS and Xpath expressions, indexing their locations, and classifying them to optimize web site modes, reducing the need for dynamic searching and improving computing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If crawlers accumulate data from accessible URLs, then data availability increases, but waste information increases and negatively impacts mining calculations
Solution Approach 1:
The patent segments the accumulated data into distinct categories by identifying different document types (e.g., product pages, category pages, informational pages) and applying type-specific extraction rules to each segment, thereby separating useful information from waste information systematically
Solution Approach 2:
The patent extracts only the relevant facet values from each document type based on pre-defined extraction rules specific to that document type, removing unwanted data before it can negatively impact the mining calculations
2Ease of operation
If all accumulated data is processed uniformly, then processing simplicity is maintained, but extraction accuracy decreases due to mixing desired and undesired data
Solution Approach 1:
The patent applies different extraction rules and criteria to different document types locally, rather than using a uniform approach globally. Each document type receives tailored processing based on its specific characteristics, improving extraction accuracy while maintaining operational simplicity through automation
3Measurement precision
If dynamic searching is performed to distinguish desired from undesired data, then extraction precision improves, but computing time and resources increase
Solution Approach 1:
The patent performs preliminary classification of documents into types and pre-defines extraction rules for each type before the actual data extraction process. This preliminary action eliminates the need for dynamic searching during extraction, maintaining high precision while significantly reducing computing time and resources
Data Source
AI summary
A method, computer system, and a computer program for analyzing a webpage content is provided. The present invention may include receiving a plurality of terms, including a plurality of structural information, derived from one or more responses to a faceted search query. The present invention may also include extracting a plurality of facet values from the plurality of terms. The present invention may then include generating an analysis of each facet value of the plurality of facet values. The present invention may further include determining a type of each facet value of the plurality of facet values based on the analysis. The present invention may also include generating an optimized web site mode interface including the analysis and displaying the optimized web site mode interface to a user.


