Webpage Classification via Multi-Element Prediction and Log Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current webpage classification technologies rely on semi-automatic methods, which are inefficient and lack scalability and timeliness, especially when dealing with large volumes of data, as they require manual review and are not suitable for rapid classification of newly generated webpages.
Innovation Solution
A fully automatic webpage classification method that parses multiple webpage elements, predicts candidate classifications based on these elements, and determines a final classification by comparing the predicted classifications, utilizing historical search logs to enhance accuracy and scalability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If semi-automatic classification method with manual review is used, then classification accuracy can be improved, but processing efficiency and timeliness deteriorate
Solution Approach 1:
The patent implements self-service by enabling the classification system to automatically learn and adapt from historical search logs without requiring manual review. The system performs multi-dimensional classification prediction and automatic result determination, making the classification process self-sufficient and eliminating the need for human intervention while maintaining high accuracy.
Solution Approach 2:
The patent applies parameter changes by transitioning from traditional single-algorithm classification to a multi-dimensional prediction approach that evaluates multiple classification candidates simultaneously. This involves changing the classification parameters to include multiple prediction dimensions (such as different webpage elements and features) and using historical search log data to dynamically adjust classification parameters, thereby achieving both high accuracy and efficiency.
2Device complexity
If traditional classification algorithms are used, then implementation simplicity is maintained, but classification accuracy deteriorates
Solution Approach 1:
The patent applies segmentation by dividing the classification process into multiple independent prediction stages, each analyzing different webpage elements (such as title, content, metadata). Each segment produces a candidate classification result, and the final classification is determined by comparing and synthesizing these segmented predictions. This segmentation approach maintains implementation simplicity while improving accuracy through multi-dimensional analysis.
3Stability of the object's composition
If manual-defined classification categories are used, then classification stability is maintained, but scalability and adaptability deteriorate
Solution Approach 1:
The patent implements dynamics by making the classification system adaptive and dynamic rather than static. The system learns from historical search logs and automatically adjusts classification categories and parameters based on actual user behavior and data patterns. This dynamic approach allows the classification system to scale and adapt to new webpage types and classification needs while maintaining stability through consistent learning mechanisms.
4Measurement precision
If two-stage classification process (algorithm + manual review) is used, then classification accuracy can be ensured, but processing time and system complexity increase
Solution Approach 1:
The patent applies the taking out principle by extracting and eliminating the manual review stage from the classification process. Instead of requiring human reviewers, the system uses multi-dimensional prediction algorithms that automatically analyze multiple webpage elements and historical search logs to determine final classifications. This extraction of the manual review component significantly reduces processing time while maintaining accuracy through enhanced automated prediction capabilities.
Data Source
AI summary
A method and device for classifying webpages are provided. The method comprises: parsing a plurality of webpage elements from a webpage to be predicted; predicting a candidate webpage classification to which the webpage to be predicted belongs respectively according to respective webpage elements; and determining a final webpage classification of the webpage to be predicted by comparing the candidate webpage classifications predicted respectively based on the respective webpage elements.


