Webpage Classification via Multi-Element Prediction and Log Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current webpage classification technologies rely on semi-automatic methods, which are inefficient and lack scalability and timeliness, especially when dealing with large volumes of data, as they require manual review and are not suitable for rapid classification of newly generated webpages.

Innovation Solution

A fully automatic webpage classification method that parses multiple webpage elements, predicts candidate classifications based on these elements, and determines a final classification by comparing the predicted classifications, utilizing historical search logs to enhance accuracy and scalability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If semi-automatic classification method with manual review is used, then classification accuracy can be improved, but processing efficiency and timeliness deteriorate

Engineering Contradiction:
Improveclassification accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements self-service by enabling the classification system to automatically learn and adapt from historical search logs without requiring manual review. The system performs multi-dimensional classification prediction and automatic result determination, making the classification process self-sufficient and eliminating the need for human intervention while maintaining high accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies parameter changes by transitioning from traditional single-algorithm classification to a multi-dimensional prediction approach that evaluates multiple classification candidates simultaneously. This involves changing the classification parameters to include multiple prediction dimensions (such as different webpage elements and features) and using historical search log data to dynamically adjust classification parameters, thereby achieving both high accuracy and efficiency.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If traditional classification algorithms are used, then implementation simplicity is maintained, but classification accuracy deteriorates

Engineering Contradiction:
Improvealgorithm simplicityVSAvoidclassification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies segmentation by dividing the classification process into multiple independent prediction stages, each analyzing different webpage elements (such as title, content, metadata). Each segment produces a candidate classification result, and the final classification is determined by comparing and synthesizing these segmented predictions. This segmentation approach maintains implementation simplicity while improving accuracy through multi-dimensional analysis.

Inventive Principle:
Principle #1Segmentation

3Stability of the object's composition

If manual-defined classification categories are used, then classification stability is maintained, but scalability and adaptability deteriorate

Engineering Contradiction:
Improveclassification stabilityVSAvoidscalability
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamics by making the classification system adaptive and dynamic rather than static. The system learns from historical search logs and automatically adjusts classification categories and parameters based on actual user behavior and data patterns. This dynamic approach allows the classification system to scale and adapt to new webpage types and classification needs while maintaining stability through consistent learning mechanisms.

Inventive Principle:
Principle #15Dynamics

4Measurement precision

If two-stage classification process (algorithm + manual review) is used, then classification accuracy can be ensured, but processing time and system complexity increase

Engineering Contradiction:
Improveclassification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies the taking out principle by extracting and eliminating the manual review stage from the classification process. Instead of requiring human reviewers, the system uses multi-dimensional prediction algorithms that automatically analyze multiple webpage elements and historical search logs to determine final classifications. This extraction of the manual review component significantly reduces processing time while maintaining accuracy through enhanced automated prediction capabilities.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10909427B2Method and device for classifying webpages
Publication Date: 2021.02.02 BEIJING QIHOOD TECHNOLOGY CO LTD
  • US10909427B2 patent drawing
  • US10909427B2 patent drawing
  • US10909427B2 patent drawing

AI summary

A method and device for classifying webpages are provided. The method comprises: parsing a plurality of webpage elements from a webpage to be predicted; predicting a candidate webpage classification to which the webpage to be predicted belongs respectively according to respective webpage elements; and determining a final webpage classification of the webpage to be predicted by comparing the candidate webpage classifications predicted respectively based on the respective webpage elements.