Web Page Quality Model Using Search Log Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing web page quality models have relatively poor accuracy due to limited sample sizes and manual rule summarization, affecting web page sorting results and user experience.
Innovation Solution
A method and apparatus for establishing a web page quality model by excavating user behavior indicators from search engine logs, calculating web page quality, and extracting quality features to automatically construct a model, improving accuracy and user experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual rules are summarized from limited samples to establish a web page quality model, then the model can be constructed with available data, but the accuracy of the established model is relatively poor
Solution Approach 1:
The system uses search engine log data that is automatically generated by user behaviors (clicks, dwell time, navigation) to train the quality model. The data serves itself by utilizing existing user interaction records without requiring manual sample collection, thereby accessing a much larger and more diverse dataset that improves model accuracy
Solution Approach 2:
The patent replaces manual rule summarization with an automated machine learning approach. Instead of manually observing and summarizing rules from limited samples, the system uses algorithms to automatically learn patterns from large-scale log data, substituting human mechanical analysis with computational processing that can handle vast quantities of data
2Measurement precision
If manual rules are used to establish the web page quality model, then the model construction process is simple, but the accuracy of calculated web page quality is relatively poor, affecting web page sorting results
Solution Approach 1:
The patent replaces simple manual rule-based systems with a complex automated machine learning system. The model establishment process uses automated algorithms to learn from large-scale log data, extracting quality indicators and features through computational processes rather than manual rule creation, thereby achieving higher accuracy despite increased system complexity
Solution Approach 2:
The patent segments the web page quality assessment into multiple independent components: quality indicators (derived from user behaviors like click-through rate and dwell time), quality features (extracted from page content and structure), and the final quality model. This segmentation allows each component to be optimized independently while contributing to the overall accuracy of the sorting system
Data Source
AI summary
A web page quality model establishment method and apparatus are disclosed. The method includes: excavating, from a search engine log, a selected user behavior indicator of each web page included in the search engine log, and calculating, according to the excavated selected user behavior indicator of each web page, web page quality of a corresponding web page; extracting, from the search engine log, a selected quality feature of each web page included in the search engine log; and establishing a web page quality model according to the web page quality and the selected quality feature of each web page included in the search engine log. Accuracy of a web page quality model established by means of this solution is relatively high, and accuracy of calculated web page quality is relatively high, thereby ensuring accuracy of a web page sorting result and user experience.
