Reading Log Analysis for User Interest Trend Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for analyzing user browsing behavior fail to effectively identify trends in document product classes, limiting network service providers' ability to offer personalized services based on user interests.
Innovation Solution
A system and method for analyzing reading logs that involves extracting relevant information, processing document content to identify keyword sets, clustering topics, calculating cohesion, classifying topics into classes, and analyzing reading trends over time to determine user interest in different topic classes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If current event analysis methods are used to collect and analyze user browsing behavior, then browsing behavior data can be gathered, but the ability to understand trends of interest in document product classes for all users is insufficient
Solution Approach 1:
The patent segments the analysis process into distinct modules: reading log extraction, interesting document filtering, document pre-processing, topic clustering, topic classification, degree of interest calculation, and reading trend analysis. This segmentation allows each module to focus on specific aspects of user behavior analysis, improving the overall precision of measuring user interest trends in document product classes.
Solution Approach 2:
The patent introduces intermediate processing steps including topic clustering and topic classification as mediators between raw browsing data and final trend analysis. These intermediary processes transform unstructured reading logs into structured topic-based insights, enabling more precise measurement of user interests in different document product classes.
2Loss of information
If detailed reading log analysis is performed on all documents, then comprehensive user behavior data is obtained, but the complexity of processing and analyzing the data increases significantly
Solution Approach 1:
The patent extracts only the most relevant information from reading logs through the interesting document filter, which selects documents based on specific criteria before further processing. This extraction principle reduces the volume of data that needs to be processed while maintaining the completeness of meaningful browsing behavior data, thereby reducing system complexity.
Solution Approach 2:
The patent performs preliminary actions including reading log extraction, interesting document filtering, and document pre-processing before the main analysis. These preliminary steps organize and prepare the data in advance, reducing the complexity of subsequent topic clustering and trend analysis operations.
3Measurement precision
If topic clustering and classification are performed on all reading data, then accurate user interest trends are identified, but the processing time and computational resources increase
Solution Approach 1:
The patent segments the large-scale data processing into time segments and processes reading logs in intervals. This segmentation allows topic clustering and classification to be performed on smaller subsets of data at a time, reducing computational burden and processing time while maintaining accurate identification of user interest trends through cumulative analysis.
Solution Approach 2:
The patent applies partial action by focusing topic clustering and classification only on interesting documents that meet specific criteria, rather than processing all reading data. This selective approach maintains precise user interest trend identification while significantly reducing the time and computational resources required for processing.
Data Source
AI summary
Methods for analyzing reading log and documents corresponding thereof are provided, including: acquiring reading log and documents corresponding thereto, wherein the reading log at least includes reading-related information about the documents within a predetermined period of time, selecting interesting document sets from the documents according to the reading log in each time segment, performing a document content pre-processing on the interesting document sets to determine keyword sets corresponding thereto for each time segment according to the interesting document sets, performing cluster calculation on the keyword sets to obtain topics and calculating cohesion of each topic, deleting topics with insufficient cohesion to obtain multiple high-relevance topics and classifying each high-relevance topic into one of predetermined topic classes according to the respective keyword sets of the high-relevance topics, obtaining reading statistics for each topic class and calculating multiple degrees of interest for each topic class during each time segment.


