Logistic Regression Model for User Interest Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods are inefficient in accurately identifying user interests from vast amounts of internet data, making it difficult for enterprises to determine potential customers and understand user needs effectively.
Innovation Solution
A method and device that utilize logistic regression models trained with manually labeled training samples and optimized parameters, combined with ROC curve evaluation, to classify text data and compute confidence scores, enabling accurate identification of user interests.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional data processing methods are used to analyze user information, then the process is simple, but the accuracy of identifying user interests and determining potential customers is insufficient
Solution Approach 1:
The patent segments the user interest identification process into multiple independent modules: data collection module, feature extraction module, classification model training module, and interest identification module. Each module handles a specific task, improving overall accuracy while maintaining manageable system complexity through modular design.
Solution Approach 2:
The patent employs a composite approach by integrating multiple algorithms and data sources: combining social media data with transaction records, integrating feature extraction with classification models, and merging multiple features (user attributes, behavior patterns, social relationships) to create a comprehensive user profile that significantly improves identification accuracy.
2Loss of information
If comprehensive user data is collected from social media and transaction records, then user understanding is more complete, but the data processing time and computational resources increase
Solution Approach 1:
The patent extracts only the most relevant features from the comprehensive user data using feature extraction algorithms. Instead of processing all raw data, it identifies and extracts key features such as user attributes, behavior patterns, and social relationship indicators, maintaining information completeness while significantly reducing processing time and computational resources.
Solution Approach 2:
The patent performs preliminary data processing and feature extraction before the actual classification and interest identification. By pre-processing the data, organizing it into structured formats, and extracting relevant features in advance, the system reduces the computational burden during the main analysis phase, thereby reducing overall processing time while maintaining data completeness.
3Manufacturing precision
If manual labeling is used to create training samples, then the quality of training data is high, but the time and labor required for sample preparation increases
Solution Approach 1:
The patent performs preliminary automated processing of training samples before manual labeling. It pre-processes raw data, performs initial feature extraction, and organizes data into structured formats that facilitate more efficient manual labeling. This preliminary action reduces the time and effort required for manual labeling while maintaining high sample quality.
Solution Approach 2:
The patent introduces an intermediary automated feature extraction system that bridges raw data and manual labeling processes. This intermediary system prepares and structures the data in a way that makes manual labeling more efficient, reducing the direct time burden on annotators while ensuring high-quality training samples through systematic feature preparation.
Data Source
AI summary
Disclosed is a method for identifying a user interest, including: obtaining training samples and test samples, the training samples being obtained by manually labeling after the corresponding topic models have been trained based on text data; extracting characteristics of the training samples and of the test samples, and computing optimal model parameters of a logistic regression model by an iterative algorithm based on the characteristics of the training samples; evaluating the logistic regression model based on the characteristics of the test samples and an area AUC under an ROC curve to train and obtain a first theme classifier; determining a theme to which the text data belongs using the first theme classifier, computing a score of the theme to which the text data belongs based on the logistic regression model, and computing a confidence score of the user being interested in the theme according to a second preset algorithm. Further disclosed are a device for identifying a user interest and a computer-readable storage medium.


