Logistic Regression Model for User Interest Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods are inefficient in accurately identifying user interests from vast amounts of internet data, making it difficult for enterprises to determine potential customers and understand user needs effectively.

Innovation Solution

A method and device that utilize logistic regression models trained with manually labeled training samples and optimized parameters, combined with ROC curve evaluation, to classify text data and compute confidence scores, enabling accurate identification of user interests.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional data processing methods are used to analyze user information, then the process is simple, but the accuracy of identifying user interests and determining potential customers is insufficient

Engineering Contradiction:
Improveaccuracy of user interest identificationVSAvoidcomplexity of data processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the user interest identification process into multiple independent modules: data collection module, feature extraction module, classification model training module, and interest identification module. Each module handles a specific task, improving overall accuracy while maintaining manageable system complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs a composite approach by integrating multiple algorithms and data sources: combining social media data with transaction records, integrating feature extraction with classification models, and merging multiple features (user attributes, behavior patterns, social relationships) to create a comprehensive user profile that significantly improves identification accuracy.

Inventive Principle:
Principle #40Composite materials

2Loss of information

If comprehensive user data is collected from social media and transaction records, then user understanding is more complete, but the data processing time and computational resources increase

Engineering Contradiction:
Improvecompleteness of user informationVSAvoiddata processing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent extracts only the most relevant features from the comprehensive user data using feature extraction algorithms. Instead of processing all raw data, it identifies and extracts key features such as user attributes, behavior patterns, and social relationship indicators, maintaining information completeness while significantly reducing processing time and computational resources.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary data processing and feature extraction before the actual classification and interest identification. By pre-processing the data, organizing it into structured formats, and extracting relevant features in advance, the system reduces the computational burden during the main analysis phase, thereby reducing overall processing time while maintaining data completeness.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If manual labeling is used to create training samples, then the quality of training data is high, but the time and labor required for sample preparation increases

Engineering Contradiction:
Improvequality of training samplesVSAvoidtraining sample preparation efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent performs preliminary automated processing of training samples before manual labeling. It pre-processes raw data, performs initial feature extraction, and organizes data into structured formats that facilitate more efficient manual labeling. This preliminary action reduces the time and effort required for manual labeling while maintaining high sample quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary automated feature extraction system that bridges raw data and manual labeling processes. This intermediary system prepares and structures the data in a way that makes manual labeling more efficient, reducing the direct time burden on annotators while ensuring high-quality training samples through systematic feature preparation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10977447B2Method and device for identifying a user interest, and computer-readable storage medium
Publication Date: 2021.04.13 PING AN TECH (SHENZHEN) CO LTD
  • US10977447B2 patent drawing
  • US10977447B2 patent drawing
  • US10977447B2 patent drawing

AI summary

Disclosed is a method for identifying a user interest, including: obtaining training samples and test samples, the training samples being obtained by manually labeling after the corresponding topic models have been trained based on text data; extracting characteristics of the training samples and of the test samples, and computing optimal model parameters of a logistic regression model by an iterative algorithm based on the characteristics of the training samples; evaluating the logistic regression model based on the characteristics of the test samples and an area AUC under an ROC curve to train and obtain a first theme classifier; determining a theme to which the text data belongs using the first theme classifier, computing a score of the theme to which the text data belongs based on the logistic regression model, and computing a confidence score of the user being interested in the theme according to a second preset algorithm. Further disclosed are a device for identifying a user interest and a computer-readable storage medium.