Cross-Network User Identification via Behavioral Feature Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for identifying users across multiple social networks are inadequate as they fail to comprehensively and accurately describe user features, leading to inaccurate association of user data sets.
Innovation Solution
A method and apparatus that extract and utilize multiple features correlated with behavioral data, such as social, spatial, temporal, and text features, to generate a classification and prediction model for identifying users across different social networks, enhancing the accuracy of user identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If only registration information and partial text content are used to model account features, then the method is simple to implement, but the user features cannot be described comprehensively and accurately
Solution Approach 1:
The patent segments user feature extraction into multiple independent modules: registration information extraction, text content extraction, behavioral data extraction (publishing frequency, interaction patterns, temporal patterns, spatial patterns). Each module extracts specific features that are then integrated to form a comprehensive user profile, resolving the contradiction between implementation simplicity and description accuracy.
Solution Approach 2:
The patent adds a new dimension of behavioral data analysis beyond traditional registration information and text content. By incorporating publishing behavior, interaction patterns, temporal patterns, and spatial patterns, the system transitions from 2D feature space (registration + text) to 4D feature space (registration + text + behavior + patterns), achieving comprehensive and accurate user description.
2Measurement precision
If multiple features correlated with behavioral data are extracted and used in classification models, then user information is enriched and prediction accuracy improves, but the system complexity increases
Solution Approach 1:
The patent divides the complex feature extraction and classification system into modular components: feature extraction module (with sub-modules for different feature types), model training module, and prediction module. This segmentation allows each component to be developed and optimized independently, managing system complexity while maintaining high prediction accuracy.
Solution Approach 2:
The patent designs a universal classification and prediction model framework that can handle multiple feature types (registration information, text content, behavioral data, patterns) through a unified interface. The model structure remains consistent regardless of the specific features input, reducing system complexity while enabling comprehensive user analysis.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention discloses a method and an apparatus for identifying a same user in multiple social networks, where the method includes: inputting accounts that are in a test set and are obtained from registered accounts in at least two different social networks, and generating at least one test set account combination from the accounts in the test set; extracting at least two different features that are of each account in the at least one test set account combination and are correlated with behavioral data of a user of the account; inputting the at least two different features that are of each account in the at least one test set account combination and are correlated with behavioral data of a user of the account into a classification and prediction model, to obtain a predicted value or a set of predicted values, which may belong to a same user, of the at least one test set account combination; and performing computation on the predicted value or the set of predicted values of the test set account combination by using an association algorithm, and outputting a computed prediction result for the test set account combination. In the foregoing manner, the present invention can describe user information comprehensively and accurately, which makes a final prediction result more accurate.