Multi-dimensional data analysis method and system based on WeChat nickname information extraction
By performing semantic coding and cluster analysis on WeChat nicknames and extracting multi-dimensional features, the problem of personalized recommendation systems failing to utilize nickname information was solved, achieving more accurate product recommendations and improving user satisfaction.
Patent Information
- Application Number
- CN202510748770.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-16
AI Technical Summary
Existing personalized recommendation systems fail to effectively utilize the multi-dimensional information contained in WeChat nicknames, resulting in insufficient recommendation accuracy.
The pre-trained BERT model is used to semantically encode WeChat nicknames, extract multi-dimensional features such as gender orientation, regional markers, interest tags, emotional tendencies, and social attributes, and use the KMeans clustering algorithm to form user groups. The recommendation system is then combined to make accurate product recommendations.
It improves the accuracy of personalized recommendations and user satisfaction, enhances the understanding of user interests and attributes through multi-dimensional analysis, and improves business results.
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data analysis and processing, and relates to a multi-dimensional data analysis method and system based on WeChat nickname information extraction. Background Art
[0002] With the popularization of social networks and mobile payments, the number of WeChat users is huge. User nicknames, as customized identifiers for users to express themselves externally, often contain implicit characteristics such as users' personal preferences, regional information, and age level.
[0003] Current personalized recommendation systems mainly rely on data such as user purchasing behavior and search history, and often ignore the information carried by the nickname text. In fact, WeChat nicknames may contain multi-dimensional information such as user gender, age, occupational hints or hobbies (such as favorite anime, ball sports, etc.) and emotional tendencies. Extracting and analyzing this information can provide supplementary features for user portraits, improving the accuracy of personalized recommendations and user satisfaction. In application scenarios such as e-commerce, content distribution, and advertising, by using nicknames to mine interests and attributes, we can better meet Tonghu's needs and improve business results. Summary of the Invention
[0004] The purpose of the present invention is to solve the above problems and provide a multi-dimensional data analysis method and system based on WeChat nickname information extraction.
[0005] To achieve the above objectives, the present invention provides a multi-dimensional data analysis method based on WeChat nickname information extraction, comprising:
[0006] Get the user's WeChat nickname and pre-process it;
[0007] The pre-trained BERT model is used to perform semantic encoding on the pre-processed nickname text, extracting multi-dimensional features including gender orientation, regional markers, interest tags, sentiment orientation, and social attributes.
[0008] The extracted feature vectors of each dimension are input into the KMeans clustering algorithm to obtain multiple user groups, and a label portrait is generated for each group based on the clustering results;
[0009] After matching user group tags with product tags, the recommendation system interface is called to achieve accurate product recommendations for different user groups.
[0010] Furthermore, the preprocessing includes word segmentation and noise removal;
[0011] The interest tag extraction is performed by matching a pre-set interest dictionary or based on semantic clustering identification;
[0012] The sentiment tendency extraction calculates the sentiment score of the nickname text based on the sentiment dictionary;
[0013] The social attribute extraction is based on the recognition of occupation or age words in nicknames.
[0014] The present invention also provides an analysis system using the multi-dimensional data analysis method based on WeChat nickname information extraction, comprising:
[0015] The acquisition module is used to obtain user nicknames from the WeChat platform and output them to the preprocessing module;
[0016] Preprocessing module, used to segment nickname text and remove noise;
[0017] The feature extraction module is used to call the BERT model to perform semantic encoding on the preprocessed nickname text, extracting multi-dimensional features such as gender orientation, regional markers, interest tags, emotional orientation, and social attributes;
[0018] The clustering analysis module clusters the extracted multi-dimensional features of users based on the KMeans algorithm to form a labeled user group portrait;
[0019] The recommendation output module is used to match the group labels of the clustering results with the product feature labels, and call the recommendation system interface to output personalized recommendation results.
[0020] Furthermore, it also includes a visualization module for visually displaying the clustering results and user tag distribution in the form of heat maps and bar charts. DETAILED DESCRIPTION
[0021] The present invention provides a method and system for multi-dimensional data analysis based on WeChat nickname information extraction, wherein the analysis system includes an acquisition module, a preprocessing module, a feature extraction module, a cluster analysis module, a visualization module and a recommendation output module.
[0022] The collection module is responsible for acquiring user nickname data from the WeChat platform interface and inputting the WeChat user nickname list into the system for analysis. The preprocessing module cleans and segmentes the collected nickname text, removing noise such as meaningless characters, emoticons, and punctuation. It also performs word segmentation on Chinese nicknames and stemming on English nicknames. The feature extraction module uses a pretrained BERT model to encode the preprocessed nickname text, obtaining a 768-dimensional semantic vector for each nickname. Based on these vectors, a rule engine or classifier is used to extract labels and probability distributions for multiple dimensions, such as gender, region, interests, sentiment, occupational cues, and age terms. The cluster analysis module inputs the user's multidimensional feature vectors into the KMeans clustering algorithm, performs offline training or online updates on all users, and divides them into several user groups (clusters). For each group, a central label is calculated and a group profile is generated. The visualization module visualizes the clustering results and label distribution, such as by using bar charts and heat maps to present the feature label distribution of each group, facilitating analysis by operations personnel. The recommendation output module uses the recommendation system interface to implement targeted recommendations based on clustered group profiles. The system maps user group tag information to product tags, sends the analysis results as input to the existing personalized recommendation engine, and outputs a list of matching products.
[0023] The data flow starts with nickname collection, is processed by various functional modules, and finally the recommendation engine outputs personalized recommendation results.
[0024] The multi-dimensional data analysis methods based on WeChat nickname information extraction include:
[0025] The present invention uses deep semantic understanding technology to extract multi-dimensional features of WeChat nicknames.
[0026] First, a pre-trained BERT (Bidirectional Encoder Representations from Transformers) model is used to semantically encode the user's nickname text to obtain a context-sensitive semantic representation. Subsequent processing of the BERT output vectors allows the identification and extraction of multiple user feature dimensions, including: gender (e.g., nicknames containing cues such as "Mr." and "Ms."), geographic markers (nicknames containing place names or dialect words), interest tags (e.g., nicknames containing hobbies such as "anime," "sneakers," and "football"), sentiment (nicknames containing positive and negative emotional terms), and social attributes (e.g., occupational cues such as "doctor" and "teacher" or age-related terms such as "teenager" and "old").
[0027] In this process, the BERT features can be classified and judged by combining domain knowledge dictionaries or secondary classifiers. In this way, the present invention converts nickname text into multi-dimensional user attribute labels, providing basic features for subsequent personalized recommendations.
[0028] In view of the extracted multi-dimensional features, the present invention further adopts a cluster analysis method to perform group profiling on user groups.
[0029] Specifically, the multi-dimensional feature vector of each user (such as gender orientation score, interest tag weight, sentiment distribution, etc.) is constructed into a feature matrix and trained using the KMeans clustering algorithm.
[0030] The algorithm automatically divides users into several labeled user groups, with each cluster corresponding to a group of users with similar nickname characteristics. The clustering process iteratively calculates cluster centers and continuously optimizes the similarity between users within each cluster, thereby achieving a stable user group classification.
[0031] Once clusters are formed, each cluster can be tagged, for example, labeling a group containing the "sneakers" interest tag as "sneaker enthusiasts." Based on this multidimensional cluster analysis, the system can construct more granular and rich user group profiles, providing a basis for personalized recommendations.
[0032] The BERT model mentioned in the embodiment can adopt a mainstream pre-trained model (such as bert-base-chinese) and use the Transformer architecture for text encoding. The system can load the BERT model through the HuggingFace Transformers library or call the open source language model interface for inference.
[0033] After the user nickname passes through the preprocessing module, it is converted into the input format of the BERT model (such as the WordPiece word segmentation input ID sequence). The BERT model outputs a semantic vector of a fixed length (such as 768).
[0034] The feature extraction module then performs a multi-dimensional analysis of the vector based on a pre-set strategy: Gender orientation can be determined by comparing cosine similarity with pre-labeled gender vectors; interest tags can be identified by matching interest keyword dictionaries or searching for nearest neighbors with interest category vectors in the semantic space; sentiment scores are calculated based on the segmented words in the nickname using a sentiment dictionary to calculate the overall sentiment orientation; and attribute tags are determined based on social attribute keywords such as "engineer," "retired," and "teenager" that appear in the nickname. Each extracted dimensional feature can be quantified as a numerical value or probability (e.g., between 0 and 1) to form the user's multi-dimensional feature vector.
[0035] The cluster analysis module feeds all user feature vectors into the KMeans algorithm. The appropriate number of clusters, K, can be determined using the silhouette coefficient or the elbow rule. During training, the model randomly initializes K cluster centers and, through iterative optimization, continuously reassigns users to the nearest center and updates the center coordinates until convergence or the set number of iterations is reached.
[0036] After clustering, the system analyzes the central features of each cluster and generates representative labels. For example, if a cluster center has high weights in the "sports" and "sneakers" interest dimensions, the user cluster can be labeled "sneaker enthusiasts." The visualization module uses heat maps to display the weight distribution of different clusters in each label dimension for operational reference.
[0037] The recommendation output module inputs the generated user group tags into the existing recommendation system interface. In implementation, the user group tags are first mapped to product attribute tags. For example, the group "Sneaker Enthusiasts" would be mapped to product tags such as "Sneakers" and "Ball Sports Equipment." The system then retrieves and sorts products with these tags, then outputs a personalized recommendation list.
[0038] For specific matching rules, tag similarity-based matching or keyword retrieval in product descriptions can be used. The system can use recommendation engine APIs (such as recommendation services based on collaborative filtering or deep learning) to use user group tags as side information and combine them with user historical behavior to generate final recommendation results. Through the above modular design and specific implementation, the present invention can fully implement multi-dimensional feature extraction, cluster analysis, and precise recommendations based on WeChat nicknames.
[0039] The multi-dimensional data analysis method and system based on WeChat nickname information extraction of the present invention can mine interests and attributes through user nicknames to improve the accuracy of personalized recommendations and user satisfaction.
[0040] The above description is merely one embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A multi-dimensional data analysis method based on WeChat nickname information extraction, characterized in that: include: Get the user's WeChat nickname and pre-process it; The pre-trained BERT model is used to perform semantic encoding on the pre-processed nickname text, extracting multi-dimensional features including gender orientation, regional markers, interest tags, sentiment orientation, and social attributes. The extracted feature vectors of each dimension are input into the KMeans clustering algorithm to obtain multiple user groups, and a label portrait is generated for each group based on the clustering results; After matching user group tags with product tags, the recommendation system interface is called to achieve accurate product recommendations for different user groups.
2. The multidimensional data analysis method based on WeChat nickname information extraction according to claim 1 is characterized in that: The preprocessing includes word segmentation and noise removal; The interest tag extraction is performed by matching a pre-set interest dictionary or based on semantic clustering identification; The sentiment tendency extraction calculates the sentiment score of the nickname text based on the sentiment dictionary; The social attribute extraction is based on the recognition of occupation or age words in nicknames.
3. An analysis system using the multi-dimensional data analysis method based on WeChat nickname information extraction according to claims 1-2, characterized in that: include: The acquisition module is used to obtain user nicknames from the WeChat platform and output them to the preprocessing module; Preprocessing module, used to segment nickname text and remove noise; The feature extraction module is used to call the BERT model to perform semantic encoding on the preprocessed nickname text, extracting multi-dimensional features such as gender orientation, regional markers, interest tags, emotional orientation, and social attributes; The clustering analysis module clusters the extracted multi-dimensional features of users based on the KMeans algorithm to form a labeled user group portrait; The recommendation output module is used to match the group labels of the clustering results with the product feature labels, and call the recommendation system interface to output personalized recommendation results.
4. The analysis system according to claim 3, characterized in that It also includes a visualization module for visualizing clustering results and user tag distribution through heat maps and histograms.