User layering method and device and storage medium

By collecting multi-source data, performing deep learning and social network analysis, and combining deep autoencoders and the DBSCAN algorithm, the problem of traditional user segmentation methods being unable to deeply explore user value and capture user changes in real time has been solved, achieving accurate user segmentation and dynamic optimization.

CN120929911APending Publication Date: 2025-11-11BEIJING QICHUANG TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511029363.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional user segmentation methods cannot deeply explore users' behavioral preferences and potential value, cannot fully reflect the interaction between users and the platform, and cannot capture user changes in real time.

Method used

Collect multi-source data, including basic user information, behavioral data, and social relationship data. Build a user value assessment model through deep learning and social network analysis. Combine deep autoencoders and DBSCAN algorithm to segment users and dynamically optimize the model through a stream processing framework.

Benefits of technology

It achieves precise user segmentation, fully reflects user value, improves operational efficiency and effectiveness, and can capture user changes in real time and optimize segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929911A_ABST
    Figure CN120929911A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of user layering methods, in particular to a user layering method and device and a storage medium, and the method specifically comprises the following steps: 1, collecting basic information data of users, including names, ages, genders, contact information and registration time, collecting behavior data of the users, and collecting social relation data of the users, cleaning the basic data and the behavior data according to the social relation data; 2, modeling a user behavior sequence; step 3, analyzing the user social network; 4, establishing a user value evaluation model; 5, layering the users based on deep learning; step 6, carrying out dynamic optimization on a layering result; step 7, applying and feeding back a layering result; according to the method, the problem of single traditional data is solved through multi-source data acquisition and processing, and accurate features are comprehensively extracted; behavior sequence modeling and social network analysis break through the limitation that only individual behaviors are concerned, and users with similar behaviors and social contact are deeply mined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of user stratification methods, and in particular to a user stratification method, apparatus, and storage medium. Background Technology

[0002] With the rapid development of the internet industry, the user base continues to expand, and user needs are becoming increasingly diversified and personalized. Against this backdrop, accurate user segmentation has become a core element for enterprises to improve operational efficiency, optimize user experience, and maximize business value.

[0003] Traditional user segmentation methods, such as simple classification based on demographic characteristics, can only superficially divide users and cannot delve into their behavioral preferences and potential value. On the other hand, segmentation methods based on a single behavioral indicator (such as visit duration or number of purchases) are too one-sided and cannot fully reflect the interaction between users and the platform. Summary of the Invention

[0004] To address the technical problem that traditional user segmentation methods are insufficient to fully reflect the interaction between users and the platform, this invention provides a user segmentation method, device, and storage medium.

[0005] The technical solution adopted in this invention is: a user segmentation method, specifically including the following steps:

[0006] Step 1: Collect basic user information data, including name, age, gender, contact information, and registration time; collect user behavior data; collect user social relationship data; and clean the basic and behavioral data.

[0007] Step 2: Model user behavior sequences;

[0008] Step 3: Analyze the user's social network;

[0009] Step 4: Establish a user value assessment model;

[0010] Step 5: Segment users based on deep learning;

[0011] Step 6: Dynamically optimize the stratification results.

[0012] Step 7: Apply and provide feedback on the stratified results.

[0013] In one embodiment, in step 1, the set U = {u1, u2, ..., u...} is used. n} represents a set of users, where each user u i The corresponding basic information vector is I i =(i i1 i i2,…,i ik ),i ij Indicates user u i The j-th basic information item;

[0014] Collect user behavior data, covering users' browsing, clicking, collection, purchase, commenting and sharing behavior on the platform, and record the time, location and target information of the behavior;

[0015] Let user u i The behavioral sequence is B i ={b i1 ,b i2 ,…,b im}, where b ij =(t ij ,a ij ,o ij ),t ij For the time when the behavior occurs, a ij For behavior type, o ij For the behavior object;

[0016] Collect users' social relationship data, including their friend lists, followed individuals, number of followers, and interaction frequency, to construct a user social network graph G = (V, E), where V is the set of user nodes, E is the set of social relationship edges between users, and the weight of each edge is w. ij Indicates user u i and u j The intensity of interaction between them;

[0017] The data cleaning methods are as follows:

[0018] Remove duplicate data, outliers, and missing values;

[0019] For missing values, a Bayesian estimation-based method is used for imputation;

[0020] Convert non-numerical features into numerical features. Let the set of behavior types be A = {a1, a2, ..., a...}. p}, then behavior a j The one-hot encoded vector is v j = (0,…,1,…,0), where the j-th bit is 1 and the rest are 0;

[0021] Feature dimensionality reduction: Principal component analysis is used to reduce the dimensionality of high-dimensional features;

[0022] Let the original feature matrix be X. n×d Where n is the number of samples and d is the feature dimension;

[0023] Calculate the covariance matrix of the characteristic matrix in This is the sample mean vector;

[0024] Performing eigenvalue decomposition on the covariance matrix C yields eigenvalues ​​λ1≥λ2≥…≥λ d and the corresponding eigenvectors e1, e2, ..., e d ;

[0025] The projection matrix E is formed by selecting the eigenvectors corresponding to the k largest eigenvalues. d×k Then the dimensionality-reduced feature matrix is ​​Y. n ×k=X×E.

[0026] In one embodiment, step 2, the specific method for modeling user behavior sequences is as follows:

[0027] The user's behavior sequence is converted into a vector representation. The Skip-Gram algorithm in the Word2Vec model is used to treat each behavior as a "word" and the user's behavior sequence as a "sentence". The vector representation of each behavior is obtained by training the model.

[0028] Let the sequence of actions be b1, b2, ..., b m The goal is to maximize the conditional probability P(b) i -c,…,b i-1 ,b i+1 ,…,b i+c |b i ), where c is the context window size, and the behavior vector v(b) is obtained through neural network training. i This vector can capture the semantic relationships between behaviors;

[0029] The similarity between different user behavior sequences is calculated using a dynamic time warping algorithm;

[0030] Let user u i The behavior sequence vector is V i =(v1,v2,…,v m ), User u j The behavior sequence vector is V j =(w1,w2,…,w n The cumulative distance D(i,j) between the two is defined as follows:

[0031] i,j)=min(D(i-1,j),D(i,j-1),D(i-1,j-1))+dis;

[0032] Where dist(v i ,w j ) is a vector v i and w jThe Euclidean distance between the two action sequences, D(m,n), is the DTW distance between them.

[0033] In one embodiment, step 3, the specific method for analyzing the user's social network, is as follows:

[0034] Calculate the degree centrality C of a user in a social network D (u i ), that is, user u i The ratio of the number of direct connections to the total number of users in the network minus 1: where deg(u i ) for user u i The degree, where N is the total number of users in the network;

[0035] Calculate the user's closeness centrality C C (u i ), that is, user u i The reciprocal of the average shortest path distance to all other users in the network: Where d(u i ,u j ) for user u i to u j The shortest path distance;

[0036] Calculate the user's betweenness centrality C B (u i That is, all shortest paths in the network pass through user u. i Ratio: Where σ st Let σ be the total number of shortest paths from s to t. st (u i ) represents the distance from s to t via u i The number of shortest paths; Louvain's algorithm is used for community detection, which uses modularity Q as the optimization objective. The formula for calculating modularity is: Where A ij k is an element of the adjacency matrix. i Let m be the degree of node i, m be the total number of edges in the network, and c be the degree of node i. i For the community to which node i belongs, δ(c i ,c j ) is an indicator function, when c i =c j It is 1 if it is true, otherwise it is 0.

[0037] In one embodiment, step 4, establishing the user value assessment model, is specifically done as follows:

[0038] Let user u i The total amount spent within time T is M.i The number of transactions is F. i The average consumption cycle is C i Then the transaction value

[0039]

[0040] Where α, β, γ are weight coefficients, determined by the analytic hierarchy process, and satisfy α + β + γ = 1;

[0041] Based on user behavior sequences and social network analysis results, consider the number of user comments, the number of shares, and the frequency of social interactions;

[0042] Let user u i The number of comments is R i The number of shares is S i The frequency of social interaction is I i Then the interactive value V I (u i ):

[0043]

[0044] Where δ,∈,ζ are weight coefficients, and δ+∈+ζ=1;

[0045] This is measured by predicting users' future behavior and value. A Long Short-Term Memory (LSTM) network is used to predict users' future spending and interaction frequency. Let the predicted future spending amount be... The predicted number of future interactions is Then the potential value V P (u i ):

[0046]

[0047] Where η and θ are weighting coefficients, and η+θ=1;

[0048] The user's overall value V(u) is calculated using a weighted summation method. i ):

[0049] V(u i )=ω T ×V T (u i )+ω I ×V I (u i )+ω P ×V P (u i );

[0050] Where ω T ,ω I ,ω PThe weights for each value dimension are determined using the entropy weight method.

[0051] In one embodiment, step 5, the specific method for user segmentation based on deep learning, is as follows:

[0052] A deep autoencoder model is constructed, consisting of an encoder and a decoder. The encoder maps the input user feature vector x to a low-dimensional latent space, obtaining the latent vector h; the decoder reconstructs the latent vector h back into the input vector. The model's loss function is the reconstruction error. By training the model, the encoder can learn the user's key features, with the latent vector h serving as the user's compressed feature representation.

[0053] The DBSCAN algorithm is used for user segmentation. The DBSCAN algorithm divides clusters by defining concepts such as core objects, density-accessible, density-reachable, and density-connected. A core object is an object containing at least MinPts samples within a radius ∈ [0, 1], and the user's feature vector is denoted as h. i Calculate the distance d(h) from each user to other users. i ,h j Users are divided into different clusters based on distance and parameters ∈ and MinPts, with each cluster representing a user hierarchy.

[0054] In one embodiment, step 6, the specific method for dynamically optimizing the stratification results, is as follows:

[0055] Real-time collection and analysis of user behavior data, transaction data, and social data;

[0056] A stream processing framework is used to process real-time data. A data update time window Δt is set. When the data in the window accumulates to a certain amount or reaches a time threshold, the update of user characteristics and value is triggered.

[0057] Incremental learning updates use an incremental learning algorithm to dynamically update the user hierarchical model. For deep autoencoder models, when new user data is input, only some parameters of the model are updated, instead of retraining the entire model, thus reducing computational overhead.

[0058] Let the new training data be Dnew and the old model parameters be θold. Update the parameters using gradient descent: Where λ is the learning rate. This is the gradient of the loss function.

[0059] The stratification results are adjusted based on the updated user characteristics and value assessment results. User stratification is adjusted by calculating the distance between a user's current feature vector and the center of their current cluster. When the distance exceeds a set threshold, the user is reassigned to the nearest cluster. Simultaneously, the clustering results are periodically evaluated using the silhouette coefficient S. Where a i Let b be the average distance between sample i and other samples in the same cluster. i The average distance between sample i and the nearest heterogeneous sample is used. When the silhouette coefficient is lower than the set threshold, the cluster analysis is re-performed to optimize the user stratification results.

[0060] In one embodiment, the specific method for applying and providing feedback on the stratified results in step 7 is as follows:

[0061] The tiered application strategy targets high-value users, providing personalized services and exclusive offers.

[0062] The system of performance feedback and model iteration establishes a tiered performance evaluation index system, including user retention rate, conversion rate, and average spending.

[0063] The beneficial effects of this invention are as follows: Compared with the prior art, this invention solves the problem of single data in traditional methods by collecting and processing multi-source data, and comprehensively extracts accurate features; behavioral sequence modeling and social network analysis break through the limitation of only focusing on individual behavior, and deeply explore users with similar behaviors and social interactions; multi-dimensional value assessment overcomes the defects of single dimensions, and can comprehensively reflect the interaction relationship between users and the platform, and accurately reflect user value; deep learning and density clustering improve the accuracy of hierarchical analysis, dynamic optimization ensures timeliness, and application feedback promotes continuous optimization, effectively solving the problems of traditional methods being one-sided, difficult to process multi-source data, and unable to capture user changes in real time, thereby improving operational efficiency and effectiveness. Attached Figure Description

[0064] Figure 1 This is a flowchart illustrating the present invention. Detailed Implementation

[0065] In the description of this invention, it should be noted that the terms "front", "up", "down", "left", "right", "vertical", "horizontal", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0066] refer to Figure 1 To address the problems existing in the background technology, this application proposes the following technical solution: a user segmentation method, specifically including the following steps:

[0067] Step 1: Collect basic user information data, including name, age, gender, contact information, and registration time; collect user behavior data; collect user social relationship data; and clean the basic and behavioral data.

[0068] In step 1, the set U = {u1, u2, ..., u} is used. n} represents a set of users, where each user u i The corresponding basic information vector is I i =(i i1 i i2 ,…,i ik ),i ij Indicates user u i The j-th basic information item;

[0069] Collect user behavior data, covering users' browsing, clicking, collection, purchase, commenting and sharing behavior on the platform, and record information such as the time, location and object of the behavior;

[0070] Let user u i The behavioral sequence is B i ={b i1 ,b i2 ,…,b im}, where b ij =(t ij, a ij ,o ij ),t ij For the time when the behavior occurs, a ij For behavior type, o ij For the behavior object;

[0071] Collect users' social relationship data, including their friend lists, followed individuals, number of followers, and interaction frequency, to construct a user social network graph G = (V, E), where V is the set of user nodes, E is the set of social relationship edges between users, and the weight of each edge is w. ij Indicates user u i and u j The intensity of interaction between them;

[0072] The data cleaning methods are as follows:

[0073] Remove duplicate data, outliers, and missing values;

[0074] For missing values, a Bayesian estimation-based method is used for imputation;

[0075] Let the missing value of a certain feature x be x. mGiven that the probability distribution model of this feature is P(x|θ), where θ is the model parameter, the parameter θ is estimated using existing data samples, and then x is calculated based on P(x|θ). m To fill in the data, make the filled data more consistent with the true distribution of the features;

[0076] Non-numerical features are converted into numerical features. For example, for gender features, 0 represents female and 1 represents male; for behavioral type features, one-hot encoding is used for conversion. Let the set of behavioral types be A = {a1, a2, ..., a...}. p}, then behavior a j The one-hot encoded vector is v j = (0,…,1,…,0), where the j-th bit is 1 and the rest are 0;

[0077] Feature dimensionality reduction: Principal component analysis (PCA) is used to reduce the dimensionality of high-dimensional features;

[0078] Let the original feature matrix be X. n×d Where n is the number of samples and d is the feature dimension;

[0079] Calculate the covariance matrix of the characteristic matrix in This is the sample mean vector;

[0080] Performing eigenvalue decomposition on the covariance matrix C yields eigenvalues ​​λ1≥λ2≥…≥λ d and the corresponding eigenvectors e1, e2, ..., e d ;

[0081] The projection matrix E is formed by selecting the eigenvectors corresponding to the k largest eigenvalues. d×k Then the dimensionality-reduced feature matrix is ​​Y. n ×k = X×E. Dimensionality reduction reduces the feature dimensions, lowers computational complexity, and preserves the main information of the data.

[0082] The above technical solution is explained as follows: This step collects multi-source data on user base, behavior, and social relationships, and processes it through cleaning, transformation, and dimensionality reduction. This comprehensively integrates the data, solving the problem of single-source data in traditional methods, extracting more accurate features, providing a reliable foundation for subsequent stratification and value assessment, and improving data quality and usability.

[0083] Step 2: Model user behavior sequences;

[0084] In step 2, the specific method for modeling user behavior sequences is as follows:

[0085] The user's behavior sequence is converted into a vector representation. The Skip-Gram algorithm in the Word2Vec model is used to treat each behavior as a "word" and the user's behavior sequence as a "sentence". The vector representation of each behavior is obtained by training the model.

[0086] Let the sequence of actions be b1, b2, ..., b m The goal is to maximize the conditional probability P(b) i -c,…,b i-1 ,b i+1 ,…,b i+c |b i ), where c is the context window size, and the behavior vector v(b) is obtained through neural network training. i This vector can capture the semantic relationships between behaviors;

[0087] The similarity between different user behavior sequences is calculated using the Dynamic Time Warping (DTW) algorithm. Let user u... i The behavior sequence vector is V i =(v1,v2,…,v m ), User u j The behavior sequence vector is V j =(w1,w2,…,w n The cumulative distance D(i,j) between the two is defined as follows:

[0088] i,j)=min(D(i-1,j),D(i,j-1),D(i-1,j-1))+dis

[0089] Where dist(v i ,w j ) is a vector v i and w j The Euclidean distance between the two action sequences is D(m,n). D(m,n) is the DTW distance between the two action sequences. The smaller the distance, the more similar the two action sequences are.

[0090] The above technical solution is explained as follows: The Word2Vec Skip-Gram algorithm vectorizes behavioral sequences, and the DTW algorithm calculates similarity. This captures semantic associations and sequence similarities of behaviors, overcoming the limitations of traditional methods that only consider single behaviors. It can identify users with similar behavioral patterns, providing behavioral dimensions for stratification and improving the accuracy of stratification.

[0091] Step 3: Analyze the user's social network;

[0092] Step 3, the specific methods for analyzing users' social networks are as follows:

[0093] Calculate the degree centrality C of a user in a social networkD (u i ), that is, user u i The ratio of the number of direct connections to the total number of users in the network minus 1: where deg(u i ) for user u i Degree centrality is defined as the degree of a user in a social network, where N is the total number of users in the network. Degree centrality reflects a user's direct influence in the social network.

[0094] Calculate the user's closeness centrality C C (u i ), that is, user u i The reciprocal of the average shortest path distance to all other users in the network: Where d(u i ,u j ) for user u i to u j The shortest path distance. Closeness centrality measures how close a user is to other users in the network.

[0095] Calculate the user's betweenness centrality C B (u i That is, all shortest paths in the network pass through user u. i Ratio: Where σ st Let σ be the total number of shortest paths from s to t. st (u i ) represents the distance from s to t via u i The number of shortest paths. Betweenness centrality reflects the mediating role of users in the network.

[0096] The Louvain algorithm is used for community detection. This algorithm uses the modularity Q as the optimization objective, and the formula for calculating the modularity is: Where A ij k is an element of the adjacency matrix. i Let m be the degree of node i, m be the total number of edges in the network, and c be the degree of node i. i For the community to which node i belongs, δ(c i ,c j ) is an indicator function, when c i =c j The value is 1 if the modularity is active, and 0 otherwise. By iteratively optimizing the modularity, the social network is divided into multiple communities, where users within each community have strong social connections.

[0097] The above technical solution is explained as follows: It calculates user social centrality and uses the Louvain algorithm to discover communities. By considering user social connections and community affiliation, it overcomes the limitation of focusing solely on individual behavior, identifies user groups with similar social attributes, supports socialized operations, and enhances the reference for layered social dimensions.

[0098] Step 4: Establish a user value assessment model;

[0099] Step 4, the specific method for establishing a user value assessment model is as follows:

[0100] Let user u i The total amount spent within time T is M. i The number of transactions is F. i The average consumption cycle is C i Then the transaction value

[0101]

[0102] Where α, β, and γ are weight coefficients, determined by the analytic hierarchy process, and satisfy α + β + γ = 1.

[0103] Interactive value: Based on user behavior sequences and social network analysis results, it considers factors such as the number of comments, shares, and frequency of social interactions. Let user u... i The number of comments is R i The number of shares is S i The frequency of social interaction is I i Then interactive value

[0104]

[0105] Where δ,∈,ζ are weight coefficients, and δ+∈+ζ=1.

[0106] Potential value: Measured by predicting future user behavior and value. A Long Short-Term Memory (LSTM) network is used to predict future user spending and interaction frequency. Let the predicted future spending amount be... The predicted number of future interactions is Then potential value

[0107]

[0108] Where η and θ are weighting coefficients, and η + θ = 1. 2. Comprehensive Value Calculation - The comprehensive value of the user is calculated using a weighted summation method.

[0109] V(u i )=ω T ×V T (u i )+ω I ×VI (u i )+ω P ×V P (u i )

[0110] Where ω T ,ω I ,ω P The weights for each value dimension are determined using the entropy weight method. The calculation steps of the entropy weight method are as follows: First, calculate the entropy value of the j-th indicator.

[0111]

[0112] in Then calculate the weights of the indicators.

[0113]

[0114] Where m represents the number of indicators.

[0115] The above technical solution can be explained as follows: It constructs multi-dimensional value encompassing transactions, interactions, and potential, and then weights these dimensions to arrive at a comprehensive value. Compared to traditional single-dimensional assessments, this more comprehensively reflects the true value and potential of users, helping businesses accurately identify high-value and high-potential users and providing a scientific basis for resource allocation and strategy formulation.

[0116] Step 5: Segment users based on deep learning;

[0117] Step 5, the specific method for user segmentation based on deep learning is as follows:

[0118] A deep autoencoder model is constructed, consisting of an encoder and a decoder. The encoder maps the input user feature vector x to a low-dimensional latent space, obtaining the latent vector h; the decoder reconstructs the input vector from the latent vector h. The model's loss function is the reconstruction error. By training the model, the encoder can learn the user's key features, with the latent vector h serving as a compressed feature representation of the user.

[0119] The DBSCAN (Density-Based Spatial Clustering for Applications and Noise) algorithm is used for user stratification. The DBSCAN algorithm divides clusters by defining concepts such as core objects, density reachability, density accessibility, and density connectivity. Core objects are defined as objects containing at least MinPts samples within a radius ∈ [0, 1], respectively. Let the user's feature vector be h. i Calculate the distance d(h) from each user to other users. i ,h j Based on distance and parameters ∈ and MinPts, users are divided into different clusters, with each cluster representing a user hierarchy.

[0120] The above technical solution is explained as follows: Features are extracted using a deep autoencoder, combined with the DBSCAN algorithm for stratification. This approach can handle high-dimensional nonlinear data, improve the accuracy and stability of stratification, overcome the sensitivity of traditional clustering to initial values, and make the stratification results more closely reflect the actual characteristics and value of the user.

[0121] Step 6: Dynamically optimize the stratification results;

[0122] Step 6, the specific method for dynamically optimizing the stratification results is as follows:

[0123] Real-time collection and analysis of user behavior, transaction, and social data. A stream processing framework (such as Apache Flink) is used to process real-time data, setting a data update time window Δt. When the data within the window accumulates to a certain amount or reaches a time threshold, updates to user characteristics and value are triggered.

[0124] Incremental learning updates dynamically update the user hierarchical model using an incremental learning algorithm. For deep autoencoder models, when new user data is input, only a subset of the model's parameters are updated, rather than retraining the entire model, thus reducing computational overhead. Let the new training data be θnew and the old model parameters be θold. The parameters are updated using gradient descent: Where λ is the learning rate. This is the gradient of the loss function.

[0125] The stratification results are adjusted based on the updated user characteristics and value assessment results. The distance between a user's current feature vector and the center of their current cluster is calculated. When the distance exceeds a set threshold, the user is reassigned to the nearest cluster. Simultaneously, the clustering results are periodically evaluated using the silhouette coefficient S. Where a i For sample a i The average distance to other samples in the same cluster, b i For sample a i The average distance to the nearest heterogeneous cluster sample. When the silhouette coefficient falls below a set threshold, the cluster analysis is re-performed to optimize the user stratification results.

[0126] The above technical solution is explained as follows: Real-time data monitoring and incremental learning are used to update the model and adjust the stratification results. This allows for timely capture of dynamic changes in user behavior, ensuring the timeliness of the stratification results. This helps businesses adjust their operational strategies promptly, avoid missing opportunities, and improve operational flexibility and effectiveness.

[0127] Step 7: Apply and provide feedback on the stratified results.

[0128] Step 7, the specific methods for applying and providing feedback on the stratified results are as follows:

[0129] The tiered application strategy targets high-value users, providing personalized services and exclusive offers such as VIP customer service and double points to improve user loyalty and retention. For mid-value users, a recommendation system pushes relevant products and services to encourage increased spending and interaction, elevating their value level. For low-value users, targeted marketing and reactivation campaigns are conducted, such as sending coupons and pushing content of interest, to stimulate user activity and spending potential.

[0130] The evaluation system for user segmentation is based on feedback and model iteration. It includes metrics such as user retention rate, conversion rate, and average spending. With an evaluation period of T, the metric values ​​for different user segments within the evaluation period are calculated and compared with the values ​​before segmentation to analyze the effectiveness of the segmentation strategy. Based on the evaluation results, the user segmentation model and application strategy are iteratively optimized. If the user retention rate of a particular segment decreases, the reasons are analyzed and the service strategy for that segment is adjusted. If the model's prediction accuracy decreases, the model's parameters and feature selection are readjusted to improve model performance.

[0131] The above technical solution is explained as follows: Strategies are developed for different tiers, and an iterative evaluation model is established. By integrating tiering with operations, feedback is used to optimize the model and strategies, improving user retention and conversion rates, maximizing the value of each tier, and promoting continuous operational optimization for the enterprise.

[0132] In summary, multi-source data collection and processing address the limitations of traditional single-source data, comprehensively extracting accurate features; behavioral sequence modeling and social network analysis overcome the limitations of focusing solely on individual behavior, uncovering users with similar behaviors and social interactions; multi-dimensional value assessment overcomes the shortcomings of single-dimensional methods, accurately reflecting user value; deep learning and density clustering improve the accuracy of hierarchical analysis, dynamic optimization ensures timeliness, and application feedback promotes continuous optimization. This effectively solves the problems of traditional methods being one-sided, difficult to process multi-source data, and unable to capture user changes in real time, thereby improving operational efficiency and effectiveness.

[0133] Although embodiments of the invention have been shown and described, the scope of the invention will be defined by the appended claims and their equivalents by those skilled in the art.

Claims

1. A method for user segmentation, characterized in that, Specifically, the following steps are included: Step 1: Collect basic user information data, including name, age, gender, contact information, and registration time; collect user behavior data; collect user social relationship data; and clean the basic and behavioral data. Step 2: Model user behavior sequences; Step 3: Analyze the user's social network; Step 4: Establish a user value assessment model; Step 5: Segment users based on deep learning; Step 6: Dynamically optimize the stratification results; Step 7: Apply and provide feedback on the stratified results.

2. The user segmentation method according to claim 1, characterized in that, In step 1, the set U = {u1, u2, ..., u} is used. n } represents a set of users, where each user u i The corresponding basic information vector is I i =(i i1 i i2 ,…,i ik ),i ij Indicates user u i The j-th basic information item; Collect user behavior data, covering users' browsing, clicking, collection, purchase, commenting and sharing behavior on the platform, and record the time, location and target information of the behavior; Let user u i The behavioral sequence is B i ={b i1 ,b i2 ,…,b im }, where b ij =(t ij ,a ij ,o ij ),t ij For the time when the behavior occurs, a ij For behavior type, o ij For the behavior object; Collect users' social relationship data, including their friend lists, followed individuals, number of followers, and interaction frequency, to construct a user social network graph G = (V, E), where V is the set of user nodes, E is the set of social relationship edges between users, and the weight of each edge is w. ij Indicates user u i and u j The intensity of interaction between them; The data cleaning methods are as follows: Remove duplicate data, outliers, and missing values; For missing values, a Bayesian estimation-based method is used for imputation; Convert non-numerical features into numerical features. Let the set of behavior types be A = {a1, a2, ..., a...}. p }, then behavior a j The one-hot encoded vector is v j = (0,…,1,…,0), where the j-th bit is 1 and the rest are 0; Feature dimensionality reduction: Principal component analysis is used to reduce the dimensionality of high-dimensional features; Let the original feature matrix be x. n×d Where n is the number of samples and d is the feature dimension; Calculate the covariance matrix of the characteristic matrix in This is the sample mean vector; Performing eigenvalue decomposition on the covariance matrix C yields eigenvalues ​​λ1≥λ2≥…≥λ d and the corresponding eigenvectors e1, e2, ..., e d ; The projection matrix E is formed by selecting the eigenvectors corresponding to the k largest eigenvalues. d×k Then the dimensionality-reduced feature matrix is ​​Y. n ×k=X×E.

3. The user segmentation method according to claim 2, characterized in that, In step 2, the specific method for modeling user behavior sequences is as follows: The user's behavior sequence is converted into a vector representation. The Skip-Gram algorithm in the Word2Vec model is used to treat each behavior as a "word" and the user's behavior sequence as a "sentence". The vector representation of each behavior is obtained by training the model. Let the sequence of actions be b1, b2, ..., b m The goal is to maximize the conditional probability P(b) i -c,…,b i-1 ,b i+1 ,…,b i+c |b i ), where c is the context window size, and the behavior vector v(b) is obtained through neural network training. i This vector can capture the semantic relationships between behaviors; The similarity between different user behavior sequences is calculated using a dynamic time warping algorithm; Let user u i The behavior sequence vector is V i =(v1,v2,…,v m ), User u j The behavior sequence vector is V j =(w1,w2,…,w n The cumulative distance D(i,j) between the two is defined as follows: i,j)=min(D(i-1,j),D(i,j-1),D(i-1,j-1))+dis; Where dist(v i ,w j ) is a vector v i and w j The Euclidean distance between the two action sequences, D(m,n), is the DTW distance between them.

4. The user segmentation method according to claim 3, characterized in that, Step 3, the specific methods for analyzing users' social networks are as follows: Calculate the degree centrality C of a user in a social network D (u i ), that is, user u i The ratio of the number of direct connections to the total number of users in the network minus 1: where deg(u i ) for user u i The degree, where N is the total number of users in the network; Calculate the user's closeness centrality C C (u i ), that is, user u i The reciprocal of the average shortest path distance to all other users in the network: Where d(u i ,u j ) for user u i to u j The shortest path distance; Calculate the user's betweenness centrality C B (u i That is, all shortest paths in the network pass through user u. i Ratio: Where σ st Let σ be the total number of shortest paths from s to t. st (u i ) represents the distance from s to t via u i The number of shortest paths; The Louvain algorithm is used for community detection. This algorithm uses the modularity Q as the optimization objective, and the formula for calculating the modularity is: Where A ij k is an element of the adjacency matrix. i Let m be the degree of node i, m be the total number of edges in the network, and c be the degree of node i. i For the community to which node i belongs, δ(c i ,c j ) is an indicator function, when c i =c j It is 1 if it is true, otherwise it is 0.

5. The user segmentation method according to claim 4, characterized in that, Step 4, the specific method for establishing a user value assessment model is as follows: Let user u i The total amount spent within time T is M. i The number of transactions is F. i The average consumption cycle is C i Then the transaction value Where α, β, γ are weight coefficients, determined by the analytic hierarchy process, and satisfy α + β + γ = 1; Based on user behavior sequences and social network analysis results, consider the number of user comments, the number of shares, and the frequency of social interactions; Let user u i The number of comments is R i The number of shares is S i The frequency of social interaction is I i Then the interactive value V I (u i ): Where δ,∈,ζ are weight coefficients, and δ+∈+ζ=1; This is measured by predicting users' future behavior and value. A Long Short-Term Memory (LSTM) network is used to predict users' future spending and interaction frequency. Let the predicted future spending amount be... The predicted number of future interactions is Then the potential value V P (u i ): Where η and θ are weighting coefficients, and η+θ=1; The user's overall value V(u) is calculated using a weighted summation method. i ): V(u i )=ω T ×V T (u i )+ω I ×V I (u i )+ω P ×V P (u i ); Where ω T ,ω I ,ω P The weights for each value dimension are determined using the entropy weight method.

6. The user segmentation method according to claim 5, characterized in that, Step 5, the specific method for user segmentation based on deep learning is as follows: A deep autoencoder model is constructed, consisting of an encoder and a decoder. The encoder maps the input user feature vector x to a low-dimensional latent space, obtaining the latent vector h; the decoder reconstructs the latent vector h back into the input vector. The model's loss function is the reconstruction error. By training the model, the encoder learns the user's key features, and the latent vector h serves as the compressed feature representation of the user. The DBSCAN algorithm is used for user segmentation. The DBSCAN algorithm divides clusters by defining the concepts of core objects, density-accessible, density-reachable, and density-connected. A core object is an object containing at least MinPts samples within a radius ∈ [0, 1], and the user's feature vector is denoted as h. i Calculate the distance d(h) from each user to other users. i ,h j Users are divided into different clusters based on distance and parameters ∈ and MinPts, with each cluster representing a user hierarchy.

7. A user segmentation method according to claim 6, characterized in that, Step 6, the specific method for dynamically optimizing the stratification results is as follows: Real-time collection and analysis of user behavior data, transaction data, and social data; A stream processing framework is used to process real-time data. A data update time window Δt is set. When the data in the window accumulates to a certain amount or reaches a time threshold, the user characteristics and value are updated. Incremental learning updates use an incremental learning algorithm to dynamically update the user hierarchical model. For deep autoencoder models, when new user data is input, only some parameters of the model are updated, instead of retraining the entire model, thus reducing computational overhead. Let the new training data be Dnew and the old model parameters be θold. Update the parameters using gradient descent: Where λ is the learning rate. The gradient of the loss function; The stratification results are adjusted based on the updated user characteristics and value assessment results. User stratification is adjusted by calculating the distance between a user's current feature vector and the center of their current cluster. When the distance exceeds a set threshold, the user is reassigned to the nearest cluster. Simultaneously, the clustering results are periodically evaluated using the silhouette coefficient S. Where a i Let b be the average distance between sample i and other samples in the same cluster. i The average distance between sample i and the nearest heterogeneous sample is used. When the silhouette coefficient is lower than the set threshold, the cluster analysis is repeated.

8. A user segmentation method according to claim 7, characterized in that, In step 7, the specific methods for applying and providing feedback on the stratified results are as follows: By employing a tiered application strategy, we provide personalized services and exclusive offers to high-value users.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and the processor communicates with the memory via the bus when the electronic device is in operation, and the machine-readable instructions, when executed by the processor, perform the steps of the user tiering method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the user layering method as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Aircraft pitching vibration self-adaptive control system

    CN121523063A