User group clustering method and apparatus, computer device and storage medium

By receiving user group clustering requests, acquiring and transforming feature data for clustering, the problem of not considering audience interests in ad placement is solved, and more accurate user group segmentation and ad resource optimization are achieved.

CN114820011BActive Publication Date: 2026-01-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110083954.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-21
Publication Date
2026-01-13
Estimated Expiration
2041-01-21

AI Technical Summary

Technical Problem

Existing advertising methods fail to adequately consider the interests of the target audience, leading to a waste of resources.

Method used

By receiving user group clustering requests, the system obtains the feature data of the initial user group data, performs feature transformation and filtering, uses the transformed feature data to perform clustering, obtains the clustering results, and performs insightful statistics.

Benefits of technology

It enables more precise user segmentation, reduces waste of advertising resources, and provides better guidance for ad placement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114820011B_ABST
    Figure CN114820011B_ABST
Patent Text Reader

Abstract

The application relates to a user group clustering method and device, computer equipment and a storage medium, which can be applied to a cloud server. The method comprises the following steps: receiving a user group clustering request, reading initial user group data carried in the user group clustering request; obtaining initial feature data of the initial user group data; performing feature transformation on the initial feature data to obtain transformed feature data; and clustering the initial user group data based on the transformed feature data to obtain a clustering result. The user group clustering method performs feature transformation after obtaining the feature data of the user group, only retains features more related to the initial user group to be clustered, and finally clusters the initial user group by using the features, so that the user group can be better clustered and divided, and the clustering result is obtained. The users in each clustering result have similar user features, which can guide the advertisement delivery and reduce the waste of advertisement resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, and storage medium for clustering user groups. Background Technology

[0002] Advertising plays a crucial role in shaping brand image, helping products establish a reputation, cultivating consumer trust and loyalty, and thus indirectly boosting sales. Advertising is an important marketing tool for advertisers.

[0003] Common advertising strategies typically utilize factors such as placement methods, exposure of placement locations, placement costs, and cost-benefit ratios to determine specific placement plans. However, these methods do not take into account the target audience, which may result in advertisements that are completely uninteresting to the recipients, leading to a waste of advertising resources. Summary of the Invention

[0004] Therefore, it is necessary to provide a user group clustering method, device, computer equipment, and storage medium that can better segment people in order to address the above-mentioned technical problems.

[0005] A user group clustering method, the method comprising:

[0006] Receive user group clustering requests and read the initial user group data carried in the user group clustering requests;

[0007] Obtain the initial feature data of the initial user group data;

[0008] The initial feature data is subjected to feature transformation to obtain the transformed feature data;

[0009] The initial user group data is clustered based on the transformed feature data to obtain clustering results.

[0010] A user group clustering method, the method comprising:

[0011] Receive user group selection instructions, acquire and display feature data of the initial user group data corresponding to the user group selection instructions on the interface;

[0012] The interface receives a selection instruction for the feature data, and the initial feature data of the initial user group data is obtained.

[0013] Upon receiving the task start request, the system performs feature transformation on the initial feature data to obtain transformed feature data. Based on the transformed feature data, the system clusters the initial user group data and displays the clustering results on the interface.

[0014] A user group clustering device, the device comprising:

[0015] The request receiving module is used to receive user group clustering requests and read the initial user group data carried in the user group clustering requests.

[0016] The feature acquisition module is used to acquire the initial feature data of the initial user group data;

[0017] The feature transformation module is used to perform feature transformation on the initial feature data to obtain transformed feature data.

[0018] The clustering module is used to cluster the initial user group data based on the transformed feature data to obtain clustering results.

[0019] In one embodiment, the above-mentioned apparatus further includes: an insight statistics module, used to perform insight statistics based on the clustering results to obtain insight statistics results corresponding to the clustering results.

[0020] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps:

[0021] Receive user group clustering requests and read the initial user group data carried in the user group clustering requests;

[0022] Obtain the initial feature data of the initial user group data;

[0023] The initial feature data is subjected to feature transformation to obtain the transformed feature data;

[0024] The initial user group data is clustered based on the transformed feature data to obtain clustering results.

[0025] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0026] Receive user group clustering requests and read the initial user group data carried in the user group clustering requests;

[0027] Obtain the initial feature data of the initial user group data;

[0028] The initial feature data is subjected to feature transformation to obtain the transformed feature data;

[0029] The initial user group data is clustered based on the transformed feature data to obtain clustering results.

[0030] The aforementioned user group clustering method, apparatus, computer equipment, and storage medium, upon receiving a user group clustering request, read the initial user group data carried in the request and obtain the initial feature data of the initial user group. Then, feature transformation is performed on the initial feature data, and the transformed feature data is used to cluster the initial user group, yielding clustering results. After obtaining the user group's feature data, the feature transformation operation retains only the features most relevant to the initial user group to be clustered. Finally, these features are used to cluster the initial user group, allowing for better user group segmentation. The resulting clusters show users in each cluster exhibiting similar user characteristics, providing guidance for ad placement and reducing the waste of advertising resources. Attached Figure Description

[0031] Figure 1 This is a diagram illustrating the application environment of a user group clustering method in one embodiment;

[0032] Figure 2 This is a flowchart illustrating the user group clustering method in another embodiment;

[0033] Figure 3 This is a flowchart illustrating a user group clustering method in a specific embodiment;

[0034] Figure 4 This is a flowchart illustrating a user group clustering method in a specific embodiment;

[0035] Figure 5(1) is a schematic diagram of the interface of the tool application entry in one embodiment;

[0036] Figure 5(2) is a schematic diagram of the interface for creating a task and selecting a task type in one embodiment;

[0037] Figure 5(3) is a schematic diagram of the interface for creating a new intelligent circle task and selecting the parent group package in one embodiment;

[0038] Figure 5(4) is a schematic diagram of the interface for setting the layer conditions in one embodiment;

[0039] Figure 5(5) is a schematic diagram of the interface for determining the task, setting parameters and submitting the task in one embodiment;

[0040] Figure 5(6) is a schematic diagram of the interface for viewing the results in one embodiment;

[0041] Figure 6 This is a structural block diagram of a user group clustering device in one embodiment;

[0042] Figure 7 Here is a block diagram of a user group clustering structure in another embodiment;

[0043] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0045] In one embodiment, such as Figure 1 As shown, a user group clustering method is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, or to a system including both a terminal and a server, and implemented through the interaction between the terminal and the server. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0046] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.

[0047] Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied to the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will all require robust system support, which can only be achieved through cloud computing.

[0048] Cloud computing is a computing model that distributes computing tasks across a large pool of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, cloud resources appear infinitely scalable, readily available, on-demand, and expandable, with payment based on usage.

[0049] As a provider of fundamental cloud computing capabilities, a cloud resource pool (referred to as a cloud platform, generally called an IaaS (Infrastructure as a Service) platform) is established. Various types of virtual resources are deployed in the resource pool for external customers to choose from. The cloud resource pool mainly includes: computing devices (virtualized machines containing operating systems), storage devices, and network devices.

[0050] In one embodiment, the user group clustering method includes the following steps S110 to S140.

[0051] Step S110: Receive user group clustering request and read the initial user group data carried in the user group clustering request.

[0052] In this embodiment, the user group clustering request is initiated by the user. In one example, an interactive interface is displayed on the front-end screen, providing the user with the option to initiate a user group clustering request. Furthermore, when initiating a user group clustering request, the user needs to provide or select the user group to be clustered. In some embodiments, if the user has specific clustering requirements, these requirements can be set on the front-end screen; for example, specifying that the user group be clustered to obtain n clusters. The user group selected by the user for clustering is the initial user group data in this embodiment.

[0053] The initial user group data includes three or more user data sets, each corresponding to one user. In one embodiment, the initial user group data includes social media accounts, mobile phone numbers, etc., of the user group; typically, one mobile phone number and one social media account correspond to only one user. Here, social media accounts are user registration accounts in various applications, such as QQ numbers, WeChat accounts, enterprise WeChat accounts, etc. When performing user group clustering, mobile phone numbers and social media accounts can be used to represent users. Understandably, in other embodiments, the user data in the initial user group data can also be other types of data.

[0054] In one embodiment, the user data in the user group data also carries a data identifier, which is used to represent the data category corresponding to the user data. For example, if the data identifier is a mobile phone number, it means that the data type of this user data is a mobile phone number; or if the data identifier is a QQ number, it means that the data type of this user data is a QQ number.

[0055] Step S120: Obtain initial feature data of the initial user group data.

[0056] Feature data refers to data associated with users, such as gender, age, occupation, or interests. In one embodiment, initial feature data includes basic attribute features: education status, geographical attributes, employment status, etc.; and interest features: business interests, information interests, video browsing preferences, etc.

[0057] In one embodiment, the initial feature data includes feature data selected by the user; in this embodiment, the user is provided with the option to select the desired features through an interactive interface, and the background receives the user-selected features as the initial feature data. In another embodiment, the initial feature data includes all features associated with the initial user data. In yet another embodiment, the initial feature data is feature data associated with the initial user group data. For example, if the type of the initial user group data is a social media account, the initial feature data includes information such as the friends of the social media account, the time spent on social media, and the time periods during which the social media account is frequently used. It is understood that in other embodiments, the initial feature data may also be feature data determined in other ways.

[0058] Further, in one embodiment, initial feature data of the initial user group is obtained from the feature repository. In one embodiment, the features stored in the feature repository include second-party features from the service provider performing user group clustering, and / or first-party features uploaded by users who have user group clustering needs; for example, in a specific embodiment, Tencent provides user group clustering services to advertisers, and Tencent's features are second-party features, while the features uploaded by advertisers are first-party features. In a specific embodiment, the feature repository is the Tencent Distribution Data Warehouse (TDW), a distributed data processing system based on a share-nothing architecture, with high availability and high scalability, used for massive data storage and analysis. It provides users with a SQL-like interface and can provide petabyte-level storage and terabyte-level computing power to meet the ever-increasing demand for massive data analysis and help discover more user value.

[0059] Step S130: Perform feature transformation on the initial feature data to obtain the transformed feature data.

[0060] Feature transformation refers to transforming initial data in a certain way to obtain the desired features. In one embodiment, feature transformation includes feature processing and filtering of the initial feature data. Feature processing includes converting the original data (which is not easy to quantify) into meaningful data (which can be quantified), or data that can be processed by a computer, so as to facilitate subsequent processing steps. Feature filtering refers to filtering out some features from the initial feature data and retaining only the required feature data.

[0061] Furthermore, in one embodiment, feature transformation includes the following specific methods: continuous variable transformation, categorical variable encoding, date variable transformation, missing value handling, and feature combination. Further, continuous variable transformation includes specific transformation methods such as continuous data standardization, continuous data transformation, or continuous data discretization. Categorical variable encoding includes categorical variable transformation and date variable transformation.

[0062] In one embodiment, feature transformation of initial feature data includes performing continuous data transformation processing on continuous features in the original feature data.

[0063] Continuous features include uncountable features such as age and height. Continuous data transformation processing of continuous features includes: continuous data standardization, continuous data transformation, and continuous data discretization. Continuous data standardization refers to transforming a continuous variable into a variable with a mean of 0 and a standard deviation of 1. Continuous data transformation refers to changing the distribution of the original data through function transformation. The purpose is to transform the data from unrelated to related, from a skewed distribution to one with greater variation after transformation, or to make the data conform to the assumptions required by the model theory, and then analyze it, for example, transforming the data to a normal distribution. In a specific embodiment, data transformation methods include: ① logarithmic function transformations such as logX, Ine, etc., x′=ln(x); ② Box-Cox transformation (a generalized power transformation method proposed by Box and Cox in 1964): a method to automatically find the optimal normal distribution transformation function. The purpose of discretizing continuous data includes facilitating the exploration of data correlations, reducing the interference of outlier data on the model, introducing nonlinearity into the model, and improving the model's predictive ability. After data discretization, features can be combined, such as changing M+N to M*N. In a specific embodiment, data discretization methods include: unsupervised discretization methods, supervised discretization methods (such as decision trees), custom rules, equal-width methods, or equal-frequency / equal-depth methods.

[0064] In another embodiment, feature transformation of the initial feature data includes encoding the categorical features in the initial feature data.

[0065] In another embodiment, feature transformation of the initial feature data includes: performing missing value replacement processing on feature data in the initial feature data that contains missing values.

[0066] In one embodiment, missing value replacement can be implemented by replacing with 0, replacing with the mean, replacing with the mode, or replacing with a prediction model. Further, in one embodiment, before performing missing value replacement, the method further includes: statistically analyzing the amount of data covered by each feature in the initial feature data, and determining whether missing value replacement is necessary based on the statistical results. In one specific embodiment, if the coverage of a feature in the initial feature data is higher than 20%, the missing value replacement step is initiated, and any of the above replacement methods (replacing with 0, replacing with the mean, replacing with the mode, or replacing with a prediction model) is selected to fill in the missing values; in another embodiment, if the coverage of a feature in the initial feature data is lower than 20%, that feature dimension is removed.

[0067] In another embodiment, feature transformation of the initial feature data includes: combining features based on the features in the initial feature data to obtain combined features.

[0068] The purpose of feature combination is to construct more and better features to improve model accuracy. In one specific embodiment, the feature combination method includes ① multiple continuous variables: addition, subtraction, multiplication, and division operations; ② multiple categorical variables: cross-combination of all values. In one embodiment, facing a large number of high-dimensional sparse features, there are a large number of cross-combination methods. Manually designing feature crosses not only requires a lot of manpower and trial costs, but also easily overlooks some important cross features. Therefore, a feature cross-processing module is introduced to perform feature cross-processing. In one specific embodiment, for the explicit construction of high-order cross features, a DCN-style cross-processing method can be used to implement the feature cross-processing process.

[0069] In another embodiment, feature transformation also includes feature extraction. Since clustering models cannot handle sequential features well, if sequential data needs to be introduced, manual feature extraction is required to create discrete or continuous features. Manual feature extraction often has many limitations; different sequential data often contain different informational patterns, and manual extraction methods struggle to truly capture these semantic connotations. In one embodiment, a network module for feature extraction from sequential data is introduced, enabling the model to automatically mine the information contained within the sequential data. In a specific embodiment, the feature extraction process is implemented using a Transformer, which can capture the relationships between people in the sequential data.

[0070] Step S140: Cluster the initial user group data based on the transformed feature data to obtain the clustering results.

[0071] The process of dividing a collection of physical or abstract objects into multiple classes composed of similar objects is called clustering. A cluster generated by clustering is a set of data objects that are similar to objects within the same cluster and different from objects in other clusters. In this embodiment, the initial user group is clustered based on transformed feature data; that is, based on user characteristics, a clustering algorithm is used to divide the target user set into different user clusters.

[0072] In one embodiment, clustering can be implemented using K-Means clustering, mean-shift clustering algorithm, DBSCAN clustering algorithm, expectation-maximization (EM) clustering using Gaussian mixture model (GMM), and hierarchical clustering methods.

[0073] In one specific embodiment, the K-Means clustering method is used to cluster the initial user group data based on the transformed feature data. K-Means clustering, or k-means clustering algorithm, is an iterative clustering analysis algorithm. Its steps are as follows: the data is pre-divided into K groups; K objects are randomly selected as initial cluster centers; the distance between each object and each seed cluster center is calculated; and each object is assigned to the nearest cluster center. The cluster centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the cluster centers are recalculated based on the existing objects in the cluster. This process is repeated until a termination condition is met. The termination condition may be that no (or a minimum number) objects are reassigned to different clusters, no (or a minimum number) cluster centers change, or the sum of squared errors reaches a local minimum.

[0074] Furthermore, in one specific embodiment, unsupervised machine learning algorithms such as k-means multi-view clustering (clustering data using multiple different descriptive methods) and spectral clustering are used to cluster the initial user group.

[0075] Clustering user groups essentially involves dividing the overall population into different subgroups, characterizing the features of each subgroup, and then, based on these features, establishing corresponding marketing strategy audience segments.

[0076] In one embodiment, after obtaining the clustering results of the user group, the method further includes: displaying the clustering results. In this embodiment, after the clustering is completed and the clustering results are obtained, the clustering results are transmitted to the front end for rendering and displayed in the interactive interface for users to view.

[0077] In another embodiment, after obtaining the clustering results of the user group, the method further includes storing the clustering results to a preset storage path. In this embodiment, the clustering results of the user group can be stored according to the preset storage path to facilitate subsequent user queries of the clustering results. Subsequently, upon receiving a clustering result query initiated by a user, the corresponding clustering results can be read from the preset storage path for display.

[0078] The aforementioned user group clustering method, upon receiving a user group clustering request, reads the initial user group data carried in the request and obtains the initial feature data of the initial user group. Then, it performs feature transformation on the initial feature data and uses the transformed feature data to cluster the initial user group, obtaining the clustering results. After obtaining the user group's feature data, the feature transformation operation retains only the features more relevant to the initial user group to be clustered. Finally, using these features to cluster the initial user group allows for better user group segmentation. The resulting clusters show users with similar user characteristics, providing guidance for ad placement and reducing the waste of advertising resources.

[0079] In one embodiment, after clustering the initial user group data based on the transformed feature data to obtain the clustering results, the method further includes: performing insight statistics based on the clustering results to obtain the insight statistics results corresponding to the clustering results.

[0080] Insight statistics: This involves performing statistical calculations on different user clusters, such as means, histograms, and concentration comparisons (group mean / sample mean), to provide a basis for drawing conclusions about user insights. Customer insight: This is the process of discovering user attributes and characteristics by observing the feature distribution of a target user group.

[0081] After clustering the user group, the user group is further clustered based on the transformed adjustment, that is, the user group is divided into different user clusters. In this embodiment, each user cluster is analyzed, which may include analyzing the feature coverage in each user cluster and generating statistical analysis results based on the feature coverage. The statistical analysis results are displayed in statistical forms such as mean, histogram, concentration comparison, etc., which can provide users with easy-to-understand data analysis conclusions for subsequent operations.

[0082] In one embodiment, the insight statistics include different functions such as industry insights, my audience, content insights, campaign analysis, and tool applications. Users can select the type of insight analysis they need through the interactive interface. In this embodiment, after obtaining the clustering results, the clustering results are fed back to the front end for display. Based on the user's selection, the clustering results are then subjected to insight statistical analysis to obtain the insight statistical analysis results, which are then transmitted to the front end for display.

[0083] In one embodiment, such as Figure 2 As shown, before clustering the initial user group data based on the transformed feature data to obtain the clustering results, S210 is also included: performing feature filtering on the transformed feature data to obtain the target feature data.

[0084] Feature filtering involves processing the transformed feature data to select some features, which are denoted as target feature data in this embodiment. Finally, the user group data is clustered using the target feature data.

[0085] Furthermore, in this embodiment, clustering the initial user group data based on the transformed feature data to obtain clustering results includes: clustering the initial user group data based on the target feature data to obtain clustering results.

[0086] In one embodiment, feature filtering is performed on the transformed feature data to obtain target feature data, including: calculating the variance of each feature dimension in the target feature data and the correlation coefficient between each pair of features; and based on the variance of all features and the correlation coefficient between each pair of features, feature filtering is performed on all features in the transformed feature data to obtain target feature data.

[0087] In statistical description, variance is used to calculate the difference between each variable (observation) and the population mean. To avoid the sum of squared deviations from the mean being zero, and considering that the sum of squared deviations from the mean is affected by sample size, statistics uses the average sum of squared deviations from the mean to describe the degree of variability of the variable. In a specific example, the population variance is calculated as follows:

[0088]

[0089] Where, σ 2 Let X be the population variance, σ be the variable, σ be the population mean, and N be the population size.

[0090] Correlation is a non-deterministic relationship, and the correlation coefficient is a measure of the degree of linear correlation between variables. Due to the different research objects, the correlation coefficient can be defined in several ways. Simple correlation coefficient: also called the correlation coefficient or linear correlation coefficient, generally represented by the letter r, is used to measure the linear relationship between two variables. In a specific embodiment, the calculation of the correlation coefficient includes:

[0091]

[0092] Where Cov(X,Y) is the covariance of X and Y, and σ X Let σ be the standard deviation of X. Y Let X be the standard deviation of Y. The correlation coefficient is calculated by dividing the covariance of X and Y by the standard deviations of X and Y, respectively.

[0093] In one embodiment, the correlation coefficient is also the special covariance after standardization, which eliminates the influence of the two variables' dimensions.

[0094] The calculation of covariance includes:

[0095] Cov(X,Y)=E[(X-μ X )(Y-μ Y )]

[0096] The covariance is calculated as follows: if there are two variables, the difference between the X value and its mean at each time step is multiplied by the difference between the Y value and its mean.

[0097] The calculation of standard deviation includes:

[0098]

[0099] The standard deviation is the square root of the arithmetic mean of the squared deviations from the mean, denoted by σ. It is also the square root of the variance of the standard deviation. The standard deviation reflects the dispersion of a dataset.

[0100] Further, in one specific embodiment, based on the variance of all features and the correlation coefficient between every two-dimensional features, feature filtering is performed on all features in the transformed feature data to obtain target feature data. This includes: retaining features in two-dimensional features with a correlation coefficient greater than a preset correlation coefficient threshold whose corresponding variance is greater than a preset variance threshold, and removing features whose corresponding variance is less than or equal to a preset variance threshold. In another embodiment, all two-dimensional features with a correlation coefficient greater than the preset correlation coefficient threshold can also be sorted according to the magnitude of their corresponding variances, and the features with the highest preset values ​​can be retained. In yet another embodiment, all two-dimensional features with a correlation coefficient less than the preset correlation coefficient threshold are removed from the transformed feature data.

[0101] In this embodiment, by calculating the correlation coefficient between features and combining it with the variance of the feature itself, it is determined whether the feature needs to be removed. This allows for filtering of the transformed feature data, thereby retaining features that are more relevant to the user group. These features can better divide the initial user group into multiple user clusters, resulting in better clustering results.

[0102] In one embodiment, please refer to... Figure 2 Before clustering the initial user group data based on the transformed feature data to obtain the clustering results, S220 is also included: converting the target feature data into a format and outputting the transformed feature data after format conversion.

[0103] In one specific embodiment, the feature data obtained after feature transformation and / or feature filtering of the initial feature data is in libsvm format (a data format). When clustering the initial user group, the feature data format needs to be converted to dense format (a data format). Subsequently, the clustering results are obtained by using the dense format feature data to cluster the initial user group data.

[0104] This application also provides an application scenario in which the above-mentioned user group clustering method is applied, such as... Figure 3 The diagram shown illustrates the user group clustering method in this embodiment. This embodiment uses user group data as a number packet as an example, including mobile phone numbers, QQ numbers, WeChat accounts, etc. Specifically, the application of this user group clustering method in this application scenario is as follows:

[0105] 1. The MI (Marketing Strategy Platform, a strategy service for user growth experts) front-end presentation layer interacts with users, providing functions such as allowing users to select number packages (the aforementioned initial user group data) and task parameters.

[0106] 2. The MI front-end service layer uses the user's selection results to generate a user group clustering task request and sends the request to the back-end system.

[0107] 3. The backend receives user group clustering task requests, records the number packets, number types, feature lists, parameters, etc., and starts subsequent tasks.

[0108] 4. Extract the original features from the feature repository and denote them as the initial feature data. Perform feature transformation on the initial feature data and output the transformed feature data.

[0109] 5. Calculate the correlation coefficient and variance of features in the feature warehouse and output the results.

[0110] 6. EMP (Elastic Modeling Platform, a collection of custom modeling tools) reads the transformed feature data, performs feature filtering based on correlation coefficients and variance, and obtains the target feature data. The target feature data output by EMP is in libsvm format.

[0111] 7. EMP converts libsvm format data to dense format.

[0112] 8. EMP calls the clustering algorithm and outputs the clustering results (user's cluster, cluster center, cluster feature name).

[0113] 9. The feature repository performs insight statistics and outputs the statistical results for each user cluster in the clustering results.

[0114] 10. The MI front-end service layer obtains clustering and insight results and converts them into the format required for front-end display.

[0115] 11. The MI front-end is used for display.

[0116] The attribution of each of the above steps is as follows: Figure 4 As shown in the figure. Among them, AMS represents the user segment.

[0117] The following is arranged according to different stages

[0118] 1. Trigger the task (MI → Task Manager)

[0119] • Input:

[0120] Seed number package

[0121] Feature list

[0122] Task configuration parameters

[0123] • Returns: Status or error message

[0124] 2. Feature Extraction and Processing (Task Manager → Feature Repository)

[0125] • Input:

[0126] wuid list

[0127] Feature list

[0128] Output:

[0129] Target feature data

[0130] Statistical information of features

[0131] 3. Trigger clustering calculation (Task Manager → Control Center)

[0132] • Input:

[0133] Model configuration information

[0134] Statistical information of features

[0135] Target feature data

[0136] Output: Clustering results

[0137] 4. Trigger task insight statistics (Task Manager → Feature Repository)

[0138] • Input:

[0139] Clustering results

[0140] Feature lists are used to gain insights into feature data.

[0141] Output: Insights into statistical results

[0142] The user group clustering method in the above embodiments displays an interface on the front end where users can select number packages and clustering task parameters. Users can start the user group clustering task on the interface, select the number packages of the user groups to be clustered, and set the clustering task parameters, such as clustering the user group into k user clusters. Users can also input the features needed for clustering. After receiving the user group clustering request, the backend obtains the initial feature data, performs feature transformation and feature filtering on the initial feature data to obtain target feature data, and performs format conversion. Based on the converted target feature data, the number packages are clustered to obtain the clustering results. Finally, insight statistical analysis is performed on each user cluster in the clustering results, the insight results are output, and the results are fed back to the front end for display. This method clusters user groups using user-selected features or features obtained from a feature repository. After feature transformation and filtering, the number packets are clustered. Because the feature transformation and filtering processes perform certain processing based on feature coverage, features with low coverage are removed, retaining only those that better distinguish number packets. This makes the target feature data more closely match the number packets to be clustered, resulting in better clustering results. Finally, insight analysis is performed on each user cluster in the clustering results, and the statistical results are fed back to users for viewing, providing them with more intuitive clustering and analysis results.

[0143] Furthermore, the aforementioned user group clustering method can be applied to advertising. Before advertising, the user group clustering method can be used to cluster and analyze the target audience of the advertisement. Then, the advertising plan can be designed more accurately based on the clustering results, reducing the waste of advertising resources.

[0144] Figure 5(1) shows a schematic diagram of the interface of the tool application entry in one embodiment. The interface includes a main menu area, a task folder selection area, a navigation area, and a tool application task details & operation area. The main menu area is divided into industry insights, my audience, content insights, campaign analysis, and tool applications. The intelligent circle tasks corresponding to the user group clustering method are under the tool application module.

[0145] Figure 5(2) shows a schematic diagram of the interface for creating a task and selecting a task type in one embodiment. Selecting "Intelligent Circle Task" under the "Tools Application" module will take you to the Intelligent Circle Task unit.

[0146] Figure 5(3) shows a schematic diagram of the interface for creating a new smart circle task and selecting the parent group package in one embodiment. The interface includes a main menu area, a task folder management area, a navigation area, and a main work area. The task folder management area reuses the current MI file management mechanism and can be hidden by clicking the arrow shown in the interface diagram. The right side of the navigation area is used to explain the current user location (tool application) and return to the main page; the left side is used to explain the steps for creating a new smart circle task, and Figure 5(3) shows step 1. The search box in the main work area is used to provide users with template search based on task name; the group package data range in the list of the main task area is used to obtain the group package information of the task status under the current user's permission range. Group packages with a group size ≥ 5 million are set as optional, and the rest are set as unselectable; the definition domains of group name, creation time, number type, group type, task status, and group size in the list should be consistent with the original definition of the current MI product. The next step in the main task area is limited to selecting a group package. The next step is available. Then, click the next step to enter the circle page. For the paginated display of the list in the main task area, the maximum number of records per page is 10.

[0147] Figure 5(4) shows a schematic diagram of the interface for setting circle conditions in one embodiment. The interface includes a main menu area, a parent group package recording area, a circle condition setting recording area, a navigation area, and a feature selection area. The right side of the navigation area is used to explain the current user location (tool application) and return to the main page. The left side of the navigation area is used to explain the working steps for creating a new smart circle task. Figure 5(4) shows step 2. The parent group package recording area includes: case 1. Single group, just record the group name. After the user clicks the group name and closes the icon, they will return directly to the previous step; case 2. Multiple groups, record multiple group names, and by default take the union of multiple group packages, and record the union rules; the user can delete a specified group package by using the close button corresponding to the group package name. After deleting the last one, they will return to the previous step. The search box in the feature selection area can provide users with fuzzy search based on the tag (feature name) name; the feature selection area is limited to a maximum of 10 features. After exceeding 10, the remaining features cannot be selected; when returning to the previous step in the feature selection area, the historical filtering record is retained; the next step in the feature selection area leads to the task confirmation interface. The circle condition setting record area allows you to set the number of sub-packages. The default selection is the suggested number, and users can select the expected number of sub-packages via a drop-down box (single selection only). The feature records in the circle condition setting record area are used to record the results of the features selected by the user, displayed in a tree structure.

[0148] Figure 5(5) shows a schematic diagram of the interface for determining tasks, setting parameters, and submitting tasks in one embodiment. The interface includes a main menu area, a navigation area, and a task information details area. The right side of the navigation area is used to indicate the current user location (Smart Circle < Tool Application) and to return to the main page; the left side of the navigation area is used to explain the steps for creating a new Smart Circle task, as shown in Figure 5(5) for step 3. The task name in the task information details area can be edited by the user; the folder selection can be used to display the folders created by the user when clicked, and the hierarchical relationship of the currently displayed folders is Smart Circle - Task Name - Task Package; the parent group package is not editable and is used to record the parent group package logic, the classification features are not editable and are used to record the selected feature information; the number of sub-packages is editable; the reminder method is a prompt channel after the task is completed, which can be implemented by calling existing functions. Clicking "Start Task" in the interface shown in Figure 5(5) displays the task execution prompt.

[0149] Figure 5(6) shows a schematic diagram of the interface for viewing results in one embodiment. The interface includes a main menu area, a navigation area, a task information details area, and a task result main interface. The right side of the navigation area is used to indicate the current user location (task name - result < smart circle < tool application) and to return to the main page. The task information details area includes: parent group package record: case 1. single group package, retain the original group package name; case 2. multiple group package, copy the task name to the newly generated parent group package; parent group size is the current group package size; sub-package features are the selected feature details; number of sub-packages is the number of sub-package tasks selected by the task.

[0150] Furthermore, the list in the main task results interface includes: a. The default rule for sub-group names is Task Name + Core Feature (Main Differential Feature) + Number (Item No. 01), which is editable. Users can set the text box to edit mode using the edit button and save after editing. b. Group Description: Describes the main feature values ​​of the current group. c. Group Size: The actual size of the current sub-group. d. Status: An enumerated value. If the status is "Not Extracted," the group is not visible on the front end; if the status is "Extracted," the group is visible on the front end. e. Type: The original group is the group directly generated after task execution; the combined group is the group adjusted by the user. f. Update Time: The last operation time of the current group. g. The checkbox next to the group name allows you to select all current sub-groups. The "Merge Groups" button in the main task results interface: This button becomes available after selecting >1 sub-group; clicking it leads to the merged interface. The "Extract Selected Groups" button in the task results interface: This button becomes available after selecting >1 sub-group; clicking it leads to the prompt interface. The "Exit" button on the main interface of the task results: Click it to return to the tool application page.

[0151] In one specific embodiment, a user group clustering method is provided. The method includes: receiving a user group selection instruction, acquiring and displaying feature data of the initial user group data corresponding to the user group selection instruction on the interface; receiving a selection instruction for the feature data on the interface to obtain the initial feature data of the initial user group data; receiving a task start request, performing feature transformation on the initial feature data to obtain transformed feature data, clustering the initial user group data based on the transformed feature data, and displaying the obtained clustering results on the interface.

[0152] The user issues a user group selection command in the interface, corresponding to the step of selecting the initial user group (parent group) as shown in Figure 5(3). After selecting the initial user group (parent group), the feature data is displayed in the interface, corresponding to the step of setting up the circle as shown in Figure 5(4). The user can select the required feature data as the initial feature data in the interface. The system receives the user's selection command for the feature data (i.e., the user clicks "Next") and generates the initial user feature data based on the selected features. The system proceeds to the task confirmation step shown in Figure 5(5). After the user clicks "Start Task", the system receives the task start request, performs feature transformation based on the initial feature data, obtains the transformed feature data, and then performs user clustering on the initial user group (parent group) based on the transformed feature data to obtain the clustering results. Finally, the clustering results are displayed in the interface, corresponding to the interface shown in Figure 5(6).

[0153] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0154] In one embodiment, such as Figure 6 As shown, a user group clustering device is provided. This device can be a software module, a hardware module, or a combination of both integrated into a computer device. Specifically, the device includes: a request receiving module 610, a feature acquisition module 620, a feature transformation module 630, and a clustering module 640, wherein:

[0155] The request receiving module 610 is used to receive user group clustering requests and read the initial user group data carried in the user group clustering requests.

[0156] The feature acquisition module 620 is used to acquire initial feature data of the initial user group data;

[0157] The feature transformation module 630 is used to perform feature transformation on the initial feature data to obtain the transformed feature data.

[0158] Clustering module 640 is used to cluster the initial user group data based on the transformed feature data to obtain the clustering results.

[0159] The aforementioned user group clustering device, upon receiving a user group clustering request, reads the initial user group data carried in the request and obtains the initial feature data of the initial user group. Then, it performs feature transformation on the initial feature data and uses the transformed feature data to cluster the initial user group, obtaining clustering results. After obtaining the user group's feature data, the feature transformation operation retains only the features more relevant to the initial user group to be clustered. Finally, these features are used to cluster the initial user group, which can better classify the user group. The resulting clusters show users in each cluster having similar user characteristics, providing guidance for ad placement and reducing the waste of advertising resources.

[0160] In one embodiment, such as Figure 7 As shown, the above device also includes: an insight statistics module 710, used to perform insight statistics based on the clustering results, and obtain the insight statistics results corresponding to the clustering results.

[0161] In one embodiment, please refer to... Figure 7 The above-mentioned device further includes: a feature filtering module 720, used to perform feature filtering on the transformed feature data to obtain target feature data; in this embodiment, the clustering module 640 is specifically used to cluster the initial user group data based on the target feature data to obtain clustering results.

[0162] In one embodiment, the feature filtering module of the above-mentioned device includes: a calculation unit for calculating the variance of each feature dimension in the target feature data and the correlation coefficient between every two features; and a filtering unit for performing feature filtering on all features in the transformed feature data based on the variance of all features and the correlation coefficient between every two features to obtain the target feature data.

[0163] In one embodiment, the feature transformation module 630 of the above-described device is specifically used to perform continuous data transformation processing on continuous features in the initial feature data.

[0164] In another embodiment, the feature transformation module 630 of the above-described device is specifically used to encode the categorical features in the initial feature data.

[0165] In another embodiment, the feature transformation module 630 of the above-described device is specifically used to perform missing value replacement processing on feature data in which there are missing values ​​in the initial feature data.

[0166] In another embodiment, the feature transformation module 630 of the above-described device is specifically used to combine features based on the features in the initial feature data to obtain the combined features.

[0167] In one embodiment, the above apparatus further includes: a format conversion module, used to convert the target feature data into a format and output the converted feature data.

[0168] Specific limitations regarding the user group clustering device can be found in the limitations of the user group clustering method described above, and will not be repeated here. Each module in the aforementioned user group clustering device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0169] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a user group clustering method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0170] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0171] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0172] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0173] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.

[0174] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0175] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0176] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A user population clustering method, characterized by, The method includes: The marketing strategy platform front end provides users with the option to select number packages and task parameters. It uses the user's selection results to generate a user group clustering request and sends the request to the back end system. The back end receives the user group clustering request, reads the initial user group data carried in the user group clustering request, and starts subsequent tasks. The feature repository extracts the original features and provides them to users to select the desired features on the interactive interface. The backend receives the features selected by the user as the initial feature data, obtains the initial feature data of the initial user group data, and obtains the initial feature data of the initial user group from the feature repository. The features stored in the feature repository include second-party features from the service provider that performs user group clustering and first-party features uploaded by users who have user group clustering needs. The feature warehouse performs feature transformation on the initial feature data to obtain transformed feature data, calculates the feature correlation coefficient and variance, outputs the results, and calculates the variance of each feature dimension in the transformed feature data, as well as the correlation coefficient between each pair of features. The elastic modeling platform reads the transformed feature data, performs feature filtering based on correlation coefficient and variance, retains features with a corresponding variance greater than a preset variance threshold among two-dimensional features with a correlation coefficient greater than a preset correlation coefficient threshold, and removes features with a corresponding variance less than or equal to a preset variance threshold. For two-dimensional features with a correlation coefficient less than the preset correlation coefficient threshold, they are all removed from the transformed feature data to obtain target feature data. The platform then calls a clustering algorithm to cluster the initial user group data based on the target feature data to obtain clustering results. The feature repository performs insight statistics. Users can select the type of insight analysis they need through the interactive interface. Based on the user's selection, the clustering results are then subjected to insight statistical analysis to obtain the insight statistical analysis results. The insight statistics include at least three types of insight analysis with different functions such as industry insights, my audience, content insights, campaign analysis, and tool applications. The marketing strategy platform obtains clustering results and insightful statistical analysis results from the front end, and converts them into the format required for front-end display. Before advertising, the target audience for the ads is clustered and analyzed, and an advertising campaign is designed based on the clustering results.

2. The user population clustering method of claim 1, wherein, The feature data obtained after feature transformation of the initial feature data is in libsvm format. When clustering the initial user group, the feature data format is converted to dense format.

3. The user population clustering method of claim 1, wherein, The intelligent circle task corresponding to the user group clustering method is located in the tool application module.

4. The user population clustering method of claim 1, wherein, The method further includes: Based on the variance of all features and the correlation coefficient between each pair of features, feature filtering is performed on all features in the transformed feature data to obtain the target feature data.

5. The user population clustering method of claim 1, wherein, The feature repository performs feature transformation on the initial feature data to obtain transformed feature data, including at least one of the following: The first step is to perform continuous data transformation processing on the continuous features in the initial feature data; The second step is to encode the categorical features in the initial feature data; The third step is to perform missing value replacement processing on the feature data with missing values ​​in the initial feature data. The fourth step involves combining the features in the initial feature data to obtain the combined features.

6. The user group clustering method according to any one of claims 3 to 5, characterized in that, Before clustering the initial user group data based on the target feature data to obtain the clustering results, the process further includes: The target feature data is converted into a new format, and the converted target feature data is output.

7. A user group clustering device, characterized in that, The device includes: The request receiving module is used to provide users with selected number packages and task parameters through the marketing strategy platform front end, generate user group clustering requests using the user selection results, send the requests to the backend system, receive the user group clustering requests, read the initial user group data carried in the user group clustering requests, and start subsequent tasks. The feature acquisition module is used to extract raw features from the feature repository, provide the user with the features to choose from on the interactive interface, receive the features selected by the user as initial feature data in the background, obtain the initial feature data of the initial user group data, and obtain the initial feature data of the initial user group from the feature repository. The features stored in the feature repository include second-party features from the service provider that performs user group clustering and first-party features uploaded by users who have user group clustering needs. The feature transformation module is used to transform the initial feature data through the feature warehouse to obtain transformed feature data, calculate the feature correlation coefficient and variance, output the results, and calculate the variance of each feature dimension in the transformed feature data, as well as the correlation coefficient between each pair of features. The clustering module is used to read transformed feature data through the elastic modeling platform, filter features based on correlation coefficient and variance, retain features with a correlation coefficient greater than a preset correlation coefficient threshold and a corresponding variance greater than a preset variance threshold, and remove features with a corresponding variance less than or equal to a preset variance threshold. For two-dimensional features with a correlation coefficient less than the preset correlation coefficient threshold, they are removed from the transformed feature data to obtain target feature data. The clustering algorithm is then called to cluster the initial user group data based on the target feature data to obtain clustering results. Insight statistics are performed through a feature warehouse. Users can select the type of insight analysis they need on the interactive interface. Based on the user's selection, the clustering results are then subjected to insight statistical analysis to obtain insight statistical analysis results. Insight statistics include at least three types of insight analysis with different functions such as industry insights, my audience, content insights, placement analysis, and tool applications. The clustering results and insight statistical analysis results are obtained through the marketing strategy platform front-end and converted into the format required for front-end display. Before advertising, the target audience for the advertisements to be placed is first clustered and analyzed, and an advertising placement plan is designed based on the clustering results.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Group information classification method and device, computer device and storage medium

    CN109858525A

  • User classification method and device, computer equipment and storage medium

    CN111553390A

  • Shop site selection method based on decision tree

    CN111985576A