User value prediction method, system and device and storage medium
By preprocessing user tags and selecting the optimal feature values using the ant colony algorithm, a generative adversarial network model was constructed, which solved the problems of irrelevant and redundant features in the user tag system, achieved accurate prediction of user value and improved enterprise operational efficiency.
Patent Information
- Application Number
- CN202510804794.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-19
AI Technical Summary
The existing user labeling system contains a large number of irrelevant and redundant features, which affect the training effect of the value estimation model and make it difficult to accurately predict user value.
By obtaining user labels, user categories and user label values for preprocessing, the ant colony algorithm is used to select the best feature values, and a generative adversarial network model is constructed for training. The generative adversarial network model competes in adversarial mode to improve prediction accuracy.
It improves data utilization and accuracy, removes irrelevant and redundant features, enhances the training effect of the value estimation model, achieves accurate user value prediction, and enhances the company's data insight and operational efficiency.
Smart Images

Figure CN120670913A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a user value prediction method, system, device and storage medium. Background Art
[0002] Each communication user has numerous tags, and over time, the number has accumulated to nearly 1,000. These tags include various profiles, including app preferences, terminal preferences, consumption preferences (e.g., video enthusiasts), behavior preferences, location preferences, and preferences for data and voice usage. Tags often contain irrelevant and redundant features, which hinder the training of value estimation models. Uncovering the potential value of users based on this vast tag system will benefit market development and empower frontline personnel. Summary of the Invention
[0003] In view of this, the purpose of the embodiments of the present invention is to provide a user value prediction method, system, device and storage medium, which can select labels with optimal features and accurately predict user value through a generative adversarial network model.
[0004] In one aspect, an embodiment of the present invention provides a method for predicting user value, comprising:
[0005] Obtaining a user tag, a user category, and a user tag value, and preprocessing the user tag value to obtain a user tag feature value; the user tag includes several types of sold items;
[0006] Based on the ant colony algorithm, a preset number of optimal feature values are selected from the user tag feature values to form a sample data set;
[0007] Training a preset model according to the sample data set to obtain a trained generative adversarial network model;
[0008] Obtaining items for sale, and predicting user value based on the items for sale and the generative adversarial network model.
[0009] Optionally, the preprocessing of the user tag value includes:
[0010] Eliminate abnormal values in the user tag values;
[0011] and / or, filling in the missing values in the user tag values;
[0012] And / or, the user tag value is standardized.
[0013] Optionally, based on an ant colony algorithm, a preset number of optimal feature values are selected from the user tag feature values to form a sample data set, including:
[0014] Calculating the correlation between the user tags and between the user tags and the user categories;
[0015] constructing an initial ant colony search space according to the user tag, the user category, and the correlation;
[0016] Perform several iterations according to the initial ant colony search space and search parameters, select a preset number of optimal feature values from the user label feature values, and form a sample data set according to the user labels, the optimal feature values and the user categories.
[0017] Optionally, calculating the correlations between the user tags and between the user tags and the user categories includes:
[0018] Calculating the distance correlation coefficient between the user tags to determine the correlation between the user tags;
[0019] A distance correlation coefficient between the user tag and the user category is calculated to determine the correlation between the user tag and the user category.
[0020] Optionally, the preset model includes a first model, and training the preset model according to the sample data set includes:
[0021] Classifying the sample data set to obtain a negative sample data set;
[0022] The negative sample data set is divided into a first training set and a first test set, the first model is trained according to the first training set, and the first model is tested according to the first test set until requirements are met.
[0023] Optionally, the preset model includes a second model, and the training of the preset model according to the sample data set includes:
[0024] Classifying the sample data set to obtain a positive sample data set;
[0025] The positive sample data set is divided into a second training set and a second test set, the second model is trained according to the second training set, and the second model is tested according to the second test set until the requirements are met.
[0026] Optionally, obtaining the product to be sold and predicting the user value based on the product to be sold and the generative adversarial network model includes:
[0027] Obtain products to be sold, and replace the products to be sold with the sold products in the sample data set to form a prediction data set;
[0028] The prediction data set is input into the generative adversarial network model to predict user value.
[0029] On the other hand, an embodiment of the present invention provides a user value prediction system, including:
[0030] The first module is used to obtain user tags, user categories and user tag values, and pre-process the user tag values to obtain user tag feature values; the user tags include several types of sold products;
[0031] The second module is configured to select a preset number of optimal feature values from the user tag feature values based on an ant colony algorithm to form a sample data set;
[0032] The third module is used to train the preset model according to the sample data set to obtain a trained generative adversarial network model;
[0033] The fourth module is used to obtain products to be sold and predict user value based on the products to be sold and the generative adversarial network model.
[0034] On the other hand, an embodiment of the present invention provides a user value prediction device, comprising:
[0035] at least one processor;
[0036] at least one memory for storing at least one program;
[0037] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0038] On the other hand, an embodiment of the present invention provides a computer-readable storage medium storing a program executable by a processor. When the program is executed by the processor, it is used to perform the above method.
[0039] The implementation of the embodiment of the present invention includes the following beneficial effects: this embodiment first pre-processes the user label value to obtain the user label feature value, thereby improving data utilization and accuracy; based on the ant colony algorithm, a preset number of optimal feature values are selected from the user label feature values to form a sample data set, remove irrelevant features and redundant features, and improve the training effect of the value estimation model; the preset model is trained according to the sample data set to obtain a trained generative adversarial network model; products to be sold are obtained, and user value is predicted based on the products to be sold and the generative adversarial network model, thereby selecting a label with the best features, and accurately predicting user value through the generative adversarial network model, further improving user experience and enterprise operating efficiency, bringing deeper data insights to the enterprise, and supporting more scientific and accurate decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1This is a schematic diagram of the steps of a user value prediction method provided by an embodiment of the present invention;
[0041] Figure 2 This is a flowchart of steps for preprocessing user tag values provided by an embodiment of the present invention;
[0042] Figure 3 This is a schematic flow chart of steps for forming a sample data set provided by an embodiment of the present invention;
[0043] Figure 4 This is a schematic diagram of a step flow for calculating correlation provided by an embodiment of the present invention;
[0044] Figure 5 This is a schematic flow chart of steps for training a first model provided by an embodiment of the present invention;
[0045] Figure 6 1 is a flowchart of steps for training a second model provided by an embodiment of the present invention;
[0046] Figure 7 This is a schematic diagram of a process flow for predicting user value provided by an embodiment of the present invention;
[0047] Figure 8 This is a structural block diagram of a user value prediction system provided by an embodiment of the present invention;
[0048] Figure 9 This is a structural block diagram of a user value prediction device provided by an embodiment of the present invention;
[0049] Figure 10 This is a structural block diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are provided for ease of description only and do not limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted based on the understanding of those skilled in the art.
[0051] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flow chart. The terms "first", "second", etc. in the specification and claims and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0053] Some technical terms in this embodiment are explained below.
[0054] Ant Colony Algorithm: This algorithm mimics the behavior of ants in their search for food to solve optimization problems. It transforms feature selection into a minimum-cost path problem in a search graph. Through multiple iterations, it searches the feature space to find user tags with minimal redundancy and maximum relevance.
[0055] Generative Adversarial Networks (GANs): The basic principle of GANs is to train a generator against a discriminator through continuous competition, ultimately reaching an equilibrium point where the discriminator cannot distinguish between the generator's generated data and real data. In this process, the generator is responsible for capturing the distribution of sample data and generating new data.
[0056] like Figure 1 As shown, an embodiment of the present invention provides a user value prediction method, including:
[0057] S100: Obtain user tags, user categories, and user tag values, and pre-process the user tag values to obtain user tag feature values; user tags include several types of sold items.
[0058] It should be noted that user tags are determined based on actual applications and are not specifically limited in this embodiment. For example, user tags include, but are not limited to, the number of primary cards, the number of secondary cards, the number of broadband users, the number of fixed-line users, and voice plans. User categories are determined based on actual applications and are not specifically limited in this embodiment. For example, user categories include, but are not limited to, business dedicated line customers, live broadband customers, and enterprise customers. The user tag value refers to the specific numerical value corresponding to the user tag.
[0059] Specifically, user tags, user categories, and user tag values are obtained, and unreasonable data such as abnormal data and blank data in the user tag values are preprocessed to obtain user tag feature values.
[0060] S200 , based on the ant colony algorithm, selecting a preset number of optimal feature values from the user tag feature values to form a sample data set.
[0061] The optimal feature value refers to the most valuable feature value among the user tag feature values, that is, the feature value that has the greatest impact on subsequent user value prediction. The preset number is determined based on actual application and is not specifically limited in this embodiment. For example, it is determined based on the number of user tag feature values and prediction accuracy.
[0062] Specifically, based on the ant colony algorithm, a preset number of optimal feature values are selected from the user label feature values, and then a sample data set is formed according to the optimal feature values, the user labels corresponding to the optimal feature values, and the user categories.
[0063] S300: Train the preset model according to the sample data set to obtain a trained generative adversarial network model.
[0064] The preset model refers to a generative adversarial network model with undetermined parameters. The model consists of a generative model G and a discriminative model D. These models compete in an adversarial mode: the discriminative model D attempts to distinguish between real positive samples and those forged by the generative model G, while the generative model G attempts to deceive the discriminative model D. The sample dataset is divided into a training set and a test set. The preset model is trained with the training set, its parameters are updated, and the trained model is tested on the test set until it meets the preset requirements, resulting in a trained generative adversarial network model.
[0065] S400: Obtain products to be sold, and predict user value based on the products to be sold and a generative adversarial network model.
[0066] Items for sale have similar user labels as those for sold items, but their potential user values are different. Specifically, we replace sold items with items for sale to form a prediction dataset. We then use this dataset and a generative adversarial network model to predict user values.
[0067] Using a cascaded feature-based generative adversarial network, we can accurately predict the potential value change of users after ordering different items. In this embodiment of the invention, if the predicted result shows no increase in user value, it is considered a decrease in value; if the predicted result is positive, it is considered an increase in user value.
[0068] Alternatively, as Figure 2 As shown, the user tag value is preprocessed, including:
[0069] S110, removing abnormal values from user tag values;
[0070] S120, and / or, filling in the missing values in the user tag values;
[0071] S130, and / or, performing standardization processing on the user tag value.
[0072] Specifically, the user tag value is preprocessed, including but not limited to the following operations:
[0073] 1) Eliminate outliers in user tag values; 2) Use zero-filling to fill in missing values in user tag values; 3) Standardize user tag values. The specific calculation is as follows:
[0074]
[0075] Here, x represents a value in the user tag value, min represents the minimum value in the user tag value, and max represents the maximum value in the user tag value.
[0076] It should be noted that the preprocessing of the user tag value includes but is not limited to the above-mentioned processing method, and the preprocessing may include one or more of the above-mentioned processing methods.
[0077] Alternatively, as Figure 3 As shown, based on the ant colony algorithm, a preset number of optimal feature values are selected from the user tag feature values to form a sample data set, including:
[0078] S210: Calculate the correlation between user tags and between user tags and user categories.
[0079] Specifically, first, the correlation between user tags and user categories is calculated based on the values of the user tags and the values of the user categories. Correlation indicates the degree of association between the two. The stronger the correlation, the greater the degree of association. The method for calculating correlation depends on the actual application and is not specifically limited in this embodiment.
[0080] S220: Construct an initial ant colony search space according to user tags, user categories, and relevance.
[0081] The ant colony search space is a directed weighted graph containing multiple nodes. Multiple nodes are formed according to user tags and user categories. The weights between nodes are determined according to the correlation, thereby constructing the initial ant colony search space.
[0082] S230 , performing several iterations based on the initial ant colony search space and search parameters, selecting a preset number of optimal feature values from the user label feature values, and forming a sample data set based on the user labels, optimal feature values, and user categories.
[0083] The search parameters and number of iterations are determined based on actual applications and are not specifically limited in this embodiment. Specifically, after performing several iterations based on the initial ant colony search space and search parameters, a preset number of user tags are selected based on the order of node pheromone values. A preset number of optimal feature values are then selected from the user tag feature values based on the user tags. Finally, a sample dataset is formed based on the selected user tags, optimal feature values, and user categories.
[0084] In a specific embodiment, let the search space be a fully connected directed weighted graph, G=<F,E> , where F={F1,F2,…,F n ,F m} are user labels and user categories; E={(F i ,F j ):F i ,F j ∈F} is the edge of the graph, edge (F i ,F j )∈E is the weight setting, specifically the inverse of the distance correlation coefficient between node i and node j. The calculation formula is as follows:
[0085]
[0086] Initially, l ants are randomly placed in n nodes. The probability of ant k going from node i to node j is calculated as:
[0087]
[0088] Where, τ i is the initial pheromone of node i, is the set of nodes that ant k has not visited when it is located at node i, α is the information heuristic factor, β is the expectation heuristic factor, and they are the main parameters controlling pheromone and correlation; η(i,j) is the path length from node i to node j, which is the inverse of the distance correlation coefficient.
[0089] After each iteration, the node pheromone is updated. The formula is as follows:
[0090]
[0091] Where, τ i (t) is the pheromone of node i at time t, ρ is the pheromone evaporation coefficient, and FC(i) is a counter that records the number of times the ant encounters feature i in one iteration.
[0092] After N iterations, the features are sorted according to the pheromone value, and the user label nodes are selected to form the user label subset.
[0093] Alternatively, as Figure 4 As shown, the correlation between user tags and between user tags and user categories is calculated, including:
[0094] S211. Calculate the distance correlation coefficient between user tags to determine the correlation between user tags;
[0095] S212: Calculate the distance correlation coefficient between the user tag and the user category to determine the correlation between the user tag and the user category.
[0096] The distance correlation coefficient represents the linear or nonlinear relationship between the two. The distance correlation coefficient is calculated according to a preset formula. The preset formula is determined according to actual application and is not specifically limited in this embodiment.
[0097] In a specific embodiment, the Pearson correlation coefficient is generally used to calculate the correlation between vectors, but the Pearson correlation coefficient can only measure whether there is a linear relationship between the feature and the category label, and cannot reflect the nonlinear relationship. In this embodiment of the present invention, the distance correlation coefficient is selected to calculate the correlation between them. The distance correlation coefficient makes up for the shortcomings of the Pearson correlation coefficient. If the distance correlation coefficient value of two vectors is 0, it is considered that the two are independent, and the larger the value, the stronger the correlation between the two. The distance correlation coefficient estimated value calculation formula is as follows:
[0098]
[0099] in, Among them, S1, S2, and S3 are:
[0100]
[0101] Where x∈R p , y∈R q , x i is the i-th value of vector x, y i is the i-th value of vector y, and n is the length of the vector.
[0102] By substituting the values of user tags and user categories into the above formula, the correlation between user tags and between user tags and user categories can be calculated.
[0103] Alternatively, as Figure 5 As shown, the preset model includes a first model, and training the preset model according to the sample data set includes:
[0104] S301, classify the sample data set to obtain a negative sample data set;
[0105] S302: Divide the negative sample data set into a first training set and a first test set, train the first model according to the first training set, and test the first model according to the first test set until the requirements are met.
[0106] Specifically, the method for classifying the sample dataset is determined based on actual application and is not specifically limited in this embodiment. The first model refers to the model before the G model is trained. The ratio of the first training set to the first test set is determined based on actual application and is not specifically limited in this embodiment. During the training of the first model, the parameters of the first model are updated based on sample similarity.
[0107] Alternatively, as Figure 6 As shown, the preset model includes a second model, and the preset model is trained according to the sample data set, including:
[0108] S303, classify the sample data set to obtain a positive sample data set;
[0109] S304: Divide the positive sample data set into a second training set and a second test set, train the second model according to the second training set, and test the second model according to the second test set until the requirements are met.
[0110] Specifically, the classification method for the sample dataset is determined based on actual application and is not specifically limited in this embodiment. The second model refers to the model before the D model is trained. The ratio of the second training set to the second test set is determined based on actual application and is not specifically limited in this embodiment. During the training of the second model, the parameters of the second model are updated based on the accuracy rate.
[0111] In a specific embodiment, to more accurately characterize user preferences using the features extracted from the above process and thereby improve the overall prediction performance, the present invention uses a triplet loss algorithm for feature classification: samples are divided into three categories: base samples, negative samples, and positive samples. Base samples are high-quality user features used as a reference; positive samples accurately characterize user behavior and provide a foundation for improving prediction performance; and negative samples are used as a reverse incentive to better establish positive samples.
[0112] All high-quality users proposed can be used as base samples A. In order to optimize its classification or generation capabilities, during the training process, the embodiment of the present invention generates positive samples P by continuously "approaching" the distribution of similar samples, and generates negative samples N by "moving away" from the distribution of heterogeneous samples.
[0113] In the user feature value prediction matrix, the item feature vector is represented as f(x)∈R d , each sample i can be mapped to the d-dimensional Euclidean space, and the triplet loss algorithm makes the user-based sample feature vector f(x a ) and the positive sample feature vector f(x p ) is closer to the negative sample feature vector f(x n ) is further. Therefore, its optimization goal is as follows:
[0114]
[0115] Among them, α is a constant that represents the boundary value of positive and negative sample pair training, and H is the set of all samples related to the user.
[0116] User value prediction is key to achieving high-quality user development, and user value calculation is primarily based on the interaction between users and products. To obtain richer and more accurate user value information, it is necessary to deeply analyze the "user-product" matrix using a generative adversarial network.
[0117] The generative adversarial network consists of a generative model G and a discriminative model D, which compete in an adversarial mode: D tries to distinguish between real positive samples and G's forged positive samples, while G tries to deceive D. The cost function of GAN is shown in the formula.
[0118]
[0119] Among them, m is the sample in the training set, U represents the user set, u n represents the nth user, φ and θ represent the parameters of the discriminant model D and the generative model G, respectively.
[0120] First, train the discriminant model D, that is, maximize the probability of correctly distinguishing true positive samples from G's forged positive samples, as shown in the formula:
[0121]
[0122] Among them, m + , m - Represent positive samples and negative samples respectively, m g Represents the positive samples forged by the generative model G.
[0123] In contrast to the discriminative model D, the generative model G aims to minimize the correct discrimination probability of the discriminative model D, that is, to deceive the discriminative model. G extracts negative samples from the sample pool according to the relevant probability distribution, and the parameter optimization of its generative model is shown in the following formula.
[0124]
[0125] Among them, g φ (u,m) calculates the sample similarity, and t is the time parameter.
[0126]
[0127] Among them, M is the sample set, and mg represents the positive samples forged by the generative model G.
[0128] Because the sample data of individual users is discrete, it is necessary to use the policy gradient method based on reinforcement learning to predict user value, and then obtain the user value classification of user-associated sales products, as shown in the following formula, where K is the size of the sample pair list that needs to be judged.
[0129]
[0130] Alternatively, as Figure 7 As shown, obtain the product to be sold, and predict the user value based on the product to be sold and the generative adversarial network model, including:
[0131] S401. Obtain products to be sold, and replace the products to be sold with the sold products in the sample data set to form a prediction data set;
[0132] S402: Input the predicted data set into the generative adversarial network model to predict user value.
[0133] Items for sale refer to items that may be recommended to users. Specifically, the items for sale are replaced by sold items in the sample dataset, while other data remains unchanged, forming a predicted dataset. This dataset is then fed into the generative adversarial network model to predict user value. User value predictions include no increase in user value and an increase in user value. A predicted result of 0 indicates no increase in user value, while a predicted result of 1 indicates an increase in user value.
[0134] The implementation of the embodiment of the present invention includes the following beneficial effects: this embodiment first pre-processes the user label value to obtain the user label feature value, thereby improving data utilization and accuracy; based on the ant colony algorithm, a preset number of optimal feature values are selected from the user label feature values to form a sample data set, remove irrelevant features and redundant features, and improve the training effect of the value estimation model; the preset model is trained according to the sample data set to obtain a trained generative adversarial network model; products to be sold are obtained, and user value is predicted based on the products to be sold and the generative adversarial network model, thereby selecting a label with the best features, and accurately predicting user value through the generative adversarial network model, further improving user experience and enterprise operating efficiency, bringing deeper data insights to the enterprise, and supporting more scientific and accurate decision-making.
[0135] The user value prediction process is described below with a specific embodiment.
[0136] S1. 5,000 tagged users are selected as samples for this embodiment. Each user contains 1,735 tags, and each user tag corresponds to a user tag value. The user tag values are preprocessed. The user tag feature values after preprocessing are as follows:
[0137] Feature 1 Feature 2 Feature 3 …… Feature 1735 User 1 0.913 0.236 -0.159 …… 0.591 User 2 0.516 -0.746 0.352 …… 0.483 User 3 -0.433 -0.519 -0.258 …… 0.649 …… …… …… …… …… …… 5000 users 0.262 -0.176 0.306 …… -0.742
[0138] S2. Calculate the distance correlation coefficients between user tags and between user tags and user categories. The correlation values obtained by calculation are as follows:
[0139] Feature 1 Feature 2 …… Feature 1735 User Category Feature 1 1 0.589 …… 0.026 0.724 Feature 2 0.589 1 …… 0.472 0.891 …… ……5 …… …… …… …… Feature 1735 0.026 0.472 …… 1 0.722 User Category 0.724 0.891 …… 0.722 1
[0140] S3. Ant colony algorithm feature selection, the parameters are set as follows: number of iterations N = 50, number of ants l = 1736, information heuristic factor α = 1, expected heuristic factor β = 1, initial pheromone τ0 = 0.2, pheromone evaporation coefficient ρ = 0.1. After the parameter setting is completed, each node is placed with an ant, and then the iteration work is started. The details of the pheromone value of each iteration node are shown in the following table:
[0141]
[0142]
[0143] Through the ant colony algorithm, the 200 features with the highest pheromone values are selected as the optimal feature values to complete the feature selection work.
[0144] S4. Form a training data set based on the selected best features, and divide the training data set into a training set and a test set in a 4:1 ratio.
[0145] S5. G model generates negative samples and updates model parameters. Assume that the number of iterations g = 1000. First, calculate the similarity between training samples. The details are as follows:
[0146] Sample 1 Sample 2 Sample 3 …… Sample 4000 Sample 1 1 0.537 0.172 …… 0.451 Sample 2 0.537 1 0.476 …… 0.318 Sample 3 0.150 0.476 1 …… 0.221 …… …… …… …… …… …… Sample 4000 0.451 0.318 0.221 …… 1
[0147] The following probability formula is used to generate negative sample data. After generating negative samples, the probability is recalculated and the G model parameters are adjusted to generate negative samples again. After repeating g times, the G model training and negative sample data are completed.
[0148]
[0149] The D model judges the model comparison accuracy and model parameter update, sets the number of iterations d = 1000, and compares the results of the generated model G with the correct probability distribution. The judgment results are as follows:
[0150] Sample 1 Sample 2 Sample 3 …… Sample 1000 1 0 1 1
[0151] Compare the predicted result distribution probability with the generated G model distribution probability, and update the D model parameters according to the predicted distribution result distribution probability. After repeating the iteration d times, the D model training can be completed.
[0152] S6. Calculate whether the trained model has converged. If so, the model training is complete. Otherwise, execute steps S4 and S5 until the model converges.
[0153] S7. User value prediction: We used the trained generative adversarial network model to select 8 sales items for a 50,000 test set to predict their value. We verified the accuracy of the model by calculating the actual user value (0 represents a decrease, 1 represents an increase). Some of the results are shown in the following table:
[0154]
[0155] In the above embodiment, there are more than 1,700 user tags in total, and there are many irrelevant and redundant features in the tags, which affect the training effect of the value estimation model. Selecting the optimal features through the feature selection method can effectively improve the accuracy and operation efficiency of the value estimation model. The traditional feature selection method uses the similarity between features and categories when selecting features. The embodiment of the present invention is based on the feature selection method of the improved ant colony algorithm, which not only considers the similarity between features and categories, but also adds the similarity calculation between features. The algorithm of the embodiment of the present invention can remove irrelevant and redundant features in user features, thereby improving the model operation results.
[0156] Using a cascaded feature-based generative adversarial network, we can accurately predict the potential value change of users after ordering different items. In this embodiment of the invention, if the predicted result shows no increase in user value, it is considered a decrease in value; if the predicted result is positive, it is considered an increase in user value.
[0157] In digital marketing, we provide algorithms for calculating user value through in-depth mining, enabling accurate assessment of user needs and value calculation. This will provide access to deeper, valuable data and promote the development of more scientific and precise marketing plans and policies. The advantages are as follows:
[0158] 1. Accurately identify user needs: By analyzing users' specific behaviors, preferences, or attributes, we can enhance user satisfaction and gain a more precise understanding of their needs and interests. This helps provide personalized product recommendations, content, or services, thereby increasing user satisfaction and loyalty.
[0159] 2. Improve marketing efficiency: After identifying the characteristics that have the greatest impact on user value, the marketing team can design targeted marketing strategies and activities, reducing resources wasted on non-target or low-value users. This not only improves conversion rates but also optimizes marketing budget allocation.
[0160] 3. Predicting User Behavior: By analyzing the relationship between user characteristics and behavior in historical data, we can build a predictive model to estimate future user behavior, such as demand intentions and churn risk. This provides a basis for taking preventive measures (such as customer retention strategies).
[0161] 4. Optimize products and services: Understanding which user characteristics best drive user value growth can help companies continuously optimize their products and services.
[0162] 5. Enhanced user segmentation capabilities: Analysis based on optimal features enables more detailed segmentation of user groups, enabling ultra-segmented market targeting. This segmentation is not only based on traditional user voice and data usage behavior, but also incorporates factors such as app usage preferences, gigabit availability, contract status, and broadband availability, enabling each segment to receive more tailored services or products.
[0163] 6. Promote dynamic value management: User value changes over time, context, and personal circumstances. By continuously monitoring and analyzing optimal characteristics, companies can adjust strategies to adapt to dynamic changes in user value and maintain long-term positive relationships with users.
[0164] like Figure 8 As shown, an embodiment of the present invention provides a user value prediction system, including:
[0165] The first module is used to obtain user tags, user categories and user tag values, and pre-process the user tag values to obtain user tag feature values; user tags include several types of sold products;
[0166] The second module is used to select a preset number of optimal feature values from the user tag feature values based on the ant colony algorithm to form a sample data set;
[0167] The third module is used to train the preset model according to the sample data set to obtain a trained generative adversarial network model;
[0168] The fourth module is used to obtain products to be sold and predict user value based on the products to be sold and the generative adversarial network model.
[0169] It can be seen that the contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0170] like Figure 9 As shown, an embodiment of the present invention provides a user value prediction device, comprising:
[0171] at least one processor;
[0172] at least one memory for storing at least one program;
[0173] When at least one program is executed by at least one processor, the at least one processor implements the above method.
[0174] Among them, the memory is a non-transient computer-readable storage medium that can be used to store non-transient software programs and non-transient computer executable programs. The memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory optionally includes a remote memory remotely arranged relative to the processor, and these remote memories can be connected to the processor via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0175] It can be seen that the contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0176] In addition, the embodiments of the present application further disclose a computer program product or computer program, which is stored in a computer-readable storage medium. The processor of a computer device can read the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device performs the above-mentioned method. Similarly, the contents of the above-mentioned method embodiment are all applicable to the present storage medium embodiment, and the functions specifically implemented by the present storage medium embodiment are the same as those of the above-mentioned method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned method embodiment.
[0177] An embodiment of the present invention further provides a computer-readable storage medium, which stores a program executable by a processor. The program executable by the processor is used to implement the above method when executed by the processor.
[0178] It is understood that all or some steps, systems in the disclosed method above can be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components can be implemented as software by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those of ordinary skill in the art, the term computer storage medium is included in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data) and is volatile and non-volatile, removable and non-removable media. Computer storage media includes but is not limited to RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, disk storage or other magnetic storage device, or can be used to store desired information and any other medium that can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0179] Specifically, see Figure 10The computer device 1000 may include an RF (Radio Frequency) circuit 1010, a memory 1020 including one or more computer-readable storage media, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a short-range wireless transmission module 1070, a processor 1080 including one or more processing cores, and a power supply 1090. It will be understood by those skilled in the art that Figure 10 The device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0180] The RF circuit 1010 can be used to receive and transmit signals during information transmission or calls. Specifically, it receives downlink information from the base station and transmits it to one or more processors 1080 for processing. Furthermore, it transmits uplink data to the base station. Typically, the RF circuit 1010 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a subscriber identity module (SIM) card, a transceiver, a coupler, an LNA (low noise amplifier), a duplexer, and the like. Furthermore, the RF circuit 1010 can communicate with the network and other devices via wireless communication. Wireless communication can utilize any communication standard or protocol, including but not limited to GSM (Global System of Mobile Communications), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, and SMS (Short Messaging Service).
[0181] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various functional applications and data processing by running the software programs and modules stored in the memory 1020. The memory 1020 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the device 1000 (such as audio data, a phone book, etc.), etc. In addition, the memory 1020 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 1020 may also include a memory controller to provide the processor 1080 and the input unit 1030 with access to the memory 1020. Although Figure 10 The RF circuit 1010 is shown, but it is understandable that it is not an essential component of the device 1000 and can be omitted as needed without changing the essence of the invention.
[0182] The input unit 1030 can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical, or trackball signal input related to user settings and function control. Specifically, the input unit 1030 may include a touch-sensitive surface 1031 and other input devices 1032. The touch-sensitive surface 1031, also known as a touch display or touchpad, can detect user touch operations on or near it (for example, operations performed by a user using a finger, stylus, or any other suitable object or accessory on or near the touch-sensitive surface 1031) and drive corresponding connected devices according to a pre-set program. Optionally, the touch-sensitive surface 1031 may include a touch detection device and a touch controller. The touch detection device detects the user's touch direction and detects the signals generated by the touch operation, transmitting the signals to the touch controller. The touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 1080. It can also receive and execute commands from the processor 1080. In addition, the touch-sensitive surface 1031 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 1031, the input unit 1030 can also include other input devices 1032. Specifically, the other input devices 1032 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, a joystick, and the like.
[0183] The display unit 1040 can be used to display information input by the user or information provided to the user and various graphical user interfaces of the control 1000, which can be composed of graphics, text, icons, videos and any combination thereof. The display unit 1040 may include a display panel 1041. Optionally, the display panel 1041 may be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like. Furthermore, the touch-sensitive surface 1031 may be covered on the display panel 1041. When the touch-sensitive surface 1031 detects a touch operation on or near it, it is transmitted to the processor 1080 to determine the type of touch event. The processor 1080 then provides corresponding visual output on the display panel 1041 according to the type of touch event. Although in Figure 10 In the embodiment, the touch-sensitive surface 1031 and the display panel 1041 are implemented as two independent components to implement input and output functions, but in some embodiments, the touch-sensitive surface 1031 and the display panel 1041 can be integrated to implement input and output functions.
[0184] The computer device 1000 may also include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor may turn off the display panel 1041 and / or the backlight when the device 1000 is moved to the ear. As a type of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the posture of the mobile phone (such as switching between horizontal and vertical screens, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can be configured in the device 1000, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be described in detail here.
[0185] Audio circuit 1060, speaker 1061, and microphone 1062 provide an audio interface between the user and device 1000. Audio circuit 1060 converts received audio data into electrical signals and transmits them to speaker 1061, which then converts them into sound signals for output. Microphone 1062, on the other hand, converts collected sound signals into electrical signals, which are then received by audio circuit 1060 and converted into audio data. The audio data is then processed by output processor 1080 and sent to another control device via RF circuit 1010, or the audio data is output to memory 1020 for further processing. Audio circuit 1060 may also include an earphone jack to allow external headphones to communicate with device 1000.
[0186] The short-range wireless transmission module 1070 may be a WIFI (wireless fidelity) module, a Bluetooth module, an infrared module, etc. The device 1000 may transmit information to a wireless transmission module provided on a competing device via the short-range wireless transmission module 1070 .
[0187] Processor 1080 is the control center of device 1000. It connects the various components of the entire control device using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 1020 and accessing data stored in memory 1020, it performs various functions of device 1000 and processes data, thereby providing overall control of the control device. Optionally, processor 1080 may include one or more processing cores. Alternatively, processor 1080 may integrate an application processor and a modem processor, with the application processor primarily handling the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 1050.
[0188] Device 1000 also includes a power supply 1090 (e.g., a battery) for supplying power to various components. Preferably, the power supply can be logically connected to processor 1080 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. Power supply 1090 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0189] Although not shown, the device 1000 may also include a camera, a Bluetooth module, etc., which will not be described in detail here.
[0190] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0191] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0192] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0193] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A user value prediction method, characterized in that: include: Obtaining user tags, user categories, and user tag values, and preprocessing the user tag values to obtain user tag feature values; The user tag includes several types of sold items; Based on the ant colony algorithm, a preset number of optimal feature values are selected from the user tag feature values to form a sample data set; Training a preset model according to the sample data set to obtain a trained generative adversarial network model; Obtaining items for sale, and predicting user value based on the items for sale and the generative adversarial network model.
2. The method according to claim 1, characterized in that The preprocessing of the user tag value includes: Eliminate abnormal values in the user tag values; and / or, filling in the missing values in the user tag values; And / or, the user tag value is standardized.
3. The method according to claim 1, characterized in that Based on the ant colony algorithm, a preset number of optimal feature values are selected from the user tag feature values to form a sample data set, including: Calculating the correlation between the user tags and between the user tags and the user categories; constructing an initial ant colony search space according to the user tag, the user category, and the correlation; Perform several iterations according to the initial ant colony search space and search parameters, select a preset number of optimal feature values from the user label feature values, and form a sample data set according to the user labels, the optimal feature values and the user categories.
4. The method according to claim 3, characterized in that The calculating the correlation between the user tags and between the user tags and the user categories includes: Calculating the distance correlation coefficient between the user tags to determine the correlation between the user tags; A distance correlation coefficient between the user tag and the user category is calculated to determine the correlation between the user tag and the user category.
5. The method according to claim 1, wherein The preset model includes a first model, and the training of the preset model according to the sample data set includes: Classifying the sample data set to obtain a negative sample data set; The negative sample data set is divided into a first training set and a first test set, the first model is trained according to the first training set, and the first model is tested according to the first test set until requirements are met.
6. The method according to claim 1, characterized in that The preset model includes a second model, and the training of the preset model according to the sample data set includes: Classifying the sample data set to obtain a positive sample data set; The positive sample data set is divided into a second training set and a second test set, the second model is trained according to the second training set, and the second model is tested according to the second test set until the requirements are met.
7. The method according to claim 1, characterized in that The obtaining of the product to be sold and predicting the user value based on the product to be sold and the generative adversarial network model includes: Obtain products to be sold, and replace the products to be sold with the sold products in the sample data set to form a prediction data set; The prediction data set is input into the generative adversarial network model to predict user value.
8. A user value prediction system, characterized in that: include: The first module is used to obtain user tags, user categories and user tag values, and preprocess the user tag values to obtain user tag feature values; The user tag includes several sold items; The second module is configured to select a preset number of optimal feature values from the user tag feature values based on an ant colony algorithm to form a sample data set; The third module is used to train the preset model according to the sample data set to obtain a trained generative adversarial network model; The fourth module is used to obtain products to be sold and predict user value based on the products to be sold and the generative adversarial network model.
9. A user value prediction device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 7 when executed by the processor.