User classification method and device, electronic equipment and storage medium

By determining the correlation coefficient of the purchased traffic packages of the car network entertainment traffic users in the user classification model, the problem of inaccurate classification in the existing technology is solved, and accurate identification and classification of user types are achieved.

CN120598591APending Publication Date: 2025-09-05CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510701151.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing user classification methods are not suitable for user classification scenarios of in-vehicle network entertainment traffic and cannot effectively identify the traffic usage needs of different users.

Method used

By obtaining user traffic usage data from the training data set, the user classification model is used to determine the correlation coefficient between the consumed traffic, arrival frequency, arrival interval time and consumption amount of the purchased traffic package, and user classification is performed. The user classification model is trained based on the correlation coefficient to determine the user type of the current user.

Benefits of technology

It achieves accurate classification of users of car network entertainment traffic, reduces the complexity of the user classification model, and improves the accuracy and efficiency of classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598591A_ABST
    Figure CN120598591A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a user classification method and device, electronic equipment and a storage medium. The method comprises the steps that a training data set is acquired; inputting the training data set into a user classification model, and determining the consumption flow of a purchased flow packet in the training data set, the account-arriving frequency of the purchased flow packet, the account-arriving interval time of the purchased flow packet, and correlation coefficients between the consumption flow, the account-arriving frequency and the account-arriving interval time of the purchased flow packet and the consumption amount of the purchased flow packet through the user classification model; classifying the plurality of users based on the correlation coefficient to train a user classification model; acquiring traffic use data of the current user, and determining a user type of the current user according to the traffic use data of the current user and the trained user classification model; wherein different user types correspond to different traffic use requirements. Therefore, the trained user classification model can accurately classify the user types of the users of the in-vehicle network entertainment traffic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a user classification method, a user classification device, an electronic device and a computer-readable storage medium. Background Art

[0002] With the development of smart cars, users are using more and more in-vehicle network entertainment traffic. Therefore, effectively classifying users of in-vehicle network entertainment traffic can help service providers of entertainment traffic packages to gain a deeper understanding of users' purchasing preferences, usage preferences and consumption capabilities, thereby promoting service providers of entertainment traffic packages to formulate refined operation strategies and reduce operating costs. In related technologies, user types are classified based on the characteristics of traditional e-commerce industry that only one product is purchased at a time, but the package packages of in-vehicle network entertainment traffic have the characteristic of automatically superimposing traffic packages every month, which is very different from the characteristics of traditional e-commerce. The method of classifying users in related technologies is not applicable to the scenario of classifying users of in-vehicle network entertainment traffic. Summary of the Invention

[0003] In view of the above problems, embodiments of the present invention are proposed to provide a user classification method, a user classification device, an electronic device, and a computer-readable storage medium that overcome the above problems or at least partially solve the above problems.

[0004] In order to solve the above problems, an embodiment of the present invention discloses a user classification method, including:

[0005] Obtaining a training data set, the training data set including traffic usage data of multiple users; the traffic usage data including the amount of money spent on a traffic package, the traffic consumed by the traffic package, the frequency of receipt of the traffic package, and the interval between receipts of the traffic package;

[0006] Inputting the training data set into a user classification model, determining, through the user classification model, the correlation coefficients of the traffic consumption of the purchased traffic package, the frequency of arrival of the purchased traffic package, and the interval between arrival of the purchased traffic package, respectively, and the consumption amount of the purchased traffic package, and classifying the multiple users based on the correlation coefficients to train the user classification model;

[0007] Obtain the traffic usage data of the current user, and determine the user type of the current user based on the traffic usage data of the current user and the trained user classification model; wherein different user types correspond to different traffic usage requirements.

[0008] Optionally, inputting the training data set into a user classification model includes:

[0009] Determining the traffic usage data of a plurality of users in the training data set whose consumption amount of the purchased traffic packages is greater than zero as a first sub-training data set;

[0010] Normalizing the first sub-training data set to obtain a normalized first sub-training data set;

[0011] The normalized first sub-training data set is input into a user classification model.

[0012] Optionally, the correlation coefficients of the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package, respectively, and the consumption amount of the purchased traffic package, determined by the user classification model, include:

[0013] Clustering the plurality of users according to the traffic usage data of the plurality of users in the first sub-training data set using the user classification model to obtain a plurality of clusters; each cluster includes at least one user;

[0014] The user classification model is used to determine, for each cluster, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package of the users in each cluster, and their correlation coefficients with the consumption amount of the purchased traffic package.

[0015] Optionally, the user classification model is used to determine, for each cluster, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package of each user in the cluster, and the correlation coefficients thereof with the consumption amount of the purchased traffic package, respectively, including:

[0016] Determining, for each cluster, using the user classification model, a first average value of the amount of consumption of the purchased traffic packages, a second average value of the traffic consumed by the purchased traffic packages, a third average value of the frequency of arrival of the purchased traffic packages, and a fourth average value of the interval between arrivals of the purchased traffic packages for users in each cluster;

[0017] The user classification model is used to determine, based on the first average value, the second average value, the third average value, the fourth average value, the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package, a first correlation coefficient between the consumed traffic of the purchased traffic package and the consumption amount of the purchased traffic package, a second correlation coefficient between the arrival frequency of the purchased traffic package and the consumption amount of the purchased traffic package, and a third correlation coefficient between the arrival interval of the purchased traffic package and the consumption amount of the purchased traffic package for each user in the cluster.

[0018] Optionally, the clusters include a first cluster, a second cluster, and a third cluster; and classifying the multiple users based on the correlation coefficients includes:

[0019] Determining, by the user classification model, based on the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient, respectively, a first weight of the consumption amount of the purchased traffic package, a second weight of the consumed traffic of the purchased traffic package, a third weight of the arrival frequency of the purchased traffic package, and a fourth weight of the arrival interval of the purchased traffic package for each user in the cluster;

[0020] determining a comprehensive score for each cluster using the user classification model according to the first weight, the second weight, the third weight, the fourth weight, the first average value, the second average value, the third average value, and the fourth average value;

[0021] Through the user classification model, the users in the first cluster are determined as the first user type, the users in the second cluster are determined as the second user type, and the users in the third cluster are determined as the third user type; the comprehensive score of the first cluster is greater than the comprehensive score of the second cluster and the comprehensive score of the third cluster; the comprehensive score of the second cluster is greater than the comprehensive score of the third cluster; the traffic usage demand corresponding to the first user type is greater than the traffic usage demand corresponding to the second user type; the traffic usage demand corresponding to the second user type is greater than the traffic usage demand corresponding to the third user.

[0022] Optionally, the user classification model determines, according to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient, a first weight of the consumption amount of the purchased traffic package, a second weight of the consumed traffic of the purchased traffic package, a third weight of the arrival frequency of the purchased traffic package, and a fourth weight of the arrival interval of the purchased traffic package for each user in the cluster, including:

[0023] determining a total coefficient according to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient by using the user classification model;

[0024] The user classification model is used to determine the first weight of the consumption amount of the purchased traffic package of the user in each cluster, the second weight of the consumed traffic of the purchased traffic package, the third weight of the arrival frequency of the purchased traffic package, and the fourth weight of the arrival interval of the purchased traffic package.

[0025] Optionally, after determining the traffic usage data of a plurality of users in the training data set whose consumption amount for purchasing the traffic packages is greater than zero as the first sub-training data set, the method further includes:

[0026] Determine the traffic usage data of multiple users in the training data set whose traffic purchase amount is zero as the second sub-training data set,

[0027] The second sub-training data set is input into the user classification model, and based on the second sub-training data set and preset classification conditions, the multiple users in the second sub-training data set are classified to train the user classification model.

[0028] Optionally, the traffic usage data further includes: usage of the free traffic package, validity period of the free traffic package, and a record of purchased traffic packages; the preset classification conditions include: a first preset condition and a second preset condition; the first preset condition is that the usage of the user's free traffic package is greater than a usage threshold; the second preset condition is that the validity period of the user's free traffic package is expired and there is a record of purchased traffic packages; the classification of the multiple users in the second sub-training data set based on the second sub-training data set and the preset classification conditions includes:

[0029] Determining at least one user in the second sub-training data set that meets a first preset condition as a third user type by using the user classification model;

[0030] Determining, by the user classification model, at least one user in the second sub-training data set that meets a second preset condition as a fourth user type; the traffic usage demand corresponding to the third user type is greater than the traffic usage demand corresponding to the fourth user type;

[0031] The user classification model is used to determine at least one user in the second sub-training data set who does not meet the first preset condition and does not meet the second preset condition as the fifth user type; the traffic usage demand corresponding to the fourth user type is greater than the traffic usage demand corresponding to the fifth user type.

[0032] An embodiment of the present invention further discloses a device for determining a user type, comprising:

[0033] an acquisition module, configured to acquire a training data set, the training data set comprising traffic usage data of a plurality of users; the traffic usage data comprising a consumption amount of a traffic package purchased, traffic consumed by the traffic package purchased, a frequency of receipt of the traffic package purchased, and a time interval between receipts of the traffic package purchased;

[0034] a training module, configured to input the training data set into a user classification model, determine, through the user classification model, a correlation coefficient between the traffic consumption of the purchased traffic package, the frequency of arrival of the purchased traffic package, and the interval between arrivals of the purchased traffic package, and the consumption amount of the purchased traffic package, respectively, and classify the multiple users based on the correlation coefficients to train the user classification model;

[0035] A determination module is used to obtain the traffic usage data of the current user, and determine the user type of the current user based on the traffic usage data of the current user and the trained user classification model; wherein different user types correspond to different traffic usage requirements.

[0036] Optionally, the training module includes:

[0037] A first determining submodule is configured to determine the traffic usage data of a plurality of users in the training data set whose consumption amount of the purchased traffic packages is greater than zero as a first sub-training data set;

[0038] a normalization module, configured to perform normalization processing on the first sub-training data set to obtain a normalized first sub-training data set;

[0039] The first input module is configured to input the normalized first sub-training data set into a user classification model.

[0040] Optionally, the training module includes:

[0041] a clustering submodule, configured to cluster the plurality of users according to the traffic usage data of the plurality of users in the first sub-training data set using the user classification model to obtain a plurality of clusters; each cluster includes at least one user;

[0042] The second determination submodule is used to determine, for each cluster, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package of the users in each cluster through the user classification model, and the correlation coefficients thereof with the consumption amount of the purchased traffic package.

[0043] Optionally, the second determining submodule includes:

[0044] a first determining unit, configured to determine, for each cluster, by using the user classification model, a first average consumption amount of the purchased traffic packages, a second average consumption amount of the purchased traffic packages, a third average frequency of arrival of the purchased traffic packages, and a fourth average interval between arrivals of the purchased traffic packages for users in each cluster;

[0045] The second determination unit is used to determine, through the user classification model and based on the first average value, the second average value, the third average value, the fourth average value, the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package of the users in each cluster, a first correlation coefficient between the consumed traffic of the purchased traffic package and the consumption amount of the purchased traffic package, a second correlation coefficient between the arrival frequency of the purchased traffic package and the consumption amount of the purchased traffic package, and a third correlation coefficient between the arrival interval of the purchased traffic package and the consumption amount of the purchased traffic package.

[0046] Optionally, the clusters include a first cluster, a second cluster, and a third cluster; and the training module includes:

[0047] a third determination submodule, configured to determine, by using the user classification model and according to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient, a first weight of the consumption amount of the purchased traffic package, a second weight of the consumed traffic of the purchased traffic package, a third weight of the arrival frequency of the purchased traffic package, and a fourth weight of the arrival interval of the purchased traffic package for each user in the cluster;

[0048] a fourth determination submodule, configured to determine a comprehensive score for each cluster according to the first weight, the second weight, the third weight, the fourth weight, the first average value, the second average value, the third average value, and the fourth average value through the user classification model;

[0049] The fifth determination submodule is used to determine the users in the first cluster as the first user type, the users in the second cluster as the second user type, and the users in the third cluster as the third user type through the user classification model; the comprehensive score of the first cluster is greater than the comprehensive score of the second cluster and the comprehensive score of the third cluster; the comprehensive score of the second cluster is greater than the comprehensive score of the third cluster; the traffic usage demand corresponding to the first user type is greater than the traffic usage demand corresponding to the second user type; the traffic usage demand corresponding to the second user type is greater than the traffic usage demand corresponding to the third user.

[0050] Optionally, the third determining submodule includes:

[0051] a third determining unit, configured to determine a total coefficient according to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient by using the user classification model;

[0052] The fourth determination unit is used to determine the first weight of the consumption amount of the purchased traffic package of each user in the cluster, the second weight of the consumed traffic of the purchased traffic package, the third weight of the arrival frequency of the purchased traffic package, and the fourth weight of the arrival interval time of the purchased traffic package through the user classification model according to the first correlation coefficient, the second correlation coefficient, the third correlation coefficient and the total coefficient.

[0053] Optionally, the training module further includes:

[0054] The sixth determining submodule is configured to determine the traffic usage data of a plurality of users whose traffic purchase amount is zero in the training data set as a second sub-training data set,

[0055] The second input submodule is used to input the second sub-training data set into the user classification model, and classify the multiple users in the second sub-training data set based on the second sub-training data set and preset classification conditions to train the user classification model.

[0056] Optionally, the traffic usage data further includes: usage of the free traffic package, validity period of the free traffic package, and a record of purchased traffic packages; the preset classification conditions include: a first preset condition and a second preset condition; the first preset condition is that the usage of the user's free traffic package is greater than a usage threshold; the second preset condition is that the validity period of the user's free traffic package is expired and there is a record of purchased traffic packages; the second input submodule includes:

[0057] a fifth determining unit, configured to determine, by using the user classification model, at least one user in the second sub-training data set that meets a first preset condition as a third user type;

[0058] a sixth determining unit, configured to determine, by using the user classification model, at least one user in the second sub-training data set that meets a second preset condition as a fourth user type; and wherein the traffic usage demand corresponding to the third user type is greater than the traffic usage demand corresponding to the fourth user type;

[0059] The seventh determination unit is used to determine at least one user in the second sub-training data set who does not meet the first preset condition and does not meet the second preset condition as the fifth user type through the user classification model; the traffic usage demand corresponding to the fourth user type is greater than the traffic usage demand corresponding to the fifth user type.

[0060] The present invention also discloses an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the user classification method described above are implemented.

[0061] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the user classification method described above are implemented.

[0062] The embodiments of the present invention include the following advantages:

[0063] In an embodiment of the present invention, a training dataset is input into a user classification model. The user classification model then determines the correlation coefficients between the traffic consumption, traffic package arrival frequency, and traffic package arrival interval of multiple users in the training dataset and the amount of traffic package purchases. The multiple users are then classified based on the correlation coefficients to train the user classification model. The current user's traffic usage data is then obtained, and the user type of the current user is determined based on the traffic usage data and the trained user classification model. Different user types correspond to different traffic usage requirements. Thus, when training the user classification model, the amount of traffic package purchases, the traffic consumption, the frequency of traffic package arrivals, and the interval between traffic package arrivals are combined for training. This allows the trained user classification model to determine the user type of users based on the frequency of traffic package arrivals, enabling accurate classification of the user types of users of in-vehicle network entertainment traffic. Furthermore, when training the user classification model, the classification of multiple users based on the correlation coefficients is performed, which is a label-free learning method, making training the user classification model simpler. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 This is a flowchart of the steps of a user classification method provided by an embodiment of the present invention;

[0065] Figure 2 This is a flowchart of the steps of a user classification method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0066] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0067] In the related art, the method of user classification is not applicable to the scenario of classifying users of car network entertainment traffic. An embodiment of the present invention provides a user classification method, the core concept of which is to input a training data set into a user classification model, and use the user classification model to determine the correlation coefficients of the traffic consumption, arrival frequency and arrival interval of the purchased traffic packages of multiple users in the training data set and the consumption amount of the purchased traffic packages, and classify multiple users based on the correlation coefficients to train the user classification model. Then, the traffic usage data of the current user is obtained, and the user type of the current user is determined based on the traffic usage data of the current user and the trained user classification model, wherein different user types correspond to different traffic usage requirements. Therefore, when training the user classification model, the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package and the arrival interval of the purchased traffic package are combined for training. When the trained user classification model determines the user type of the user, it is based on the arrival frequency of the purchased traffic package. It can accurately classify the user types of users of the car network entertainment traffic. Moreover, when training the user classification model, multiple users are classified based on the correlation coefficient, which is unlabeled learning, making training the user classification model simpler.

[0068] Reference Figure 1 , shows a flowchart of a user classification method provided by an embodiment of the present invention, the method may specifically include the following steps:

[0069] Step 101, obtain a training data set, which includes traffic usage data of multiple users; the traffic usage data includes the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval time of the purchased traffic package.

[0070] In an embodiment of the present invention, before obtaining a training data set, sample data of multiple users may be collected first. The sample data may include: the amount of money spent on the user's traffic package, the traffic consumed by the purchased traffic package, the frequency of arrival of the purchased traffic package, the interval time for arrival of the purchased traffic package, the real-name authentication time, the usage of the free traffic package, the validity period of the free traffic package, the purchase record of the traffic package, etc. In order to make the user classification model more consistent with the actual business scenario and enhance the prediction effect of the user classification model, it is necessary to filter the sample data to obtain a training data set. When filtering the sample data, users and corresponding sample data whose real-name authentication time is no more than thirty days may be eliminated based on the real-name authentication time, so that the sample data of users who have not been real-name authenticated or have a short real-name authentication time will not participate in the training of the user classification model. In the present invention, the traffic package can be a traffic package for the car network entertainment traffic.

[0071] In the present invention, in order to ensure the accuracy and timeliness of the trained user classification model, sample data within a preset time period from now can be selected for the sample data of users retained after screening. For example, sample data within one year from now can be selected as the training data set, wherein the training data set may include traffic usage data of multiple users, and the traffic usage data may include the consumption amount of purchased traffic packages, the consumed traffic of purchased traffic packages, the arrival frequency of purchased traffic packages, and the arrival interval of purchased traffic packages.

[0072] In an embodiment of the present invention, the consumption amount of the traffic package purchased can be obtained by obtaining the subscription amount of the traffic package, multiplied by the ratio of the number of times the traffic package is superimposed in the statistical time period to the total number of times the traffic package is superimposed, wherein the total number of superimposed times is the number of months in the time period from the time the traffic package takes effect to the time it expires. The statistical time period can be a preset time period for reference when filtering out the training data set from the sample data. The traffic consumed by the purchased traffic package is the sum of the traffic consumed by the traffic packages purchased in the statistical time period. The frequency of arrival of the purchased traffic package is the total number of times the traffic packages purchased are superimposed in the statistical time period. The arrival interval of the purchased traffic package is the shortest interval between the arrival time of the purchased traffic package and the current time interval.

[0073] Step 102: Input the training data set into the user classification model, and determine the correlation coefficients of the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package with the consumption amount of the purchased traffic package through the user classification model, and classify the multiple users based on the correlation coefficients to train the user classification model.

[0074] In an embodiment of the present invention, after obtaining a training data set, the training data set can be input into a user classification model. The user classification model can determine the correlation coefficients of the consumed traffic, the arrival frequency, and the arrival interval of the purchased traffic package in the training set data and the consumption amount of the purchased traffic package, and classify multiple users in the training data set based on the correlation coefficient to train the user classification model.

[0075] In one embodiment, inputting the training data set into the user classification model may include: determining the traffic usage data of multiple users in the training data set whose consumption amount for purchasing traffic packages is greater than zero as the first sub-training data set; normalizing the first sub-training data set to obtain a normalized first sub-training data set; and inputting the normalized first sub-training data set into the user classification model.

[0076] In the present invention, the traffic usage data of multiple users in the training data set whose consumption amount for purchasing traffic packages is greater than zero can be determined as the first training data set, that is, the traffic usage data of multiple users in the training data set who have purchased traffic packages can be determined as the first training data set. Then, the "max-min normalization" method is used to normalize the first sub-training data set to obtain the normalized first sub-training data set. The present invention can normalize the first sub-training data set by the following formula (1):

[0077]

[0078] Among them, x norm represents the first sub-training dataset after normalization, X represents the consumption amount (M) of the purchased traffic package, the consumed traffic (C) of the purchased traffic package, the frequency of the purchase of the traffic package (F), and the interval time (R) of the purchase of the traffic package in the first sub-training dataset before normalization, x min represents the corresponding minimum value in the first sub-training data set before normalization, x max Represents the corresponding maximum values ​​in the first sub-training dataset before normalization.

[0079] In order to better understand and illustrate the embodiments of the present invention, the subsequent M in the embodiments of the present invention represents the consumption amount of the purchased traffic package, C represents the consumed traffic of the purchased traffic package, F represents the arrival frequency of the purchased traffic package, and R represents the arrival interval of the purchased traffic package.

[0080] After obtaining the normalized first sub-training data set, the normalized first sub-training data set can be input into the user classification model. Then, when training the user classification model, the normalized first sub-training data set is used to improve the training efficiency and performance of the user classification model.

[0081] It should be noted that in the present invention, when determining the training data set from the sample data, the test data set can be determined in the same manner and preset ratio as that of determining the training data set, and then the normalized first sub-test data set can be determined in the same manner as that of determining the normalized first sub-training data set. The first sub-test data set is used to test the trained user classification model.

[0082] In one embodiment, determining the correlation coefficients of the traffic consumption of a purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package, respectively, with the consumption amount of the purchased traffic package, through a user classification model, can include: clustering multiple users according to the traffic usage data of multiple users in a first sub-training data set through a user classification model to obtain multiple clusters; the clusters include at least one user; and determining the correlation coefficients of the traffic consumption of a purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package for users in each cluster, respectively, with the consumption amount of the purchased traffic package, through a user classification model.

[0083] In the present invention, after the normalized first sub-training dataset is input into the user classification model, the user classification model can cluster the multiple users based on their M, C, F, and R in the first sub-training dataset to obtain multiple clusters, each of which includes at least one user. The user classification model can then determine, for each cluster, the correlation coefficient between the three dimensions C, F, and R and M in each cluster. This allows clustering to be performed during training of the user classification model, improving its performance while reducing its computational complexity.

[0084] In the present invention, when the user classification model clusters multiple users to obtain multiple clusters, a data point can be randomly selected from the first training data set as a cluster center. One data point corresponds to one user. This can be understood as randomly selecting one user from the first training data set as a cluster center. Then, for each data point in the first training data set, the shortest distance to the selected cluster center is calculated. The shortest distance from each data point to the cluster center can be calculated using the following formula (2):

[0085]

[0086] Among them, d(z) represents the shortest distance from data point z to the cluster center, x represents the data point, c i Represents the cluster center.

[0087] Then, the probability of each data point being selected as the next cluster center is calculated based on the square of the shortest distance, where the probability of each data point being selected as the next cluster center can be calculated using the following formula (3):

[0088]

[0089] Where P(x) represents the probability that data point x is selected as the next cluster center, D represents the first training data set, d(z) represents the shortest distance from data point z to the cluster center, and d(y) represents the distance from all data points in the first training data set to the cluster center.

[0090] Then, based on the distribution of the probability of each data point being selected as the next cluster center, the next cluster center is randomly selected until a preset number of cluster centers are selected. In the present invention, the preset number can be three, that is, three cluster centers are selected. Each point in the first training data set is then assigned to the cluster to which the closest cluster center belongs. The mean of the data points in each cluster is calculated, and the cluster center is updated based on the mean until the updated cluster center no longer changes or a preset number of iterations is reached. Each data point in the first training data set is then assigned to the cluster to which the closest updated cluster center belongs, thereby obtaining multiple clusters.

[0091] In one embodiment, the user classification model is used to determine, for each cluster, the traffic consumption of the traffic packages purchased by users in each cluster, the frequency of arrival of the traffic packages purchased, and the interval time for arrival of the traffic packages purchased, as well as the correlation coefficients thereof with the consumption amount of the traffic packages purchased, including: determining, for each cluster, the first average value of the consumption amount of the traffic packages purchased by users in each cluster, the second average value of the traffic consumption, the third average value of the frequency of arrival of the traffic packages purchased, and the fourth average value of the interval time for arrival of the traffic packages purchased; and determining, based on the first average value, the second average value, the third average value, the fourth average value, the consumption amount of the traffic packages purchased, the traffic consumption of the traffic packages purchased, the frequency of arrival of the traffic packages purchased, and the interval time for arrival of the traffic packages purchased, the first correlation coefficient between the traffic consumption of the traffic packages purchased and the consumption amount of the traffic packages purchased, the second correlation coefficient between the frequency of arrival of the traffic packages purchased and the consumption amount of the traffic packages purchased, and the third correlation coefficient between the interval time for arrival of the traffic packages purchased and the consumption amount of the traffic packages purchased, for users in each cluster based on the first average value, the second average value, the third average value, the fourth average value, the consumption amount of the traffic packages purchased, the traffic consumption of the traffic packages purchased, the frequency of arrival of the traffic packages purchased, and the interval time for arrival of the traffic packages purchased.

[0092] In the present invention, when the user classification model determines the correlation coefficients of the consumption flow of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package of the users in each cluster, respectively, with the consumption amount of the purchased traffic package, the user classification model can calculate the average value of M, C, F, and R of the users in each cluster for each cluster, that is, the first average value of the consumption amount of the purchased traffic package. The second average value of the consumed traffic of the purchased traffic package The third average value of the frequency of arrival of the purchased traffic package and the fourth average value of the arrival interval of the purchased traffic package The user classification model then determines the first correlation coefficient between the traffic consumption of the purchased traffic package and the consumption amount of the purchased traffic package, the second correlation coefficient between the frequency of arrival of the purchased traffic package and the consumption amount of the purchased traffic package, and the third correlation coefficient between the interval time of arrival of the purchased traffic package and the consumption amount of the purchased traffic package for users in each cluster based on the first average value, the second average value, the third average value and the fourth average value, the consumption amount of the purchased traffic package, the consumption traffic of the purchased traffic package, the frequency of arrival of the purchased traffic package and the interval time of arrival of the purchased traffic package. The present invention can calculate the correlation coefficients of the consumption traffic of the purchased traffic package, the frequency of arrival of the purchased traffic package and the interval time of arrival of the purchased traffic package, and the consumption amount of the purchased traffic package by the following formula (4):

[0093]

[0094] Among them, PR represents the correlation coefficient between the traffic consumption of the purchased traffic package, the frequency of the traffic package arrival, and the interval between the traffic package arrivals and the amount of the traffic package purchase, and X i Indicates the values ​​of C, F, and R. Represents the mean corresponding to C, F, and R, M i Represents the various values ​​of M, Represents the mean corresponding to M.

[0095] The first correlation coefficient PR of the traffic consumption of the purchased traffic package and the amount of the purchased traffic package can be calculated by the above formula (4): CM The second correlation coefficient PR between the frequency of receipt of the purchased traffic package and the amount of consumption of the purchased traffic package FM The third correlation coefficient PR between the arrival interval of the purchased data package and the consumption amount of the purchased data package RM. Thus, the present invention calculates the correlation coefficients of the traffic consumption, traffic arrival frequency and traffic arrival interval of the traffic packages purchased by users in each cluster through the average of the traffic consumption amount, traffic consumption, traffic arrival frequency and traffic arrival interval of the traffic packages purchased by users in each cluster and the traffic consumption amount of the traffic packages purchased, so that the determined correlation coefficient is more accurate. Further, when the user classification model is trained by the correlation coefficient, the trained user classification model is more accurate and the complexity of the user classification model is reduced.

[0096] In one embodiment, the clusters may include a first cluster, a second cluster, and a third cluster; classifying multiple users based on the correlation coefficient may include: determining, through a user classification model, according to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient, a first weight of the consumption amount of the purchased traffic package, a second weight of the consumed traffic of the purchased traffic package, a third weight of the arrival frequency of the purchased traffic package, and a fourth weight of the arrival interval time of the purchased traffic package for each user in the cluster; and classifying, through a user classification model, according to the first weight, the second weight, the third weight, the fourth weight, the first average value, the second average value, the third average value ... The mean and the fourth mean are used to determine the comprehensive score of each cluster; the users in the first cluster are determined as the first user type, the users in the second cluster are determined as the second user type, and the users in the third cluster are determined as the third user type through the user classification model; the comprehensive score of the first cluster is greater than the comprehensive score of the second cluster and the comprehensive score of the third cluster; the comprehensive score of the second cluster is greater than the comprehensive score of the third cluster; the traffic usage demand corresponding to the first user type is greater than the traffic usage demand corresponding to the second user type; the traffic usage demand corresponding to the second user type is greater than the traffic usage demand corresponding to the third user.

[0097] In an embodiment of the present invention, the preset number of cluster centers can be three, so the corresponding number of clusters is also three, namely the first cluster, the second cluster and the third cluster. The user classification model can be based on the correlation coefficient of the traffic consumption of the purchased traffic package, the arrival frequency of the purchased traffic package and the arrival interval of the purchased traffic package of the users in each cluster, and the consumption amount of the purchased traffic package, to determine the weights corresponding to the consumption amount of the purchased traffic package, the traffic consumption of the purchased traffic package, the arrival frequency of the purchased traffic package and the arrival interval of the purchased traffic package of the users in each cluster, and determine the comprehensive score of each cluster, and then classify the users in the first cluster, the users in the second cluster and the users in the third cluster by the comprehensive score, so that the user classification model classifies the users in the cluster based on the comprehensive score, and the classification result obtained is more accurate.

[0098] For example: the calculated comprehensive score of the first cluster is 0.53, the comprehensive score of the second cluster is 0.40, and the comprehensive score of the third cluster is 0.13. Then, based on the size of the comprehensive score, the users in the first cluster are determined as the first user type, and the first user type can be important value users. The users in the second cluster are determined as the second user type, and the second user type can be important retention users. The users in the third cluster are determined as the third user type, and the third user type can be important development users.

[0099] In one embodiment, the user classification model is used to determine the first weight of the consumption amount of the purchased traffic package, the second weight of the consumed traffic of the purchased traffic package, the third weight of the arrival frequency of the purchased traffic package, and the fourth weight of the arrival interval of the purchased traffic package for users in each cluster according to the first correlation coefficient, the second correlation coefficient and the third correlation coefficient. It can include: determining the total coefficient according to the first correlation coefficient, the second correlation coefficient and the third correlation coefficient through the user classification model; determining the first weight of the consumption amount of the purchased traffic package, the second weight of the consumed traffic of the purchased traffic package, the third weight of the arrival frequency of the purchased traffic package, and the fourth weight of the arrival interval of the purchased traffic package through the user classification model according to the first correlation coefficient, the second correlation coefficient, the third correlation coefficient and the total coefficient.

[0100] In the present invention, the weights corresponding to the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package are calculated based on the correlation coefficients of the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package with the consumption amount of the purchased traffic package, so that the obtained weights are not set arbitrarily based on demand, but are determined according to the user's usage data, so that the obtained weights are more accurate, and further the user classification model is trained to classify users based on the weights, so that the user types determined by the trained user classification model are more accurate.

[0101] In the present invention, the total coefficient can be calculated by the following formula (5):

[0102] PR total =|PR MM |+|PR CM |+|PR FM |+|PR EM | Formula (5)

[0103] Among them, PR total is the total coefficient, |PR MM | is 1, PR MMIndicates the correlation coefficient between the amount of consumption of the purchased traffic package and itself, PR CM The first correlation coefficient between the traffic consumed by the purchased traffic package and the amount of money spent on the purchased traffic package, PR FM The second correlation coefficient, PR, represents the frequency of receipt of traffic packages and the amount of money spent on purchasing traffic packages. RM The third correlation coefficient between the arrival interval of the purchased data package and the consumption amount of the purchased data package.

[0104] In the present invention, the corresponding weight can be calculated by the following formula (6):

[0105]

[0106] Among them, a is the first weight of the consumption amount of purchasing traffic package, PR total is the total coefficient, b is the second weight of the traffic consumption of the purchased traffic package, c is the third weight of the arrival frequency of the purchased traffic package, d is the fourth weight of the arrival interval of the purchased traffic package, PR CM The first correlation coefficient between the traffic consumed by the purchased traffic package and the amount of money spent on the purchased traffic package, PR FM The second correlation coefficient, PR, represents the frequency of receipt of traffic packages and the amount of money spent on purchasing traffic packages. RM The third correlation coefficient between the arrival interval of the purchased data package and the consumption amount of the purchased data package.

[0107] In the present invention, the comprehensive score of the cluster can be calculated by the following formula (7):

[0108]

[0109] Among them, H represents the comprehensive score of the cluster, a is the first weight of the amount of money spent on the purchase of the data package, b is the second weight of the traffic consumed by the purchase of the data package, c is the third weight of the frequency of the arrival of the data package, and d is the fourth weight of the arrival interval of the data package. Indicates the first average value of the amount of money spent on purchasing a data package. Indicates the second average value of traffic consumption of purchased traffic packages. The third average value of the frequency of arrival of purchased data packages, The fourth average value of the interval between receipt of purchased data packages

[0110] In an embodiment of the present invention, after determining the traffic usage data of multiple users in the training data set whose consumption amount for purchasing traffic packages is greater than zero as the first sub-training data set, it may also include: determining the traffic usage data of multiple users in the training data set whose consumption amount for purchasing traffic is equal to zero as the second sub-training data set, inputting the second sub-training data set into the user classification model, and classifying the multiple users in the second sub-training data set based on the second sub-training data set and preset classification conditions to train the user classification model.

[0111] In the present invention, after determining the first sub-training data set, the traffic usage data of multiple users in the training data set whose consumption amount for purchased traffic is equal to zero can be determined as the second sub-training data set, that is, the traffic usage data of multiple users in the training data set who do not hold purchased traffic packages are determined as the second sub-training data set, and then the second sub-training data set is input into the user classification model. The user classification model classifies multiple users in the second sub-training data set based on the second training data set and preset classification conditions to train the user classification model, so that the trained user classification model classifies users who do not hold purchased traffic packages according to the preset classification conditions, and will not classify based on the correlation coefficient, and then the trained user classification model adopts different classification methods for different users, so that the trained user classification model is more comprehensive and the trained user classification model is more accurate in classifying users.

[0112] After obtaining the second sub-training data set, the present invention can normalize the second sub-training data set to obtain a normalized second sub-training data set. After obtaining the normalized second sub-training data set, the normalized second sub-training data set can be input into the user classification model. Then, when training the user classification model, the normalized second sub-training data set is used, thereby improving the training efficiency and performance of the user classification model.

[0113] It should be noted that in the present invention, when determining the training data set from the sample data, the test data set can be determined in the same manner and preset ratio as that of determining the training data set, and then a normalized second sub-test data set can be determined in the same manner as that of determining the normalized second sub-training data set. The second sub-test data set is used to test the trained user classification model.

[0114] In one embodiment, the traffic usage data also includes: the usage of the free traffic package, the validity period of the free traffic package and the purchase record of the traffic package; the preset classification conditions include: a first preset condition and a second preset condition; the first preset condition is that the usage of the user's free traffic package is greater than the usage threshold; the second preset condition is that the validity period of the user's free traffic package is invalid and there is a purchase record of the traffic package; the classification of the multiple users in the second sub-training data set based on the second sub-training data set and the preset classification conditions includes: determining at least one user in the second sub-training data set that meets the first preset condition as a third user type through the user classification model; determining at least one user in the second sub-training data set that meets the second preset condition as a fourth user type through the user classification model; the traffic usage demand corresponding to the third user type is greater than the traffic usage demand corresponding to the fourth user type; determining at least one user in the second sub-training data set that does not meet the first preset condition and does not meet the second preset condition as a fifth user type through the user classification model; the traffic usage demand corresponding to the fourth user type is greater than the traffic usage demand corresponding to the fifth user type.

[0115] For example: the first preset condition is that the user's usage of the free traffic package is greater than 80%, which means that the user does not hold the purchased traffic package, but the usage of the free traffic package is large, which also means that users who meet the first preset condition have a greater demand for traffic packages, have greater potential to purchase traffic packages, and have higher user value. The user classification model can determine the users who meet the first preset condition in the second training data set as the third user type, and the third user type is an important development user. The second preset condition is that the validity period of the free traffic package has expired and there is a record of purchasing the traffic package. The record of purchasing the traffic package can be a record of purchasing the traffic package within a set time period, such as a record of purchasing the traffic package within 1.5 years. When the user meets the second preset condition, it means that the user's free traffic package has expired and there is an act of purchasing the traffic package. At the same time, there is no traffic package available, which also means that the user has certain purchasing potential and can be saved. Then, the user classification model can determine the user who meets the second preset condition as the fourth user type, and the fourth user type is an important retained user. Finally, the user classification model can determine the remaining users, that is, the remaining users who do not belong to the first user type, the second user type, the third user type and the fourth user type, as the fifth user type. The fifth user type is a general value user, indicating that the user of this user type has a low usage rate of the traffic package and has no great purchasing motivation. The present invention accurately classifies users who meet the preset conditions according to the refined preset conditions when training the user classification model, so that the trained user classification model is more comprehensive and the trained user classification model is more accurate in classifying users.

[0116] In an embodiment of the present invention, after training the user classification model, the trained user classification model can be subjected to an index test based on a test data set, i.e., the first test data set and the second test data set mentioned above. After the test data set is input into the trained user classification model, it can be determined whether the accuracy of the user type determined by the user classification model meets the standard. If the standard is met, it can be determined that the training of the user classification model is completed. If it does not meet the standard, the training will continue.

[0117] Step 103: Obtain the traffic usage data of the current user, and determine the user type of the current user based on the traffic usage data of the current user and the trained user classification model; wherein different user types correspond to different traffic usage requirements.

[0118] In an embodiment of the present invention, after the user classification model is trained, the traffic usage data of the current user can be obtained, and the traffic usage data of the current user can be input into the trained user classification model to determine the user type of the current user through the trained user classification model.

[0119] In the present invention, after determining the user type of the current user, corresponding traffic packages can be recommended for different user types, or corresponding marketing activities can be formulated, so as to improve the conversion rate of traffic packages and increase the sales of traffic packages.

[0120] In the present invention, the retention of each user type can be observed after a period of time, for example, sixty days, when the users continue with the previous business model, and the user retention rate reference value is set to 0.9. That is, after a period of time, when the average retention rate of each user type of the user classification model reaches 0.9 or above, it indicates that the user type determined by the trained classification model has a certain stability and the determination result is reliable. If the retention rate does not reach 0.8, the user classification model can be retrained. When retraining, the latest traffic usage data can be used for training to improve the accuracy and timeliness of the user classification model.

[0121] In an embodiment of the present invention, a training data set is obtained, and the training data set includes traffic usage data of multiple users; the traffic usage data includes the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package; the training data set is input into a user classification model, and the user classification model is used to determine the correlation coefficients of the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package, respectively, with the consumption amount of the purchased traffic package, and multiple users are classified based on the correlation coefficients to train the user classification model; the traffic usage data of the current user is obtained, and the user type of the current user is determined according to the traffic usage data of the current user and the trained user classification model; wherein different user types correspond to different traffic usage requirements. Therefore, when training the user classification model, the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package and the arrival interval of the purchased traffic package are combined for training. When the trained user classification model determines the user type of the user, it is based on the arrival frequency of the purchased traffic package. It can accurately classify the user types of users of the car network entertainment traffic. Moreover, when training the user classification model, multiple users are classified based on the correlation coefficient, which is unlabeled learning, making training the user classification model simpler.

[0122] It should be noted that for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0123] Reference Figure 2 , shows a structural block diagram of a device for determining a user type provided by an embodiment of the present invention, which may specifically include the following modules:

[0124] Acquisition module 201 is used to acquire a training data set, wherein the training data set includes traffic usage data of multiple users; the traffic usage data includes the amount of money spent on the purchased traffic package, the traffic consumed by the purchased traffic package, the frequency of the receipt of the purchased traffic package, and the interval between the receipt of the purchased traffic package;

[0125] A training module 202 is configured to input the training data set into a user classification model, determine, through the user classification model, a correlation coefficient between the traffic consumption of the purchased traffic package, the frequency of arrival of the purchased traffic package, and the interval between arrivals of the purchased traffic package, and the amount of the traffic package consumed, and classify the multiple users based on the correlation coefficients to train the user classification model;

[0126] The determination module 203 is used to obtain the traffic usage data of the current user, and determine the user type of the current user based on the traffic usage data of the current user and the trained user classification model; wherein different user types correspond to different traffic usage requirements.

[0127] In one embodiment, the training module 202 includes:

[0128] A first determining submodule is configured to determine the traffic usage data of a plurality of users in the training data set whose consumption amount of the purchased traffic packages is greater than zero as a first sub-training data set;

[0129] a normalization module, configured to perform normalization processing on the first sub-training data set to obtain a normalized first sub-training data set;

[0130] The first input module is configured to input the normalized first sub-training data set into a user classification model.

[0131] In one embodiment, the training module 202 includes:

[0132] a clustering submodule, configured to cluster the plurality of users according to the traffic usage data of the plurality of users in the first sub-training data set using the user classification model to obtain a plurality of clusters; each cluster includes at least one user;

[0133] The second determination submodule is used to determine, for each cluster, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package of the users in each cluster through the user classification model, and the correlation coefficients thereof with the consumption amount of the purchased traffic package.

[0134] In one embodiment, the second determining submodule includes:

[0135] a first determining unit, configured to determine, for each cluster, by using the user classification model, a first average consumption amount of the purchased traffic packages, a second average consumption amount of the purchased traffic packages, a third average frequency of arrival of the purchased traffic packages, and a fourth average interval between arrivals of the purchased traffic packages for users in each cluster;

[0136] The second determination unit is used to determine, through the user classification model and based on the first average value, the second average value, the third average value, the fourth average value, the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package of the users in each cluster, a first correlation coefficient between the consumed traffic of the purchased traffic package and the consumption amount of the purchased traffic package, a second correlation coefficient between the arrival frequency of the purchased traffic package and the consumption amount of the purchased traffic package, and a third correlation coefficient between the arrival interval of the purchased traffic package and the consumption amount of the purchased traffic package.

[0137] In one embodiment, the clusters include a first cluster, a second cluster, and a third cluster; the training module 202 includes:

[0138] a third determination submodule, configured to determine, by using the user classification model and according to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient, a first weight of the consumption amount of the purchased traffic package, a second weight of the consumed traffic of the purchased traffic package, a third weight of the arrival frequency of the purchased traffic package, and a fourth weight of the arrival interval of the purchased traffic package for each user in the cluster;

[0139] a fourth determination submodule, configured to determine a comprehensive score for each cluster according to the first weight, the second weight, the third weight, the fourth weight, the first average value, the second average value, the third average value, and the fourth average value through the user classification model;

[0140] The fifth determination submodule is used to determine the users in the first cluster as the first user type, the users in the second cluster as the second user type, and the users in the third cluster as the third user type through the user classification model; the comprehensive score of the first cluster is greater than the comprehensive score of the second cluster and the comprehensive score of the third cluster; the comprehensive score of the second cluster is greater than the comprehensive score of the third cluster; the traffic usage demand corresponding to the first user type is greater than the traffic usage demand corresponding to the second user type; the traffic usage demand corresponding to the second user type is greater than the traffic usage demand corresponding to the third user.

[0141] In one embodiment, the third determining submodule includes:

[0142] a third determining unit, configured to determine a total coefficient according to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient by using the user classification model;

[0143] The fourth determination unit is used to determine the first weight of the consumption amount of the purchased traffic package of each user in the cluster, the second weight of the consumed traffic of the purchased traffic package, the third weight of the arrival frequency of the purchased traffic package, and the fourth weight of the arrival interval time of the purchased traffic package through the user classification model according to the first correlation coefficient, the second correlation coefficient, the third correlation coefficient and the total coefficient.

[0144] In one embodiment, the training module 202 further includes:

[0145] The sixth determining submodule is configured to determine the traffic usage data of a plurality of users whose traffic purchase amount is zero in the training data set as a second sub-training data set,

[0146] The second input submodule is used to input the second sub-training data set into the user classification model, and classify the multiple users in the second sub-training data set based on the second sub-training data set and preset classification conditions to train the user classification model.

[0147] In one embodiment, the traffic usage data further includes: usage of the free traffic package, validity period of the free traffic package, and a record of purchased traffic packages; the preset classification conditions include: a first preset condition and a second preset condition; the first preset condition is that the usage of the user's free traffic package is greater than a usage threshold; the second preset condition is that the validity period of the user's free traffic package is expired and there is a record of purchased traffic packages; the second input submodule includes:

[0148] a fifth determining unit, configured to determine, by using the user classification model, at least one user in the second sub-training data set that meets a first preset condition as a third user type;

[0149] a sixth determining unit, configured to determine, by using the user classification model, at least one user in the second sub-training data set that meets a second preset condition as a fourth user type; and wherein the traffic usage demand corresponding to the third user type is greater than the traffic usage demand corresponding to the fourth user type;

[0150] The seventh determination unit is used to determine at least one user in the second sub-training data set who does not meet the first preset condition and does not meet the second preset condition as the fifth user type through the user classification model; the traffic usage demand corresponding to the fourth user type is greater than the traffic usage demand corresponding to the fifth user type.

[0151] In an embodiment of the present invention, an acquisition module is used to obtain a training data set, which includes traffic usage data of multiple users; the traffic usage data includes the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package; the training module is used to input the training data set into a user classification model, and determine the correlation coefficients of the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package with the consumption amount of the purchased traffic package through the user classification model, and classify multiple users based on the correlation coefficients to train the user classification model; the determination module is used to obtain the traffic usage data of the current user, and determine the user type of the current user according to the traffic usage data of the current user and the trained user classification model; wherein different user types correspond to different traffic usage requirements. Therefore, when training the user classification model, the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package and the arrival interval of the purchased traffic package are combined for training. When the trained user classification model determines the user type of the user, it is based on the arrival frequency of the purchased traffic package. It can accurately classify the user types of users of the car network entertainment traffic. Moreover, when training the user classification model, multiple users are classified based on the correlation coefficient, which is unlabeled learning, making training the user classification model simpler.

[0152] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0153] An embodiment of the present invention further provides an electronic device, including:

[0154] It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned user classification method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0155] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned user classification method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0156] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0157] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, embodiments of the present invention may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0158] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0159] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0161] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0162] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0163] The above is a detailed introduction to a user classification method, a user classification device, an electronic device and a computer-readable storage medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A user classification method, characterized in that: include: Obtaining a training data set, wherein the training data set includes traffic usage data of multiple users; The traffic usage data includes the amount of money spent on the purchased traffic package, the traffic consumed by the purchased traffic package, the frequency of receipt of the purchased traffic package, and the interval between receipts of the purchased traffic package; Inputting the training data set into a user classification model, determining, through the user classification model, the correlation coefficients of the traffic consumption of the purchased traffic package, the frequency of arrival of the purchased traffic package, and the interval between arrival of the purchased traffic package, respectively, and the consumption amount of the purchased traffic package, and classifying the multiple users based on the correlation coefficients to train the user classification model; Obtain the traffic usage data of the current user, and determine the user type of the current user based on the traffic usage data of the current user and the trained user classification model; wherein different user types correspond to different traffic usage requirements.

2. The user classification method according to claim 1, characterized in that: The step of inputting the training data set into the user classification model comprises: Determining the traffic usage data of a plurality of users in the training data set whose consumption amount of the purchased traffic packages is greater than zero as a first sub-training data set; Normalizing the first sub-training data set to obtain a normalized first sub-training data set; The normalized first sub-training data set is input into a user classification model.

3. The user classification method according to claim 2, characterized in that: The correlation coefficients of the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package determined by the user classification model and the consumption amount of the purchased traffic package, respectively, include: Clustering the plurality of users according to the traffic usage data of the plurality of users in the first sub-training data set using the user classification model to obtain a plurality of clusters; each cluster includes at least one user; The user classification model is used to determine, for each cluster, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package of the users in each cluster, and their correlation coefficients with the consumption amount of the purchased traffic package.

4. The user classification method according to claim 3, characterized in that: The method of determining, for each cluster, the consumed traffic volume of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package for each user in the cluster using the user classification model, and the correlation coefficients thereof with the consumption amount of the purchased traffic package, respectively, includes: Determining, for each cluster, using the user classification model, a first average value of the amount of consumption of the purchased traffic packages, a second average value of the traffic consumed by the purchased traffic packages, a third average value of the frequency of arrival of the purchased traffic packages, and a fourth average value of the interval between arrivals of the purchased traffic packages for users in each cluster; The user classification model is used to determine, based on the first average value, the second average value, the third average value, the fourth average value, the consumption amount of the purchased traffic package, the consumed traffic of the purchased traffic package, the arrival frequency of the purchased traffic package, and the arrival interval of the purchased traffic package, a first correlation coefficient between the consumed traffic of the purchased traffic package and the consumption amount of the purchased traffic package, a second correlation coefficient between the arrival frequency of the purchased traffic package and the consumption amount of the purchased traffic package, and a third correlation coefficient between the arrival interval of the purchased traffic package and the consumption amount of the purchased traffic package for each user in the cluster.

5. The user classification method according to claim 4, characterized in that: The clusters include a first cluster, a second cluster, and a third cluster; and classifying the plurality of users based on the correlation coefficients includes: Determining, by the user classification model, based on the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient, respectively, a first weight of the consumption amount of the purchased traffic package, a second weight of the consumed traffic of the purchased traffic package, a third weight of the arrival frequency of the purchased traffic package, and a fourth weight of the arrival interval of the purchased traffic package for each user in the cluster; determining a comprehensive score for each cluster using the user classification model according to the first weight, the second weight, the third weight, the fourth weight, the first average value, the second average value, the third average value, and the fourth average value; Through the user classification model, the users in the first cluster are determined as the first user type, the users in the second cluster are determined as the second user type, and the users in the third cluster are determined as the third user type; the comprehensive score of the first cluster is greater than the comprehensive score of the second cluster and the comprehensive score of the third cluster; the comprehensive score of the second cluster is greater than the comprehensive score of the third cluster; the traffic usage demand corresponding to the first user type is greater than the traffic usage demand corresponding to the second user type; the traffic usage demand corresponding to the second user type is greater than the traffic usage demand corresponding to the third user.

6. The user classification method according to claim 5, characterized in that: The method of determining, by the user classification model according to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient, a first weight of the consumption amount of the purchased traffic package, a second weight of the consumed traffic of the purchased traffic package, a third weight of the arrival frequency of the purchased traffic package, and a fourth weight of the arrival interval of the purchased traffic package for each user in the cluster, respectively, includes: determining a total coefficient according to the first correlation coefficient, the second correlation coefficient, and the third correlation coefficient by using the user classification model; The user classification model is used to determine the first weight of the consumption amount of the purchased traffic package of the user in each cluster, the second weight of the consumed traffic of the purchased traffic package, the third weight of the arrival frequency of the purchased traffic package, and the fourth weight of the arrival interval of the purchased traffic package.

7. The user classification method according to claim 2, characterized in that: After determining the traffic usage data of a plurality of users in the training data set whose consumption amount for purchasing the traffic package is greater than zero as the first sub-training data set, the method further includes: Determine the traffic usage data of multiple users in the training data set whose traffic purchase amount is zero as the second sub-training data set, The second sub-training data set is input into the user classification model, and based on the second sub-training data set and preset classification conditions, the multiple users in the second sub-training data set are classified to train the user classification model.

8. The user classification method according to claim 7, characterized in that: The traffic usage data also includes: usage of the free traffic package, validity period of the free traffic package, and purchase record of the traffic package; the preset classification conditions include: a first preset condition and a second preset condition; the first preset condition is that the usage of the free traffic package of the user is greater than the usage threshold; the second preset condition is that the validity period of the free traffic package of the user is expired and there is a purchase record of the traffic package; the classification of the multiple users in the second sub-training data set based on the second sub-training data set and the preset classification conditions includes: Determining at least one user in the second sub-training data set that meets a first preset condition as a third user type by using the user classification model; Determining, by the user classification model, at least one user in the second sub-training data set that meets a second preset condition as a fourth user type; the traffic usage demand corresponding to the third user type is greater than the traffic usage demand corresponding to the fourth user type; The user classification model is used to determine at least one user in the second sub-training data set who does not meet the first preset condition and does not meet the second preset condition as the fifth user type; the traffic usage demand corresponding to the fourth user type is greater than the traffic usage demand corresponding to the fifth user type.

9. A user classification device, characterized in that: include: An acquisition module, configured to acquire a training data set, wherein the training data set includes traffic usage data of multiple users; The traffic usage data includes the amount of money spent on the purchased traffic package, the traffic consumed by the purchased traffic package, the frequency of receipt of the purchased traffic package, and the interval between receipts of the purchased traffic package; a training module, configured to input the training data set into a user classification model, determine, through the user classification model, a correlation coefficient between the traffic consumption of the purchased traffic package, the frequency of arrival of the purchased traffic package, and the interval between arrivals of the purchased traffic package, and the consumption amount of the purchased traffic package, respectively, and classify the multiple users based on the correlation coefficients to train the user classification model; A determination module is used to obtain the traffic usage data of the current user, and determine the user type of the current user based on the traffic usage data of the current user and the trained user classification model; wherein different user types correspond to different traffic usage requirements.

10. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of the user classification method according to any one of claims 1 to 8 are implemented.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the user classification method according to any one of claims 1 to 8 are implemented.