User Classification Model Optimization Method, System, Electronic Device and Storage Medium
By obtaining user and device behavior data, building user portraits based on the user tag system and optimizing the user classification model, the problem of low user classification accuracy in the existing technology is solved, and more efficient user classification is achieved.
Patent Information
- Application Number
- CN202311295728.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-08
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-10-08
AI Technical Summary
When classifying users, the existing user classification model mainly relies on the user's own information, and the data source is single, resulting in low classification accuracy.
通过获取用户数据和与用户相关的设备行为数据,基于用户标签体系进行特征抽取,构建用户画像数据,并将其输入用户分类模型以获得分类结果。 Integrate user classification results into the user tag system to optimize user classification model to improve classification accuracy.
Improve the accuracy of user classification, expand the label and data source of user portraits by integrating user and device behavior data, and continuously optimize the user classification model.
Smart Images

Figure CN117370834B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a user classification model optimization method, system, electronic device and storage medium. Background Art
[0002] Operators have hundreds of millions of users, and different users have different consumption levels and consumption preferences. In order to improve the recommendation effect of products and services for users and increase user retention rate, enterprises will classify users through artificial intelligence technology and adopt different marketing strategies. However, most of the current user classification models only focus on the user's own information and only obtain information related to the classification target. The data source is single, resulting in low accuracy in user classification. Summary of the invention
[0003] The main purpose of the embodiments of the present application is to propose a user classification model optimization method, system, electronic device and storage medium, aiming to improve the accuracy of user classification.
[0004] To achieve the above purpose, an embodiment of the present application provides a method for optimizing a user classification model, comprising the following steps:
[0005] Obtain user data and device behavior data related to the user;
[0006] Extracting features from the user data and the device behavior data based on a user tag system to obtain tag feature data;
[0007] Determine user portrait data according to the tag feature data;
[0008] Inputting the user portrait data into a user classification model to obtain a user classification result;
[0009] The user classification result is integrated into the user label system, and the user classification model is optimized based on the integrated user label system.
[0010] In some embodiments, the user classification model optimization method further includes the following steps:
[0011] Performing data exploratory analysis on the device behavior data;
[0012] According to the analysis results of the data exploratory analysis, the device behavior data is cleaned, missing values are filled, and outliers are removed.
[0013] In some embodiments, inputting the user portrait data into a user classification model to obtain a user classification result comprises the following steps:
[0014] Input the user portrait data into the dimension feature extraction module to obtain multiple dimension feature values;
[0015] Input the dimension feature values into the feature quantization module to obtain quantization feature values;
[0016] Input multiple quantization feature values into the classification module to obtain a user classification result.
[0017] In some embodiments, the step of inputting the dimension feature values into the feature quantization module to obtain quantization feature values includes the following steps:
[0018] Determine the numerical interval where the dimension feature value is located;
[0019] Determine the quantization feature value corresponding to the numerical interval as the quantization feature value of the dimension feature value;
[0020] Among them, the interval threshold of the numerical interval is obtained through the following steps:
[0021] Use the equal-number discretization method to divide each dimension sample, and determine the first threshold of multiple numerical intervals according to the division result;
[0022] Use the density clustering method to cluster each dimension sample, and determine the second threshold of multiple data intervals according to the clustering result;
[0023] Perform weighted calculation on the first threshold and the second threshold of each data interval to obtain the interval threshold.
[0024] In some embodiments, the step of inputting multiple quantization feature values into the classification module to obtain a user classification result includes the following steps:
[0025] Obtain the dimension weight of each dimension;
[0026] Perform weighted calculation according to the dimension weights of multiple dimensions and the corresponding dimension feature values to obtain a user comprehensive score;
[0027] Determine the user classification result according to the user comprehensive score;
[0028] Among them, the dimension weight is obtained through the following steps:
[0029] Perform temporal stability analysis on the overall characteristics of each dimension sample respectively, and adjust the dimension weight according to the stability analysis result;
[0030] Perform balance analysis on the data distribution of each dimension sample, and adjust the dimension weight according to the balance analysis result.
[0031] In some embodiments, the step of inputting the user portrait data into the user classification model to obtain a user classification result further includes the following steps:
[0032] Determine a target eigenvalue from multiple dimensional eigenvalues;
[0033] When the target eigenvalue is greater than a preset value and the quantization eigenvalue corresponding to the target eigenvalue is less than a lower limit value, modify the quantization eigenvalue corresponding to the target eigenvalue to the lower limit value.
[0034] In some embodiments, the step of inputting the user portrait data into the user classification model to obtain a user classification result further includes the following steps:
[0035] Perform periodic numerical judgment on each dimensional eigenvalue of the user;
[0036] When there is a dimensional eigenvalue whose discrete degree of periodic numerical distribution is greater than a preset discrete degree, limit the quantization eigenvalue corresponding to the dimensional eigenvalue to be less than an upper limit value.
[0037] To achieve the above object, another aspect of the embodiments of the present application proposes a user classification model optimization system, including:
[0038] A first module, configured to obtain user data and device behavior data related to the user;
[0039] A second module, configured to perform feature extraction on the user data and the device behavior data based on a user label system to obtain label feature data;
[0040] A third module, configured to determine user portrait data according to the label feature data;
[0041] A fourth module, configured to input the user portrait data into a user classification model to obtain a user classification result;
[0042] A fifth module, configured to fuse the user classification result into the user label system and optimize the user classification model based on the fused user label system.
[0043] To achieve the above object, another aspect of the embodiments of the present application proposes an electronic device, which includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the user classification model optimization method described above.
[0044] To achieve the above object, another aspect of the embodiments of the present application proposes a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the user classification model optimization method described above.
[0045] The user classification model optimization method, system, electronic device and storage medium proposed in this application obtain user data and device behavior data related to the user, extract features from the user data and device behavior data based on the user label system to obtain label feature data, and then determine user portrait data according to the label feature data. The user portrait data incorporates the behavior data of devices related to the user, improving the accuracy of the user portrait data and thus the accuracy of subsequent user classification. The user portrait data is input into the user classification model to obtain a user classification result, and then the user classification result is integrated into the user label system, and the user classification model is optimized based on the integrated user label system. By continuously integrating the user label system, the labels and data sources of the user portrait are expanded, and then the user classification model is optimized, thereby improving the accuracy of user classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flowchart of the user classification model optimization method provided by an embodiment of this application;
[0047] Figure 2 is a schematic diagram of the user label system provided by an embodiment of this application;
[0048] Figure 3 is a schematic diagram of the user classification model optimization process provided by an embodiment of this application;
[0049] Figure 4 is a flowchart of the user classification model optimization method provided by another embodiment of this application;
[0050] Figure 5 is Figure 1 a flowchart of step S104 in
[0051] Figure 6 is Figure 5 a flowchart of step S302 in
[0052] Figure 7 is Figure 5 a flowchart of step S303 in
[0053] Figure 8 is Figure 1 a flowchart of step S104 in
[0054] Figure 9 is Figure 1 a flowchart of step S104 in
[0055] Figure 10 is a visualization schematic diagram of the quantization feature values of the feature values in each dimension provided by an embodiment of the present invention;
[0056] Figure 11 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present application. Specific embodiments
[0057] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0058] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar users and do not have to be used to describe a specific order or sequence.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0060] First, several terms involved in the present application are analyzed:
[0061] Artificial intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.
[0062] User Profile refers to a series of structured or unstructured data descriptions of user attributes. User attributes usually include basic attributes, social attributes, behavioral attributes, psychological attributes, etc. The construction of User Profile usually follows principles such as data authenticity, uniqueness, dynamics, and applicability, and is the core process of big data analysis services.
[0063] The embodiments of the present application provide a method, a system, an electronic device, and a storage medium for optimizing a user classification model, aiming to improve the accuracy of user classification and grading.
[0064] The method, system, electronic device, and storage medium for optimizing the user classification model provided by the embodiments of the present application will be specifically described through the following embodiments. First, the method for optimizing the user classification model in the embodiments of the present application will be described.
[0065] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0066] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0067] The method for optimizing the user classification model provided by the embodiments of the present application relates to the field of artificial intelligence technology. The method for optimizing the user classification model provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the method for optimizing the user classification model, etc., but is not limited to the above forms.
[0068] This application can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, users, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0069] It should be noted that in each specific embodiment of this application, when it comes to relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, user permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain sensitive personal information of users, separate permission or separate consent from users will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the separate permission or separate consent of users, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.
[0070] Figure 1 It is an optional flowchart of the user classification model optimization method provided by the embodiments of this application. Figure 1 The method in [it] may include but is not limited to steps S101 to S106.
[0071] Step S101, obtain user data and device behavior data related to the user;
[0072] Step S102, perform feature extraction on the user data and device behavior data based on the user label system to obtain label feature data;
[0073] Step S103, determine user portrait data according to the label feature data;
[0074] Step S104, input the user portrait data into the user classification model to obtain a user classification result;
[0075] Step S105, fuse the user classification result into the user label system, and optimize the user classification model based on the fused user label system.
[0076] Steps S101 to S105 illustrated in the embodiments of the present application obtain user data and device behavior data related to the user, extract features from the user data and device behavior data based on the user label system to obtain label feature data, and then determine user portrait data according to the label feature data. The user portrait data integrates the behavior data of devices related to the user, improving the accuracy of the user portrait data and further improving the accuracy of subsequent user classification. The user portrait data is input into the user classification model to obtain a user classification result, and then the user classification result is fused into the user label system, and the user classification model is optimized based on the fused user label system. The label and data sources of the user portrait are expanded through the continuously fused user label system, and then the user classification model is optimized to improve the accuracy of user classification.
[0077] In step S101 of some embodiments, current Internet products contain user behavior and device behavior data. However, since the user and the device are not associated, the analysis dimension is too single. For example, the prediction of the Internet product value of a user not only depends on the information related to the user account, but also is related to the behavior information of the user device. Therefore, in addition to directly collecting user data generated by the user account through the application degree, it is also necessary to query the device ID related to the user and obtain device behavior data through the device ID. In addition, since the data provided by the product side may have abnormal data such as duplicates, errors, and illegality, the log data is preprocessed before being stored in the database.
[0078] Furthermore, the cleaned log data can be stored in a Hive table for data modeling development. In addition, some configuration resource data can be written into MySQL, mainly including mapping data such as label mapping, behavior coding, and their names. The data in Clickhouse, Hbase, and Elasticsearch is mainly used for developing visualization and microservice output.
[0079] In step S102 of some embodiments, refer to Figure 2, the user tag system contains multiple pre-defined user tags. During the definition of the user tag system, since the original values of the feature data are generated by users and each user has different behavior habits, the feature data results of many tags will be too discrete. Exemplarily, hours are divided into late night, working period, off-work period, and prime evening period; a week is divided into working days and non-working days; cities are divided into first-tier, new first-tier, and second-tier, etc. These discrete data are somewhat difficult to analyze the characteristics of user groups. Not only is it difficult for the model to converge, but also the interpretation of user characteristics is cumbersome. Therefore, it is necessary to compare the key data and dimensions from various collected data and integrate them into the existing user tag system. In addition, the traditional tag system is only built at the user mobile phone number level and lacks tags for user information such as broadband and devices. Therefore, it is necessary to add tags such as devices, broadband, and uid in the business logs to the tag system. For example, if a user is not active on the APP side but their device is active, the user is actually valuable, rather than simply judging the activity of a single device.
[0080] After obtaining the user data and device behavior data of multiple users, based on the above user tag system, feature extraction is performed on the user data and device behavior data to obtain the tag feature data corresponding to each tag. The feature extraction process of user data and device behavior data can be implemented using natural language processing technology and information extraction technology based on deep neural networks.
[0081] In step S103 of some embodiments, a user portrait of each user is constructed through the extracted tag feature data. The user portrait usually contains a series of user information, such as basic information: name, gender, age, occupation, etc.; behavior habits: time, frequency, and method of using the product, etc.; interests and hobbies: types of content liked, activities, entertainment methods, etc. The user portrait is a series of structured or unstructured data descriptions of user attributes, and user attributes usually include basic attributes, social attributes, behavioral attributes, psychological attributes, etc. In this embodiment, multi-dimensional portraits of the user's behavior characteristics are made for hierarchical classification to ensure the rationality of the data results.
[0082] In steps S104 and S105 of some embodiments, after obtaining the user portrait data, the user portrait data is input into the user classification model to obtain the user classification result. The user classification model can be designed according to the actual business classification requirements. For example, the user classification model can be the degree of contribution to a certain product, which are divided into high, medium, and low levels, etc. Through the user classification model, different user groups can be divided according to the behavioral feature portraits of different users, and then the differential value brought by different users can be quantified, so as to give full play to the value of each level of users and achieve the product goal.
[0083] The user classification model can be trained using a deep neural network. Refer to Figure 3 , the construction and optimization process of the user classification model is as follows: Collect a large amount of user log data, including user data such as user number behavior and user account behavior, and device behavior data. Preprocess and store the collected data in a database. Conduct exploratory data analysis (EDA) on the behavior data in the database, extract relevant feature data such as behavior types, active days, maximum behavior times, and non-login behaviors, and update them to the label system at the same time. Build user portraits for each user based on the feature data related to the user label system. Design and initialize the user classification model, that is, the user hierarchical classification model. Input multiple user portrait data and classification labels as training samples into the user classification model to build the user classification model. After the user classification model is trained, after evaluating and optimizing the classification results output by the user classification model, it can be launched on the platform for product operators to use. During the use of the model, the user classification results output by the model can be updated to the user label system, and for classification problems, the parameters of the model can be automatically adjusted and optimized or secondary modeling can be performed according to the latest label feature data.
[0084] In some embodiments, according to the evaluation results of the model, when optimizing parameters or performing secondary modeling for problems, if the output results of the model do not conform to logic, the weight value of the feature dimension in the model can be adjusted or re-modeled by combining user feature data.
[0085] In some embodiments, the result output by the user classification model can be the user value. The reason for re-integrating the user value into the label system as a label is that the label data is dynamically updated, and the historical user ratings may not be accurate enough due to the update and iteration of the product. Therefore, the user value evaluation model also needs to be changed accordingly according to the update of the label, and at the same time, it is iterated into the label system as a label.
[0086] In some embodiments, when the model is launched, it can predict the value of users. Specifically, different assessment periods can be set according to the attributes of different products to periodically assess each user of each product, and the data results are written into the database. The data written into HBase is used for the output of basic microservices. For example, if the product party provides a user ID, this embodiment can quickly feedback the corresponding rating information and some feature information of the user to the product party according to the primary key search; considering that the storage cost of Elasticsearch is relatively high, but its multi-dimensional search performance for data is good, therefore, the data written into Elasticsearch is used for the microservice that provides the export of user grouping data, and this part of the fields will be more targeted than the data written into HBase.
[0087] It should be noted that considering that the assessment period may differ from the cycle of user behavior habits. Since the behavior habits of each user are different, it is difficult for the model to cover the behavior habit cycles of every user. Therefore, the embodiments of the present invention propose a data smoothing mechanism, which can incorporate the scoring values of the user in the previous two cycles when the model evaluates the user score to determine the final user value. This mechanism is an important mechanism for improving the rationality of the model.
[0088] Please refer to Figure 4 , in some embodiments, the method for optimizing the user classification model according to the embodiments of the present invention includes, but is not limited to, steps S201 and S202:
[0089] Step S201, perform exploratory data analysis on the device behavior data;
[0090] Step S202, according to the analysis results of the exploratory data analysis, for the device behavior data.
[0091] In this embodiment, after obtaining the device behavior data, it is necessary to first perform exploratory data analysis to preliminarily analyze the mutual relationship between variables and the relationship between variables and predicted values, and perform preliminary processing on the data, such as: handling data anomalies and missing values, etc., so that the structure and characteristics of the data set are more reliable for the next prediction problem. EDA refers to a data analysis method that explores existing data (especially the original data obtained from surveys or observations) with as few prior assumptions as possible, and explores the structure and laws of the data through means such as graphing, tabulation, equation fitting, and calculating characteristic quantities.
[0092] In the log data, the product side will provide different device behavior data, and there are many dimensions, but not every log and dimension data can provide value. Therefore, through the EDA exploration and analysis of the behavior data, reasonable characteristic data output is formed, such as analyzing the discreteness, missing situation, trend, etc. of the analysis indicators, and then performing data cleaning, missing value filling, and outlier removal on the analysis results, etc., to ensure the rationality of the characteristic data.
[0093] Please refer to Figure 5 , in some embodiments, in step S104, the step of inputting the user portrait data into the user classification model to obtain the user classification result may include, but is not limited to, steps S301 to S303:
[0094] Step S301, input the user portrait data into the dimension feature extraction module to obtain multiple dimension feature values;
[0095] Step S302, input the dimension feature values into the feature quantization module to obtain quantized feature values;
[0096] Step S303: Input multiple quantified feature values into the classification module to obtain the user classification result.
[0097] In some embodiments, input the user portrait data into the dimension feature extraction module to extract the dimension feature values of multiple required analysis dimensions. Exemplarily, referring to multiple dimensions including but not limited to whether there is a core behavior, the actual situation of the behavior, the total number of behaviors, the number of behavior types, the number of active days, the maximum number of behaviors per day, the types of active time points, and the number of days without logging in. Among them, the dimension of whether there is a core behavior measures the basic requirements of high-quality users, and the dimensions of the actual situation of the behavior, the total number of behaviors, the number of behavior types, the number of active days, the maximum number of behaviors per day, and the types of active time points are all for the behavior consumption of the product, and the number of days without logging in is a reverse dimension. The dimension feature value is the feature vector representation of the dimension attribute value. For example, the feature vector representation of the total number of behaviors being 5 times is the dimension feature value.
[0098] In some embodiments, the feature quantification module is mainly used to quantify and score each dimension feature value. Since there are many independent characteristic behaviors in user behavior, such as the number of active days, consumption value, consecutive number of days without logging in, etc. The score quantization between dimensions is to reasonably describe the level of the user's behavior in each dimension. In this embodiment, a combination of mathematical methods such as density clustering, population discretization, and data truncation method is used to handle the problem of score quantization between dimensions.
[0099] In some embodiments, the classification module is a classifier used to comprehensively evaluate the scores of each dimension of the user. In the comprehensive evaluation process, a weighted algorithm can be used to synthesize the quantified feature values of each dimension, and the weight values of each dimension can be determined by methods such as expert scoring method, entropy value method, fuzzy evaluation, data envelopment, etc., or the expert scoring method and the dynamic weight adjustment method based on the evaluation results can be used to output the final score result of the user.
[0100] Please refer to Figure 6 , in some embodiments, in step S302, the step of inputting the dimension feature value into the feature quantification module to obtain the quantified feature value may include but not limited to steps S401 to S402:
[0101] Step S401: Determine the numerical interval where the dimension feature value is located;
[0102] Step S402: Determine the quantified feature value corresponding to the numerical interval as the quantified feature value of the dimension feature value;
[0103] Among them, the interval threshold of the numerical interval is obtained through the following steps:
[0104] Use the equal-population discretization method to divide each dimension sample, and determine the first threshold of multiple numerical intervals according to the division result;
[0105] Use the density clustering method to cluster the samples of each dimension, and determine the second threshold of multiple data intervals according to the clustering results;
[0106] Perform weighted calculation on the first threshold and the second threshold of each data interval to obtain the interval threshold.
[0107] In this embodiment, the quantization eigenvalue represents the scoring result of the dimensional eigenvalue of the user. By defining the score of the numerical interval of each dimension, after obtaining the dimensional eigenvalue, the corresponding score is determined by judging the numerical interval where the dimensional eigenvalue is located.
[0108] Furthermore, this embodiment provides multiple strategies to determine the score of the numerical interval of the dimension. Taking a user who has consumed 71 integral yuan in this product within one month as an example, how many points can this user get in the dimension of integral consumption, that is, what is the quantization eigenvalue, specifically as follows:
[0109] Strategy 1, confirm by the expert scoring method, that is, score according to the expert threshold, as shown in Table 1:
[0110] Table 1 Expert Scoring Table
[0111] Grade Division Above 80 points 60~80 40~60 20~40 0~20 Total Times x>=150 70<=x<150 30<=x<70 10<=x<30 x<10
[0112] Strategy 2, determine the score of the numerical interval through the statistical analysis algorithm, specifically as follows:
[0113] Intercept 95% of the data in the middle of the dimension sample as the evaluation data, mainly to reduce the influence of extreme values.
[0114] Use the equal-number discretization method to divide the dimension sample, and determine the first threshold of multiple numerical intervals according to the division results, that is, determine the first threshold according to the equal proportion of the number of people. For example, it is stipulated that 25% of the users score below 25 points, and 25% - 50% of the users score between 25 - 50 points. Then the upper limit and lower limit of the dimensional eigenvalues of all users below 25 points in 25% of the samples are the interval thresholds, that is, the two endpoint values of the numerical interval. This method may cause users with the same behavior value or similar behavior values to be in two intervals, resulting in large score differences.
[0115] Use the density clustering method to cluster the samples of each dimension, and determine the second threshold of multiple data intervals according to the clustering results, that is, preset the number of numerical intervals, and use the clustering algorithm to divide users with similar scores into the same interval. This method can ensure that users with similar dimensional eigenvalues are in the same category to the greatest extent, but this method is easily affected by data skewness and may cause 80% or more of the users to be in the same interval.
[0116] Construct a feature quantization module for user dimension scores based on the combination of the above three algorithms. The combination strategy is as follows: on the premise of data interception, the first threshold and the second threshold of each interval with the same number of intervals are respectively output through the equal-number discretization method and the density clustering method, and then the first threshold and the second threshold are weighted and calculated to obtain the interval threshold of each interval. Multiple numerical intervals of a dimension can be determined through the interval threshold of each interval. Among them, the weights in the weighting process need to be measured in combination with real data.
[0117] Please refer to Figure 7 , in some embodiments, in step S303, the step of inputting multiple quantization feature values into the classification module to obtain the user classification result may further include, but is not limited to, steps S501 to S503:
[0118] Step S501, obtain the dimension weight of each dimension;
[0119] Step S502, perform weighted calculation according to the dimension weights of multiple dimensions and the corresponding dimension feature values to obtain the user comprehensive score;
[0120] Step S503, determine the user classification result according to the user comprehensive score;
[0121] Among them, the dimension weight is obtained through the following steps:
[0122] Perform a temporal stability analysis on the overall characteristics of each dimension sample respectively, and adjust the dimension weight according to the stability analysis result;
[0123] Perform a balance analysis on the data distribution of each dimension sample, and adjust the dimension weight according to the balance analysis result.
[0124] In this embodiment, a weight value is assigned to each dimension in the classification module. Exemplarily, the weight values of whether there is a core behavior, the real situation of the behavior, the total number of behaviors, the number of behavior types, the number of active days, the maximum number of behaviors per day, the types of active time points, the number of days without logging in, and the user consumption value are 0.07, 0.3, 0.105, 0.14, 0.14, 0.07, 0.035, 0.14 respectively, and the sum of the weight values of all dimensions is 1.
[0125] The weight assignment strategy can be to first use the expert scoring method to determine the initial weight values of different dimensions (the sum of the weights is 1), and then construct a management platform to correct the weight values based on the feedback of the model evaluation results.
[0126] The weight optimization strategy is as follows:
[0127] Based on the stability tuning strategy, a periodic evaluation is conducted on the overall scores of each dimension. For dimensions with stable scores, their weights will relatively increase, while for indicators with large fluctuations, their weights will relatively decrease.
[0128] Based on the balance tuning strategy, the relative distribution of users with the overall score of a certain dimension is relatively reasonable. For dimensions with reasonable score distributions, their weights will relatively increase, while for dimensions with unreasonable score distributions, their weights will relatively decrease.
[0129] Please refer to Figure 8 , in some embodiments, in step S104, the step of inputting the user portrait data into the user classification model to obtain the user classification result may further include, but is not limited to, steps S601 to S602:
[0130] Step S601, determining the target eigenvalue from multiple dimension eigenvalues;
[0131] Step S602, when the target eigenvalue is greater than the preset value and the quantization eigenvalue corresponding to the target eigenvalue is less than the lower limit value, modify the quantization eigenvalue corresponding to the target eigenvalue to the lower limit value.
[0132] In this embodiment, considering that the data operation situations of different products are different, there may be some particularly important dimensions. For example, for recharge-based APPs, the product side may be more concerned about the user's recharge situation because the user's recharge behavior is often periodic and the APP will be used only when there is a need. Therefore, when designing the user classification model, important dimensions can be defined, and the corresponding target eigenvalues can be identified from multiple dimension eigenvalues through the defined dimensions. When the target eigenvalue is greater than the preset value and the quantization eigenvalue corresponding to the target eigenvalue is less than the lower limit value, modify the quantization eigenvalue corresponding to the target eigenvalue to the lower limit value, so as to ensure the lower limit of the quantization eigenvalue of the target eigenvalue. For example, when the recharge amount feature of a user within the assessment period reaches a certain value, the lower limit of the user's score is ensured. The one-vote veto mechanism adopted by the model in the above process is a flexible dimension and an important mechanism to improve the rationality of the model.
[0133] Please refer to Figure 9 , in some embodiments, in step S104, the step of inputting the user portrait data into the user classification model to obtain the user classification result may further include, but is not limited to, steps S701 to S702:
[0134] Step S701, making a periodic numerical judgment on each dimension eigenvalue of the user;
[0135] Step S702, when there is a dimension eigenvalue whose discrete degree of the periodic numerical distribution is greater than the preset discrete degree, restrict the quantization eigenvalue corresponding to the dimension eigenvalue to be less than the upper limit value.
[0136] In this embodiment, considering that when rating users, it is often a periodic assessment, and this cycle value is often greater than 1 day. However, the model's result should not return a high score for a user throughout the cycle just because the user has a particularly high behavior score on only one day during the entire cycle. Exemplarily, if a user has been using only one behavior, the score should not be too high because the user may only have abnormal consumption behavior during the activity period. Therefore, this embodiment uses a saturation mechanism for the scores of the metric dimensions in the user classification model to perform periodic numerical judgments on each dimensional feature value of the user. When there is a dimensional feature value with a discrete degree of periodic numerical distribution greater than the preset discrete degree, it indicates that the dimensional feature value of the user at some time is too different from that at other times. Then, the quantization feature value corresponding to the dimensional feature value is restricted to be less than the upper limit value, thereby ensuring the upper limit of the user's score. The discrete degree of the periodic numerical distribution can be measured by variance or standard deviation. This mechanism is one of the mechanisms to ensure the rationality of the model results and can be flexibly used in combination with actual data.
[0137] According to some embodiments of the present invention, the embodiments of the present invention have the following beneficial effects:
[0138] The embodiments of the present invention integrate the label system to construct a user portrait, and then optimize the user classification model according to the user portrait data, which can efficiently handle the adjustment of parameters and secondary modeling, and the accuracy of the user classification result is high.
[0139] The embodiments of the present invention construct a variety of mechanism strategies during the user hierarchical classification modeling, including a score saturation mechanism, a one-vote veto mechanism, a data smoothing mechanism, etc., which improves the rationality of the model.
[0140] The embodiments of the present invention integrate multiple algorithms such as the truncation method, the equal-number discretization method, and the density clustering method during the process of quantifying the dimensional scores, which improves the rationality of the dimensional score quantization. In addition, referring to Figure 10 , the user classification model can visually output the quantization scores of each dimensional feature value of the user, which is convenient for analyzing the user.
[0141] The embodiments of the present invention propose a stability-based tuning strategy and a balance-based tuning strategy during the weight optimization stage of the classifier, which effectively improves the weight rationality between different dimensions.
[0142] The embodiments of the present application also provide a user classification model optimization system, including:
[0143] The first module is used to obtain user data and device behavior data related to the user;
[0144] The second module is used to perform feature extraction on the user data and device behavior data based on the user label system to obtain label feature data;
[0145] The third module is used to determine user portrait data according to label feature data;
[0146] The fourth module is used to input the user portrait data into a user classification model to obtain a user classification result;
[0147] The fifth module is used to integrate the user classification result into the user label system and optimize the user classification model based on the integrated user label system.
[0148] It can be understood that the content in the above embodiments of the user classification model optimization method is applicable to the embodiments of this system. The functions specifically implemented by the embodiments of this system are the same as those of the above embodiments of the user classification model optimization method, and the beneficial effects achieved are also the same as those of the above embodiments of the user classification model optimization method.
[0149] An embodiment of this application also provides an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the above user classification model optimization method.
[0150] Please refer to Figure 11 , Figure 11 which illustrates the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0151] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this application;
[0152] A memory 902, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the user classification model optimization method of the embodiments of this application;
[0153] An input / output interface 903, which is used to implement information input and output;
[0154] A communication interface 904 for implementing communication and interaction between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0155] A bus 905 for transmitting information between various components of the device (such as a processor 901, a memory 902, an input / output interface 903, and a communication interface 904);
[0156] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 achieve communication connections with each other inside the device through the bus 905.
[0157] The embodiment of the present application also provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned user classification model optimization method.
[0158] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0159] The user classification model optimization method, system, electronic device, and storage medium proposed in the present application obtain user data and device behavior data related to the user, extract feature data from the user data and device behavior data based on the user tag system to obtain tag feature data, and then determine user portrait data according to the tag feature data. The user portrait data integrates the behavior data of devices related to the user, improves the accuracy of the user portrait data, and further improves the accuracy of subsequent user classification. Input the user portrait data into the user classification model to obtain a user classification result, and then fuse the user classification result into the user tag system, and optimize the user classification model based on the fused user tag system. Expand the tags and data sources of the user portrait through the continuously fused user tag system, and then optimize the user classification model to improve the accuracy of user classification.
[0160] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0161] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.
[0162] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0163] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0164] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar users, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0165] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the associated relationship of associated users, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated users before and after are an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0166] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0167] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0168] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0169] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0170] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.
Claims
1. A user classification model optimization method, It is characterized in that The following steps are involved: Acquire user data and device behavior data related to the user, wherein the device behavior data is information acquired based on the user's device ID; Extracting features from the user data and the device behavior data based on a user tag system to obtain tag feature data, wherein the user tag system includes a plurality of defined user tags; Determine user portrait data by performing structured or unstructured data description on user attributes according to the tag feature data; Inputting the user portrait data into a user classification model to obtain a user classification result; The user classification result is integrated into the user tag system as a user tag, and the user classification model is optimized based on tag feature data of the new user tag in the integrated user tag system.
2. The user classification model optimization method according to claim 1, It is characterized in that The user classification model optimization method also includes the following steps: Performing data exploratory analysis on the device behavior data; According to the analysis results of the data exploratory analysis, the device behavior data is cleaned, missing values are filled, and outliers are removed.
3. The user classification model optimization method according to claim 1, It is characterized in that The step of inputting the user portrait data into the user classification model to obtain the user classification result comprises the following steps: Inputting the user portrait data into a dimensional feature extraction module to obtain multiple dimensional feature values; Inputting the dimension feature value into a feature quantization module to obtain a quantized feature value; The plurality of quantitative feature values are input into a classification module to obtain a user classification result.
4. The user classification model optimization method according to claim 3, It is characterized in that The step of inputting the dimension feature value into a feature quantization module to obtain a quantized feature value comprises the following steps: Determine the numerical range of the dimension characteristic value; Determine the quantized feature value corresponding to the numerical interval as the quantized feature value of the dimensional feature value; The interval threshold of the numerical interval is obtained by the following steps: The samples of each dimension are divided by using the equal number of discretization methods, and the first thresholds of multiple value intervals are determined according to the division results; Clustering each dimension sample using a density clustering method, and determining a second threshold value of multiple data intervals according to the clustering result; A weighted calculation is performed on the first threshold and the second threshold of each data interval to obtain an interval threshold.
5. The user classification model optimization method according to claim 4, It is characterized in that Inputting the plurality of quantitative feature values into a classification module to obtain a user classification result comprises the following steps: Get the dimension weight of each dimension; Perform weighted calculation based on the dimension weights of multiple dimensions and the corresponding dimension feature values to obtain the user's comprehensive score; Determine a user classification result according to the user comprehensive score; The dimension weights are obtained by the following steps: Conduct temporal stability analysis on the overall characteristics of each dimension sample, and adjust the dimension weight according to the stability analysis results; Perform a balance analysis on the data distribution of each dimension sample and adjust the dimension weight according to the balance analysis results.
6. The method for optimizing the user classification model according to claim 3, wherein, the step of inputting the user portrait data into the user classification model to obtain the user classification result further includes the following steps: determining a target feature value from multiple dimensional feature values; when the target feature value is greater than a preset value and the quantization feature value corresponding to the target feature value is less than a lower limit value, modifying the quantization feature value corresponding to the target feature value to the lower limit value.
7. The method for optimizing the user classification model according to claim 3, wherein, the step of inputting the user portrait data into the user classification model to obtain the user classification result further includes the following steps: performing periodic numerical judgment on each dimensional feature value of the user; when there is a dimensional feature value whose discrete degree of the periodic numerical distribution is greater than a preset discrete degree, restricting the quantization feature value corresponding to the dimensional feature value to be less than an upper limit value.
8. A user classification model optimization system, wherein, it includes: a first module, configured to obtain user data and device behavior data related to the user, wherein the device behavior data is information obtained based on the user's device ID; a second module, configured to perform feature extraction on the user data and the device behavior data based on the user label system to obtain label feature data, wherein the user label system includes multiple defined user labels; a third module, configured to determine user portrait data by performing structured or unstructured data description on user attributes according to the label feature data; a fourth module, configured to input the user portrait data into the user classification model to obtain the user classification result; a fifth module, configured to fuse the user classification result into the user label system as a user label, and optimize the user classification model based on the label feature data of the new user label in the fused user label system.
9. An electronic device, wherein, the electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it realizes the steps of the method for optimizing the user classification model according to any one of claims 1 to 7.
10. A storage medium, which is a computer-readable storage medium for computer-readable storage, wherein, the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to realize the steps of the method for optimizing the user classification model according to any one of claims 1 to 7.
Citation Information
Patent Citations
User portraying method based on mass cross-screen behavior data
CN106980663A
Financial user classification method based on user portrait model
CN110490729A