User data processing method and device and nonvolatile computer readable storage medium
By binning user data and determining concept drift indicators, the problem of inaccurate user evaluation when concept drift occurs in data distribution is solved, and the effect of user data processing is improved.
Patent Information
- Application Number
- CN202311584929.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art cannot accurately evaluate users when the data distribution concept drifts, resulting in a decrease in the effectiveness of user data processing.
By binning the user data according to the binning threshold, multiple data binning is acquired, and the conceptual drift indicators of user characteristics are determined based on the difference between the first data bin and the corresponding second data binning, so as to evaluate the user.
In the case of concept drift in data distribution, users can be accurately evaluated and the effectiveness of user data processing can be improved.
Smart Images

Figure CN120045596A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to a user data processing method, a user data processing device, and a non-volatile computer-readable storage medium. Background Art
[0002] In many application scenarios, the samples of user data are often affected by factors such as the external environment and policy adjustments, resulting in changes in the data distribution of the samples (i.e., data drift), and even changes in the mapping relationship between the data distribution and the target variable used for user evaluation (i.e., concept drift). In this case, the effect of the model used for user analysis, classification, and evaluation may fluctuate or even decay.
[0003] In the related art, the user data processing methods for dealing with the changed data distribution include PSI (Population Stability Index), KS (Kolmogorov-Smirnov) test, Kullback-Leibler divergence, Euclidean distance, Wasserstein distance, etc. Summary of the Invention
[0004] The inventors of the present disclosure found the following problems in the above related art: It is impossible to accurately evaluate users in the case of concept drift in the data distribution, resulting in a reduction in the effect of user data processing.
[0005] In view of this, the present disclosure proposes a user data processing technical solution, which can accurately evaluate users in the case of concept drift in the data distribution, thereby improving the effect of user data processing.
[0006] According to some embodiments of the present disclosure, there is provided a user data processing method, including: binning the user data in the first user data set according to a binning threshold to obtain a plurality of first data bins; binning the second data set according to the binning threshold to obtain a plurality of second data bins, where the first user data set and the second user data set correspond to the same user feature; determining a concept drift index of the user feature according to the difference between the first data bin and its corresponding second data bin; and evaluating the user according to the concept drift index.
[0007] In some embodiments, determining a concept drift index of the user feature according to the difference between the first data bin and its corresponding second data bin includes: determining the concept drift index according to the difference between the data distribution of the first data bin and the data distribution of the second data bin.
[0008] In some embodiments, the difference includes the WOE (Weight Of Evidence) difference.
[0009] In some embodiments, determining a concept drift metric based on the difference between a first data bin and its corresponding second data bin includes: determining weights based on the proportion of the data volume of the first data bin in the first dataset and the proportion of the data volume of the second data bin in the second dataset; weighting the difference using the weights; and determining a concept drift metric based on the weighted difference.
[0010] In some embodiments, determining weights based on the proportion of the data volume of the first data bin in the first dataset and the proportion of the data volume of the second data bin in the second dataset includes: determining a first weight based on the proportion of the data volume of the first data bin in the first dataset; determining a second weight based on the proportion of the data volume of the second data bin in the second dataset; and determining the weights based on the weighted average of the first weight and the second weight.
[0011] In some embodiments, binning the user data in a first user dataset to obtain a plurality of first data bins includes: dividing the first user dataset into a plurality of first data subsets; determining candidate binning thresholds according to a first binning condition for dividing the first data subsets into a plurality of second data subsets; when each of the plurality of second data subsets satisfies a second binning condition, determining the plurality of second data subsets that satisfy the second binning condition as a plurality of first data bins, and determining the candidate binning thresholds corresponding to the plurality of second data subsets that satisfy the second binning condition as binning thresholds.
[0012] In some embodiments, determining candidate binning thresholds according to a first binning condition includes: when there is a second data subset among the plurality of second data subsets that does not satisfy the second binning condition, adjusting the first binning condition; re-determining candidate binning thresholds according to the adjusted first binning condition; re-dividing the first data subsets into a plurality of second data subsets according to the re-determined candidate binning thresholds; repeating the above adjustment step, determination step, and division step until each of the re-divided plurality of second data subsets satisfies the second binning condition.
[0013] In some embodiments, the first binning condition includes at least one of the proportion of the data volume of the first data subset in the first dataset being greater than a first threshold or the number of data bins desired; the second binning condition includes that the relationship between the user data in the second data subset and the target variable for user evaluation conforms to a preset relationship.
[0014] In some embodiments, the preset relationship is determined according to the user evaluation service to be performed.
[0015] In some embodiments, the preset relationship includes at least one of monotonically increasing, monotonically decreasing, first monotonically increasing and then monotonically decreasing, or first monotonically decreasing and then monotonically increasing.
[0016] In some embodiments, evaluating a user according to the concept drift metric includes: determining whether to select user features to evaluate the user according to the concept drift metric.
[0017] According to some other embodiments of the present disclosure, there is provided a user data processing device, including: a binning unit, configured to bin user data in a first user data set according to a binning threshold to obtain a plurality of first data bins, and bin a second data set according to the binning threshold to obtain a plurality of second data bins, where the first user data set and the second user data set correspond to the same user feature; a determining unit, configured to determine a concept drift metric of the user feature according to the difference between the first data bin and its corresponding second data bin; and an evaluating unit, configured to evaluate the user according to the concept drift metric.
[0018] In some embodiments, the determining unit determines the concept drift metric according to the difference between the data distribution of the first data bin and the data distribution of the second data bin.
[0019] In some embodiments, the difference includes the WOE difference.
[0020] In some embodiments, the determining unit determines a weight according to the proportion of the data volume of the first data bin in the first data set and the proportion of the data volume of the second data bin in the second data set, and weights the difference by using the weight; and determines the concept drift metric according to the weighted difference.
[0021] In some embodiments, the determining unit determines a first weight according to the proportion of the data volume of the first data bin in the first data set, determines a second weight according to the proportion of the data volume of the second data bin in the second data set, and determines the weight according to the weighted average of the first weight and the second weight.
[0022] In some embodiments, the binning unit divides the first user data set into a plurality of first data subsets, determines a candidate binning threshold according to a first binning condition for dividing the first data subsets into a plurality of second data subsets, and when each of the plurality of second data subsets satisfies a second binning condition, determines the plurality of second data subsets that satisfy the second binning condition as a plurality of first data bins, and determines the candidate binning threshold corresponding to the plurality of second data subsets that satisfy the second binning condition as the binning threshold.
[0023] In some embodiments, when there is a second data subset that does not meet the second binning condition among multiple second data subsets, the binning unit adjusts the first binning condition, determines a candidate binning threshold again according to the adjusted first binning condition, and re-partitions the first data subset into multiple second data subsets according to the re-determined candidate binning threshold. Repeat the above adjustment step, determination step, and partitioning step until each of the re-partitioned multiple second data subsets meets the second binning condition.
[0024] In some embodiments, the first binning condition includes at least one of the proportion of the data volume of the first data subset in the first data set being greater than the first threshold or the number of data bins to be obtained; the second binning condition includes that the relationship between the user data in the second data subset and the target variable for user evaluation conforms to a preset relationship.
[0025] In some embodiments, the preset relationship is determined according to the user evaluation service to be performed.
[0026] In some embodiments, the preset relationship includes at least one of monotonically increasing, monotonically decreasing, first monotonically increasing and then monotonically decreasing, or first monotonically decreasing and then monotonically increasing.
[0027] In some embodiments, the evaluation unit determines whether to select user features to evaluate the user according to the concept drift index.
[0028] According to still some embodiments of the present disclosure, there is provided a user data processing device, including: a memory; and a processor coupled to the memory, the processor being configured to execute the user data processing method in any of the above embodiments based on instructions stored in the memory device.
[0029] According to still some other embodiments of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the user data processing method in any of the above embodiments is implemented.
[0030] In the above embodiments, the same binning threshold is used to perform binning processing on two sets of data corresponding to the same user feature, and the concept drift index is determined according to the difference between the two sets of binning. In this way, it is possible to accurately evaluate the user in the case of concept drift in the data distribution, thereby improving the effect of user data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0032] Referring to the drawings, the present disclosure can be more clearly understood from the following detailed description:
[0033] Figure 1 Flowchart showing some embodiments of the user data processing method of the present disclosure;
[0034] Figure 2 Flowchart showing some embodiments of step 110;
[0035] Figure 3 Block diagram showing some embodiments of the user data processing apparatus of the present disclosure;
[0036] Figure 4 Block diagram showing some other embodiments of the user data processing apparatus of the present disclosure;
[0037] Figure 5 Block diagram showing still some other embodiments of the user data processing apparatus of the present disclosure. Detailed implementation manners
[0038] Now, various exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps set forth in these embodiments, numerical expressions and values do not limit the scope of the present disclosure.
[0039] Meanwhile, it should be understood that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationship.
[0040] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way serves as a limitation to the present disclosure and its application or use.
[0041] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods and devices should be regarded as part of the specification.
[0042] In all the examples shown and discussed herein, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.
[0043] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0044] The above-mentioned various user data processing methods for dealing with changes in data distribution can only measure the changes in data distribution and cannot measure the changes in the mapping relationship between user data and the target variables used by users for user evaluation.
[0045] Therefore, for the measurement of concept drift, the IV (Information Value) or KS index can be used to determine whether concept drift has occurred. However, judging only based on the changes in the IV or KS index cannot measure the degree of concept drift with a unified standard in different scenarios. In this way, the evaluation of users based on user data cannot be applied to different scenarios, resulting in poor effects in user data processing.
[0046] That is to say, methods such as PSI and distance calculation can be used to detect the stability and distribution differences of user characteristics in the user evaluation model. However, these methods can only evaluate whether the distribution of a single user characteristic has changed and do not consider the association between user characteristics and evaluation results. Therefore, these indicators are vulnerable to fluctuations in the user group. In the case of large fluctuations in the user group, although these indicators can reflect the data drift, they cannot determine whether concept drift has occurred.
[0047] Although the IV value after binning can be used to evaluate whether concept drift has occurred, uneven binning may lead to outliers. The method based on the IV value ignores the weights of binning (such as the sample proportion), resulting in poor evaluation effects.
[0048] In addition, the method of evaluating concept drift based on the decay of the IV value cannot determine a unified concept drift standard through the magnitude of the decay of the IV value. In this way, the evaluation of users based on user data cannot be applied to different scenarios, resulting in poor effects in user data processing.
[0049] To address the above technical problems, the present disclosure proposes a method for determining a concept drift detection index and a method for evaluating users. For example, this method can not only intuitively evaluate the change in the mapping relationship between user characteristics and the target variable used for user evaluation, but also adopt the analysis method of conditional inference trees to solve the technical problem of outliers caused by uneven binning in the method based on the IV value, thereby improving the effect of user data analysis.
[0050] For example, a method for evaluating the degree of concept drift proposed in the present disclosure bins two user data sets according to the same binning threshold, and evaluates the concept drift between the two user data sets based on the product of the woe difference of each data bin and the weight corresponding to each data bin. In this way, the concept drift index can intuitively reflect the change in the mapping relationship between user data and the target variable used for user evaluation. Moreover, in this method, the change degree of woe for each data bin is weighted, enabling the measurement of the degree of concept drift for different scenarios on a unified standard, thereby improving the effect of user data processing.
[0051] The technical solution of the present disclosure can analyze and solve the reasons for the attenuation of the model effect in the development work of an actual user classification model (such as the risk level of users, the interest groups they belong to, etc.); after the model is launched, it can monitor the user data and the effect of the model, thereby improving the processing effect of user data.
[0052] For example, the technical solution of the present disclosure can be implemented through the following embodiments.
[0053] Figure 1 The flowchart showing some embodiments of the user data processing method of the present disclosure.
[0054] As Figure 1 shown, in step 110, according to the binning threshold, the user data in the first user data set is binned to obtain a plurality of first data bins. The first user data set and the following second user data set correspond to the same user feature.
[0055] For example, user features may include various attribute data that can characterize user features, such as the age of the user, regional information, living habit information, consumption situation information, etc.
[0056] In some embodiments, the user database may include user data corresponding to various user features. Data that can be used for user evaluation can be extracted from the user data, such as removing primary keys such as user IDs and creation times in the user data and retaining the values therein; removing duplicate data rows in the database.
[0057] In some embodiments, for each user feature, the user data is divided into a first user data set (such as a training set) and a second user data set (such as a test set). For example, it can be divided according to the time information of the user data.
[0058] For example, for each user feature in the training set, methods such as conditional inference trees can be used for binning, and it is ensured that the relationship between the user data in each data bin after binning and the target variable for user evaluation (such as the risk trend of the user, the change in the user's health status, the change in the likelihood of the user performing a specified activity, etc.) conforms to a preset relationship. For example, as the value of the user data increases, the target variable for user evaluation monotonically increases or decreases.
[0059] In this way, the relationship between the binned user data and the target variable can conform to the logic of the user evaluation service, thereby improving the processing effect of user data.
[0060] In some embodiments, the first user data set can be binned using methods such as conditional tree monotonicity binning method, equal-frequency binning method, or equal-distance binning method.
[0061] For example, it can be throughFigure 2 The embodiments in
[0062] Figure 2 The flowchart showing some embodiments of step 110.
[0063] As Figure 2 shown, in step 1110, the first user data set is divided into a plurality of first data subsets.
[0064] In some embodiments, the first user data set can be divided into a plurality of first data subsets as a plurality of initial nodes. For example, the splitting criteria (i.e., the first binning condition) of the initial nodes can be initialized. The number of user data in the initial nodes can be determined according to the requirements of user evaluation, and each node corresponds to a data bin.
[0065] In step 1120, according to the first binning condition, a candidate binning threshold is determined for dividing the first data subset into a plurality of second data subsets.
[0066] In some embodiments, the first binning condition includes at least one of the proportion of the data volume of the first data subset in the first data set being greater than the first threshold or the number of data bins to be obtained.
[0067] For example, the first binning condition can include that the proportion of the data volume of the user data of each node in the total data volume of the corresponding user characteristics is not less than the threshold (such as 5% etc.), or the data volume of the user data of each node is greater than the minimum data volume.
[0068] In this way, it can be ensured that there is sufficient user data in each data bin, improving the accuracy of the concept drift index and thus the accuracy of user evaluation.
[0069] In some embodiments, splitting processing can be performed on the initial nodes. For example, according to the set splitting criteria of the initial nodes, the user data of the initial nodes is divided into two parts to split into two sub-nodes (i.e., two second data subsets). The number of user data of each sub-node must be greater than the minimum data volume.
[0070] In step 1130, it is judged that each of the plurality of second data subsets satisfies the second binning condition. If satisfied, step 1140 is executed; if not satisfied, step 1150 is executed.
[0071] In some embodiments, the second binning condition includes that the relationship between the user data in the second data subset and the target variable for user evaluation conforms to a preset relationship. For example, the preset relationship is determined according to the user evaluation service to be performed. The preset relationship can include at least one of monotonically increasing, monotonically decreasing, first monotonically increasing and then monotonically decreasing, or first monotonically decreasing and then monotonically increasing.
[0072] For example, determine whether the monotonicity of the user data in each child node (corresponding to each second data bin) after splitting conforms to the business logic for the target variable; in the case of non - compliance (such as inability to ensure monotonicity), perform a backtracking operation to adjust the splitting criteria; in the case of compliance and the data volume of each second data bin being greater than the minimum sample, stop splitting. Recursively operate on each initial node until all the obtained second data bins meet the splitting criteria and then terminate.
[0073] In this way, it can make the relationship between the binned user data and the target variable conform to the logic of the user evaluation service, thereby improving the processing effect of the user data.
[0074] In step 1140, determine multiple first data bins from multiple second data subsets that meet the second binning condition, and determine the binning thresholds from the candidate binning thresholds corresponding to the multiple second data subsets that meet the second binning condition.
[0075] In step 1150, in the case where there are second data subsets among the multiple second data subsets that do not meet the second binning condition, adjust the first binning condition and return to step 1120.
[0076] For example, according to the adjusted first binning condition, re - determine the candidate binning thresholds; according to the re - determined candidate binning thresholds, re - divide the first data subset into multiple second data subsets; repeat the above adjustment steps, determination steps, and division steps until each of the re - divided multiple second data subsets meets the second binning condition.
[0077] After binning the first data set, it is possible to Figure 1 continue to evaluate the user based on the user data through the remaining steps in
[0078] As Figure 1 shown, in step 120, perform binning processing on the second data set according to the binning thresholds in step 110 to obtain multiple second data bins. The first user data set and the second user data set correspond to the same user feature.
[0079] In step 130, determine the concept drift index of the user feature according to the difference between the first data bin and its corresponding second data bin.
[0080] In some embodiments, determine the concept drift index according to the difference between the data distribution of the first data bin and the data distribution of the second data bin. For example, the difference between the data distributions includes the WOE difference.
[0081] For example, calculate the WOE after binning of the user data of each user feature in the training set. The number of user features in the training set is D, and each user feature j is finally divided into N j first data bins, j ∈ D, representing the j-th user feature. The proportion of the data volume of the user data in each first data bin is The WOE of each first data bin is
[0082] For example, for the test set, based on the binning thresholds of each first data bin in the training set in the above steps, bin the user data of the corresponding user feature in the test set. The proportion of the data volume of the user data in each second data bin after binning is The WOE of each second data bin is
[0083] In some embodiments, determine weights according to the proportion of the data volume of the first data bin in the first data set and the proportion of the data volume of the second data bin in the second data set; use the weights to weight the differences; and determine the concept drift index according to the weighted differences.
[0084] For example, determine the first weight according to the proportion of the data volume of the first data bin in the first data set; determine the second weight according to the proportion of the data volume of the second data bin in the second data set; and determine the weight according to the weighted average of the first weight and the second weight.
[0085] For example, calculate the concept drift index of each user feature j:
[0086]
[0087] In the above embodiments, Index j measures the change in the WOE of the test set and the training set for the user feature j in the same interval; and uses the proportion of the data volume of each data bin as the weight to measure the proportion of the data volume affected by this change. In this way, through the adjustment of the weight, Index j can adapt to various different scenarios, thereby improving the accuracy of user evaluation and the effect of user data processing.
[0088] In the above embodiments, Index j The larger it is, the greater the change in the mapping relationship between the user data and the target variable of the user feature j in the training set and the test set, the more unstable the user feature, and the less suitable for evaluating the user. In this way, using Index j can screen the user features for user evaluation, thereby improving the accuracy of user evaluation and the effect of user data processing.
[0089] In step 140, the user is evaluated according to the concept drift index.
[0090] In some embodiments, according to the concept drift index, it is determined whether to select user features to evaluate the user. For example, when the concept drift index is greater than the drift threshold, the user features may not be used to evaluate the user; when the concept drift index is less than or equal to the drift threshold, the user features may be used to evaluate the user.
[0091] For example, the user evaluation may include classifying the user by using the user data corresponding to the user features to determine the user type to which the user belongs. For example, the user type may include whether the user belongs to a specified area, whether the user has a specified behavior, whether the user belongs to a user with risks, etc.
[0092] In this way, using Index j The user features for user evaluation can be screened, thereby improving the accuracy of user evaluation and the effect of user data processing.
[0093] In the above embodiments, after binning two data sets using the same binning threshold, the concept drift between the two data sets is evaluated according to the product of the difference in woe of each data bin and the corresponding weight. In this way, the concept drift can be intuitively reflected on a unified standard, thereby improving the accuracy of user estimation and the effect of user data processing.
[0094] Figure 3 A block diagram showing some embodiments of the user data processing apparatus of the present disclosure.
[0095] As Figure 3 shown, the user data processing apparatus 3 includes: a binning unit 31 for binning the user data in the first user data set according to the binning threshold to obtain a plurality of first data bins, binning the second data set according to the binning threshold to obtain a plurality of second data bins, where the first user data set and the second user data set correspond to the same user feature; a determination unit 32 for determining the concept drift index of the user feature according to the difference between the first data bin and its corresponding second data bin; an evaluation unit 33 for evaluating the user according to the concept drift index.
[0096] In some embodiments, the determination unit 32 determines the concept drift index according to the difference between the data distribution of the first data bin and the data distribution of the second data bin.
[0097] In some embodiments, the difference includes the WOE difference.
[0098] In some embodiments, the determination unit 32 determines weights based on the proportion of the data volume of the first data bin in the first data set and the proportion of the data volume of the second data bin in the second data set, and uses the weights to weight the differences; according to the weighted differences, a concept drift index is determined.
[0099] In some embodiments, the determination unit 32 determines a first weight according to the proportion of the data volume of the first data bin in the first data set, determines a second weight according to the proportion of the data volume of the second data bin in the second data set, and determines the weight according to the weighted average of the first weight and the second weight.
[0100] In some embodiments, the binning unit 31 divides the first user data set into multiple first data subsets, determines candidate binning thresholds according to the first binning condition for dividing the first data subsets into multiple second data subsets, and when each of the multiple second data subsets satisfies the second binning condition, determines the multiple second data subsets that satisfy the second binning condition as multiple first data bins, and determines the candidate binning thresholds corresponding to the multiple second data subsets that satisfy the second binning condition as the binning thresholds.
[0101] In some embodiments, when there are second data subsets that do not satisfy the second binning condition among the multiple second data subsets, the binning unit 31 adjusts the first binning condition, re-determines the candidate binning thresholds according to the adjusted first binning condition, re-divides the first data subsets into multiple second data subsets according to the re-determined candidate binning thresholds, and repeats the above adjustment steps, determination steps, and division steps until each of the re-divided multiple second data subsets satisfies the second binning condition.
[0102] In some embodiments, the first binning condition includes at least one of the proportion of the data volume of the first data subset in the first data set being greater than a first threshold or the desired number of data bins; the second binning condition includes that the relationship between the user data in the second data subset and the target variable for user evaluation conforms to a preset relationship.
[0103] In some embodiments, the preset relationship is determined according to the user evaluation service to be performed.
[0104] In some embodiments, the preset relationship includes at least one of monotonically increasing, monotonically decreasing, first monotonically increasing and then monotonically decreasing, or first monotonically decreasing and then monotonically increasing.
[0105] In some embodiments, the evaluation unit 33 determines whether to select user features to evaluate the user according to the concept drift index.
[0106] Figure 4 The block diagram showing some other embodiments of the user data processing device of the present disclosure.
[0107] As Figure 4 shown, the user data processing apparatus 4 of this embodiment includes: a memory 41 and a processor 42 coupled to the memory 41. The processor 42 is configured to execute the user data processing method in any one of the embodiments of the present disclosure based on instructions stored in the memory 41.
[0108] Among them, the memory 41 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, a database, and other programs.
[0109] Figure 5 Block diagrams showing still other embodiments of the user data processing apparatus of the present disclosure.
[0110] As Figure 5 shown, the user data processing apparatus 5 of this embodiment includes: a memory 510 and a processor 520 coupled to the memory 510. The processor 520 is configured to execute the user data processing method in any one of the foregoing embodiments based on instructions stored in the memory 510.
[0111] The memory 510 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory stores, for example, an operating system, application programs, a boot loader, and other programs.
[0112] The user data processing apparatus 5 may further include an input / output interface 530, a network interface 540, a storage interface 550, etc. These interfaces 530, 540, 550 and the memory 510 and the processor 520 may be connected through a bus 560, for example. Among them, the input / output interface 530 provides connection interfaces for input / output devices such as a display, a mouse, a keyboard, a touch screen, a microphone, and a speaker. The network interface 540 provides connection interfaces for various networking devices. The storage interface 550 provides connection interfaces for external storage devices such as an SD card and a USB flash drive.
[0113] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transitory storage media including, but not limited to, a magnetic disk memory, a CD-ROM, an optical memory, etc., which contain computer-usable program code.
[0114] So far, the user data processing method, user data processing apparatus, and non-volatile computer-readable storage medium according to the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details well known in the art are not described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.
[0115] The methods and systems of the present disclosure may be implemented in many ways. For example, the methods and systems of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.
[0116] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustration only and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for processing user data, comprising: binning the user data in the first user dataset according to binning thresholds to obtain a plurality of first data bins; binning the second dataset according to the binning thresholds to obtain a plurality of second data bins, wherein the first user dataset and the second user dataset correspond to the same user feature; determining a concept drift metric of the user feature according to the difference between a first data bin and its corresponding second data bin; evaluating the user according to the concept drift metric.
2. The method for processing user data according to claim 1, wherein, the determining the concept drift metric of the user feature according to the difference between a first data bin and its corresponding second data bin includes: determining the concept drift metric according to the difference between the data distribution of the first data bin and the data distribution of the second data bin.
3. The method for processing user data according to claim 2, wherein, the difference includes the Weight of Evidence (WOE) difference.
4. The method for processing user data according to any one of claims 1-3, wherein, the determining the concept drift metric according to the difference between a first data bin and its corresponding second data bin includes: determining weights according to the proportion of the data volume of the first data bin in the first dataset and the proportion of the data volume of the second data bin in the second dataset; weighting the difference by using the weights; determining the concept drift metric according to the weighted difference.
5. The method for processing user data according to claim 4, wherein, the determining weights according to the proportion of the data volume of the first data bin in the first dataset and the proportion of the data volume of the second data bin in the second dataset includes: determining a first weight according to the proportion of the data volume of the first data bin in the first dataset; determining a second weight according to the proportion of the data volume of the second data bin in the second dataset; determining the weights according to the weighted mean of the first weight and the second weight.
6. The method for processing user data according to any one of claims 1-3, wherein, the binning the user data in the first user dataset to obtain a plurality of first data bins includes: dividing the first user dataset into a plurality of first data subsets; determining candidate binning thresholds according to first binning conditions for dividing the first data subsets into a plurality of second data subsets; when each of the plurality of second data subsets satisfies second binning conditions, determining the plurality of second data subsets that satisfy the second binning conditions as the plurality of first data bins, and determining the candidate binning thresholds corresponding to the plurality of second data subsets that satisfy the second binning conditions as the binning thresholds.
7. The method for processing user data according to claim 6, wherein, the determining candidate binning thresholds according to first binning conditions includes: In the case where there is a second data subset among the multiple second data subsets that does not meet the second binning condition, adjust the first binning condition; According to the adjusted first binning condition, re-determine the candidate binning threshold; According to the re-determined candidate binning threshold, re-partition the first data subset into multiple second data subsets; Repeat the above adjustment step, determination step, and partitioning step until each of the re-partitioned multiple second data subsets meets the second binning condition.
8. The user data processing method according to claim 6, wherein: The first binning condition includes at least one of the proportion of the data volume of the first data subset in the first data set being greater than a first threshold, or the number of data bins to be obtained; The second binning condition includes that the relationship between the user data in the second data subset and the target variable for user evaluation conforms to a preset relationship.
9. The user data processing method according to claim 8, wherein, The preset relationship is determined according to the user evaluation service to be performed.
10. The user data processing method according to claim 8, wherein, The preset relationship includes at least one of monotonically increasing, monotonically decreasing, first monotonically increasing and then monotonically decreasing, or first monotonically decreasing and then monotonically increasing.
11. The user data processing method according to any one of claims 1-3, wherein, The evaluating the user according to the concept drift index includes: Determining whether to select the user feature to evaluate the user according to the concept drift index.
12. A user data processing device, comprising: A binning unit, configured to perform binning processing on the user data in the first user data set according to a binning threshold to obtain a plurality of first data bins, and perform binning processing on the second data set according to the binning threshold to obtain a plurality of second data bins, where the first user data set and the second user data set correspond to the same user feature; A determining unit, configured to determine the concept drift index of the user feature according to the difference between the first data bin and its corresponding second data bin; An evaluating unit, configured to evaluate the user according to the concept drift index.
13. A user data processing device, comprising: A memory; and A processor coupled to the memory, the processor being configured to execute the user data processing method according to any one of claims 1-11 based on instructions stored in the memory.
14. A non-volatile computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the user data processing method according to any one of claims 1-11.