Data drift detection method and device, equipment and storage medium
By acquiring and dividing user data in multiple scenarios and detecting data drift using prediction models, the problem of difficulty in detecting data drift in multiple scenarios is solved in the prior art, and efficient drift detection of user data in multiple scenarios is achieved.
Patent Information
- Application Number
- CN202311562238.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to perform drift detection on data in multiple scenarios at the same time, resulting in attenuation of model effects.
By obtaining user data in multiple scenarios of the current service, dividing it into training data sets and test data sets according to preset proportions, and inputting the test data into pre-trained prediction model to determine the model indicators of each scenario. When the model indicators are greater than the threshold, it is judged that the data drift occurs in user data.
It realizes the drift detection of user data in multiple scenarios in the current service at the same time, avoids complex modeling and parameter adjustment processes, and can find the global optimality through the interaction of multiple scenarios and avoids falling into the local optimality.
Smart Images

Figure CN120030397A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of data processing technology, and in particular to a data drift detection method, device, equipment and storage medium. Background Art
[0002] The data distribution in business applications is often affected by changes in external environment, policy adjustments and other factors, resulting in data drift. When training models based on these data, the model effect will be attenuated. Therefore, it is necessary to detect and monitor data distribution changes during the modeling process and when monitoring the model online to avoid model attenuation caused by data drift.
[0003] In the prior art, data distribution changes can be monitored based on methods such as group stability indicators and distance calculations. The ability of the entire data set to distinguish between training samples and test samples can also be evaluated through supervised adversarial verification methods to determine whether the distribution of the test set data has changed significantly compared to the training set data. Compared with group stability indicators and distance calculations, supervised adversarial verification methods can measure whether the distribution of the test set has changed from the perspective of the entire data set.
[0004] In the process of implementing the present invention, the inventors found that there are at least the following technical problems in the prior art:
[0005] When the current business has multiple scenarios, it is impossible to perform drift detection on the data of multiple scenarios at the same time. Summary of the invention
[0006] The present invention provides a data drift detection method, device, equipment and storage medium to realize simultaneous drift detection of user data in multiple scenarios within a current business.
[0007] In a first aspect, an embodiment of the present invention provides a data drift detection method, including:
[0008] Acquire user data in multiple scenarios of the current business, and divide the user data into a training data set and a test data set according to a preset ratio;
[0009] Inputting each test data in the test data set into a pre-trained prediction model, respectively, so that the prediction model determines the prediction score of each test data corresponding to each scenario, wherein the prediction model is trained by the training data set;
[0010] For each of the scenarios, a model index corresponding to the scenario is determined according to the predicted score of each of the test data corresponding to the scenario, and when it is determined that the model index is greater than an index threshold, it is determined that data drift has occurred in the user data corresponding to the scenario.
[0011] In a second aspect, an embodiment of the present invention further provides a data drift detection device, including:
[0012] An acquisition module, used to acquire user data in multiple scenarios of the current business, and divide the user data into a training data set and a test data set according to a preset ratio;
[0013] A prediction module, used for inputting each test data in the test data set into a pre-trained prediction model, so that the prediction model determines the prediction score of each test data corresponding to each scenario, wherein the prediction model is trained by the training data set;
[0014] An execution module is used to determine, for each of the scenarios, a model index corresponding to the scenario based on the predicted scores of the test data corresponding to the scenario, and determine that data drift occurs in the user data corresponding to the scenario when it is determined that the model index is greater than an index threshold.
[0015] In a third aspect, an embodiment of the present invention further provides a computer device, the computer device comprising:
[0016] one or more processors;
[0017] a storage device for storing one or more programs,
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the data drift detection method as described in any one of the first aspects.
[0019] In a fourth aspect, an embodiment of the present invention further provides a storage medium comprising computer executable instructions, wherein the computer executable instructions, when executed by a computer processor, are used to perform the data drift detection method as described in any one of the first aspects.
[0020] The embodiments of the above invention have the following advantages or beneficial effects:
[0021] An embodiment of the present invention provides a data drift detection method, comprising: obtaining user data in multiple scenarios of a current business, dividing the user data into a training data set and a test data set according to a preset ratio; inputting each test data in the test data set into a pre-trained prediction model, so that the prediction model determines the prediction score of each test data corresponding to each scenario, wherein the prediction model is trained by the training data set; for each scenario, determining a model index corresponding to the scenario according to the prediction score of each test data corresponding to the scenario, and determining that data drift occurs in the user data corresponding to the scenario when it is determined that the model index is greater than an index threshold. The above technical solution can firstly obtain user data belonging to multiple scenarios in the current business, and divide the user data into a training data set and a test data set. Secondly, each test data in the test data set can be input into the prediction model trained based on the training data set. The user data includes user characteristics, scene labels and time labels. Therefore, the prediction model trained based on the training data set can be used to determine the prediction score of each test data corresponding to each scenario, so as to determine the probability that the test data is a sample obtained in the second time period with a time label in the scenario to which it belongs. Then, for each scenario, the model index corresponding to the scenario can be determined according to the prediction score of each test data corresponding to the scenario. The larger the model index of the scenario, the greater the change of the test data obtained in the second time period with a time label in the test data corresponding to the scenario compared with the test data obtained with a time label of the first time period, that is, the test data corresponding to the scenario has data drift, and then it can be determined that the user data corresponding to the scenario has data drift, so as to realize drift detection of user data of multiple scenarios in the current business at the same time. Compared with the supervised adversarial verification method, it avoids the complex modeling and parameter adjustment process, and can find the global optimum with the help of the interaction of multiple scenarios to avoid falling into the local optimum. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flow chart of a data drift detection method provided by an embodiment of the present invention;
[0023] Figure 2 A flow chart of another data drift detection method provided by an embodiment of the present invention;
[0024] Figure 3 A schematic diagram of a multi-task learning model in another data drift detection method provided by an embodiment of the present invention;
[0025] Figure 4 A schematic diagram of the structure of a data drift detection device provided by an embodiment of the present invention;
[0026] Figure 5A schematic diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0028] It should be mentioned before discussing exemplary embodiments in more detail that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe various operations (or steps) as sequential processes, many operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of various operations can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to methods, functions, procedures, subroutines, subprograms, etc. In addition, the embodiments in the present invention and the features in the embodiments can be combined with each other without conflict.
[0029] Methods based on group stability indicators and distance calculation are usually used to detect the stability and distribution differences of features in a data set. However, these methods can only evaluate whether the distribution of a single feature has changed. They do not consider the association between features and therefore cannot evaluate the overall distribution changes of the data set.
[0030] The adversarial verification method can evaluate the distribution changes of the data set from the level of feature sets. Compared with methods such as group stability indicators and distance calculation, it can convert a single feature dimension into a feature set dimension, and can more directly reflect the changes in the data distribution of the entire data set. However, the data for actual applications usually comes from multiple different scenarios. Traditional machine learning methods such as xgboost can only perform adversarial verification on each scenario separately to measure whether the distribution of the data set as a whole or in a certain scenario has changed. If you need to evaluate whether the data distribution of different scenarios has changed, you need to perform adversarial verification on each scenario separately, which causes duplication and redundancy of the model and increases the time cost.
[0031] Therefore, the present application proposes a data drift detection method to detect data drift in different scenarios within a business.
[0032] The data drift detection method proposed in the present application will be described in detail below with reference to diagrams and embodiments.
[0033] Figure 1The flowchart of a data drift detection method provided by an embodiment of the present invention is applicable to the case where it is necessary to perform scenario-based drift detection on user data within a service. The method can be executed by a data drift detection device, which can be implemented by software and / or hardware. Figure 1 Said method specifically comprises the following steps:
[0034] Step 110: Obtain user data in multiple scenarios of the current business, and divide the user data into a training data set and a test data set according to a preset ratio.
[0035] Among them, the current business is a specific business in actual applications, such as risk assessment business, sales evaluation business, etc. The scenario can be understood as a customer group or a sub-business. For example, for the risk assessment business, multiple customer groups of different age groups and different risk assessment sub-businesses can all be scenarios corresponding to the risk assessment business. User data can be understood as user behavior data and user portraits, etc., and both user behavior data and user portraits can include multiple user characteristics.
[0036] Specifically, for the current business, the current business corresponds to multiple customer groups and multiple sub-businesses. Therefore, user data can be obtained in multiple scenarios of the current business, and the obtained user data belongs to different scenarios, that is, different customer groups or sub-businesses. In order to detect drift of user data, the user data includes historical user data obtained in a first time period and time-sensitive user data obtained in a second time period. It can be understood that the first time period is earlier than the second time period. Furthermore, the historical user data and time-sensitive user data included in the user data can be scattered, and then the user data can be divided into a training data set and a test data set according to a preset ratio. The training data set and the test data set both include historical user data and time-sensitive user data.
[0037] In practical applications, user data includes user features, scenario labels, and time labels. User features may include the number of historical orders placed by the user, the number of historical order returns, etc. The scenario label of the user data is determined by the scenario to which the user data belongs. The scenario label of the user data is marked as a dummy variable. When the user data is obtained from n scenarios of the current business, the scenario label of the user data obtained from scenario 1 is [1, 0, 0, ..., 0], and the scenario label of the user data obtained from scenario 2 is [0, 1, 0, ..., 0]. Similarly, the scenario label of the user data obtained from scenario n is [0, 0, 0, ..., 1]. The time label of the user data is determined by the time period in which the user data is obtained. When the user data is obtained in the first time period, it can be determined that the user data is a negative sample, and then the time label of the user data can be determined to be 0. When the user data is obtained in the second time period, it can be determined that the user data is a positive sample, and then the time label of the user data can be determined to be 1.
[0038] In an embodiment of the present invention, user data of the current business is obtained, and the user data belongs to multiple different scenarios, that is, user data belonging to multiple different scenarios is obtained within the current business, and data division of the user data is implemented to obtain a training data set and a test data set for data drift detection.
[0039] Step 120: Input each test data in the test data set into a pre-trained prediction model, so that the prediction model determines the prediction score of each test data corresponding to each scenario.
[0040] The prediction model is obtained by training the training data set.
[0041] Specifically, each test data in the test data set is input into the pre-trained prediction model respectively. Since the prediction model is trained by the training data set, each training data in the training data set includes user features, scene labels and time labels. Therefore, the prediction score of each test data corresponding to each scene can be determined according to the prediction model. The prediction score is any value between 0 and 1.
[0042] The prediction score of the test data corresponding to the scenario indicates the probability that the test data is a sample obtained in the second time period.
[0043] In the embodiment of the present invention, based on the prediction model obtained by training the training data set, the prediction score of each test data corresponding to each scenario is determined to determine the probability that the test data is a sample obtained in the second time period.
[0044] Step 130: For each of the scenarios, determine the model index corresponding to the scenario according to the predicted score of each of the test data corresponding to the scenario, and determine that data drift occurs in the user data corresponding to the scenario when it is determined that the model index is greater than an index threshold.
[0045] Specifically, after determining the prediction scores of the test data corresponding to each scenario, for each scenario, the model metrics of the scenario can be determined based on the prediction scores and time tags of all the test data corresponding to the scenario. The model metrics can be used as the data drift metrics for the scenario. The larger the model metrics of the scenario, the greater the change in the test data obtained in the second time period compared to the test data obtained in the first time period among the test data corresponding to the scenario. When it is determined that the model metrics of the scenario are greater than the metric threshold, it can be determined that the change in the test data obtained in the second time period compared to the test data obtained in the first time period among the test data corresponding to the scenario is greater, that is, it can be determined that the test data corresponding to the scenario has data drift. Furthermore, it can be determined that the user data corresponding to the scenario has data drift, thus realizing the simultaneous drift detection of the user data of multiple scenarios in the current business.
[0046] In the embodiments of the present invention, it is realized to determine the model metrics corresponding to the scenario according to the prediction scores of the test data corresponding to the scenario, and when it is determined that the model metrics are greater than the metric threshold, it is determined that the test data corresponding to the scenario has data drift. Furthermore, it can be determined that the user data corresponding to the scenario has data drift, thus realizing the simultaneous drift detection of the user data of multiple scenarios in the current business.
[0047] The data drift detection method provided by an embodiment of the present invention includes: obtaining user data in multiple scenarios of the current business, and dividing the user data into a training data set and a test data set according to a preset ratio; inputting each test data in the test data set into a pre-trained prediction model, so that the prediction model determines the prediction score of each test data corresponding to each scenario, wherein the prediction model is trained by the training data set; for each scenario, determining the model index corresponding to the scenario according to the prediction score of each test data corresponding to the scenario, and determining that data drift occurs in the user data corresponding to the scenario when it is determined that the model index is greater than an index threshold. The above technical solution can firstly obtain user data belonging to multiple scenarios in the current business, and divide the user data into a training data set and a test data set. Secondly, each test data in the test data set can be input into the prediction model trained based on the training data set. The user data includes user characteristics, scene labels and time labels. Therefore, the prediction model trained based on the training data set can be used to determine the prediction score of each test data corresponding to each scenario, so as to determine the probability that the test data is a sample obtained in the second time period with a time label in the scenario to which it belongs. Then, for each scenario, the model index corresponding to the scenario can be determined according to the prediction score of each test data corresponding to the scenario. The larger the model index of the scenario, the greater the change of the test data obtained in the second time period with a time label in the test data corresponding to the scenario compared with the test data obtained in the first time period, that is, the test data corresponding to the scenario has data drift, and then it can be determined that the user data corresponding to the scenario has data drift, so as to realize drift detection of user data of multiple scenarios in the current business at the same time. Compared with the supervised adversarial verification method, it avoids the complex modeling and parameter adjustment process, and can find the global optimum with the help of the interaction of multiple scenarios to avoid falling into the local optimum.
[0048] Figure 2A flowchart of another data drift detection method provided in an embodiment of the present invention. The embodiment of the present invention can be applied to situations where scene-based drift detection of data within an application is required. Based on the above embodiments, the embodiment of the present invention adds "determining the statistical value of each user feature in the user data based on each training data in the training data set; normalizing the user data according to the statistical value of each user feature to obtain normalized user data" after dividing the user data into a training user data set and a test user data set according to a preset ratio. Before inputting each test data in the test data set into a pre-trained prediction model, "constructing a multi-task learning model according to the number of scenarios in the current business" and "training the multi-task learning model based on the training data set to obtain the prediction model" are added. After determining that the user data corresponding to the scenario has data drift, "for the user data corresponding to the scenario, sorting the user data obtained in the first time period based on the predicted score of the user data, and performing equal-frequency binning according to the sorting result; determining a score threshold according to the binning result, and determining the user data with the predicted score greater than the score threshold as the target user data corresponding to the scenario" is added. The explanations of the terms that are the same as or corresponding to the above embodiments are not repeated here. See Figure 2 , the data drift detection method provided by the embodiment of the present invention includes:
[0049] Step 210: Obtain user data in multiple scenarios of the current business, and divide the user data into a training data set and a test data set according to a preset ratio.
[0050] The user data includes user data acquired in a first time period and user data acquired in a second time period, and the first time period is earlier than the second time period.
[0051] Specifically, user data can be obtained in multiple scenarios of the current business. According to the acquisition time, the user data includes historical user data obtained in a first time period and time-sensitive user data obtained in a second time period. The first time period is earlier than the second time period. The historical user data obtained in the first time period is a negative sample with a time label of 0, and the time-sensitive user data obtained in the second time period is a positive sample with a time label of 1. Then, the historical user data and time-sensitive user data included in the user data are scattered, and then the user data is divided into a training data set and a test data set according to a first preset ratio. The training data set and the test data set both include historical user data and time-sensitive user data. The training data set is used to train the prediction model. In actual applications, in order to train a more accurate prediction model, the user data can also be divided into a training data set, a verification data set, and a test data set according to a second preset ratio. For example, the user data can be divided into a training data set, a verification data set, and a test data set according to 7:1.5:1.5. The training data set, the verification data set, and the test data set all include historical user data and time-sensitive user data.
[0052] In actual applications, historical user data can be obtained in a first time period before the current moment, and time-sensitive user data can be obtained in a second time period before the current moment. The first time period is earlier than the second time period, that is, the second time period is closer to the current moment than the first time period, and the time-sensitive user data obtained in the second time period is more time-sensitive than the historical user data obtained in the first time period. For example, historical user data can be user data obtained within the training sample time for training the actual model, and time-sensitive user data can be OOT (Out of Time) user data, that is, user data obtained outside the training sample time for verifying the effect of the actual model.
[0053] The user data may further include a scene tag indicating the scene to which the user data belongs. As described in the first embodiment, when the user data is obtained from n scenes of the current service, the scene tag of the user data obtained from scene 1 is [1, 0, 0, ..., 0], the scene tag of the user data obtained from scene 2 is [0, 1, 0, ..., 0], and so on. The scene tag of the user data obtained from scene n is [0, 0, 0, ..., 1].
[0054] In an embodiment of the present invention, user data of the current business is obtained, and the user data belongs to multiple different scenarios, that is, user data belonging to multiple different scenarios is obtained within the current business, and data division of the user data is implemented to obtain a training data set, a verification data set and a test data set for data drift detection.
[0055] Step 220: Determine the statistical value of each user feature in the user data based on each of the training data in the training data set; and normalize the user data according to the statistical value of each of the user features to obtain normalized user data.
[0056] The user characteristics include the user's behavioral characteristics, which may be the user's historical order number, historical order return number, etc. The statistical value of the user characteristics may be the mean and standard deviation of the user characteristics.
[0057] Specifically, the original user data often contains noise, and the original user data can be processed to obtain data that can be used for prediction model training. Specifically, first, the user labels and time information contained in the user data need to be removed, and then the duplicate rows contained in the user data need to be removed, and the missing user features in the user data need to be filled with 0 values. Furthermore, the user data can be normalized to eliminate the influence of the unit and scale differences between the various user features contained in the user data, so that the process of finding the optimal solution of the prediction model becomes smooth and easy to converge.
[0058] Therefore, based on each training data in the training data set, the statistical value of each user feature in the user data can be determined, for example, the mean μ and standard deviation σ of each user feature in the training data can be determined, and the mean μ and standard deviation σ of each user feature in the training data can be determined as the mean μ and standard deviation σ of each user feature in the user data, and each user data obtained in multiple scenarios of the current business can be normalized according to the mean μ and standard deviation σ of each user feature to obtain normalized user data. The user feature x in the user data is converted to obtain a normalized user feature Z, and then the user data consisting of the normalized user features corresponding to all the user features in the user data is determined as the normalized user data, thereby achieving normalization processing of the user data.
[0059] In an embodiment of the present invention, after determining the statistical value of each user feature of the user data based on the training data set, normalization processing is performed on each user data obtained in multiple scenarios of the current business based on the statistical value of each user feature to obtain normalized user data.
[0060] Step 230: construct a multi-task learning model according to the number of the scenarios in the current business.
[0061] The multi-task learning model consists of an input layer, a multi-gate mixture of experts layer (Multi-Gate Mixture-of-ExpersLayer, MMoE layer), a tower layer, and an output layer. The MMoE layer consists of Expert modules and Gate modules, and the number of Gate modules is determined by the number of scenarios in the current business.
[0062] Specifically, after determining the number of scenarios in the current business, the number of Gate modules in the multi-task learning model can be determined according to the number of scenarios in the current business, thereby constructing a multi-task learning model suitable for multi-scenario data drift detection.
[0063] Figure 3 A schematic diagram of a multi-task learning model in another data drift detection method provided by an embodiment of the present invention, such as Figure 3 As shown in Figure 2, the multi-task learning model includes the input layer, MMoE layer, Tower layer ( Figure 3 TowerLayer 1…TowerLayer n) and the output layer ( Figure 3 OutputLayer 1…OutputLayer n in the MMoE layer). The input layer is used to input user data into the MMoE layer. The MMoE layer is used to determine the weighted sum of the probability of each Gate module being selected for each Expert module. The Tower layer and the output layer are used to determine the predicted score of each user data corresponding to each scenario. The number of Expert modules in the MMoE layer is N. Each Expert module is composed of several layers of fully connected layers, and the number of hidden layer units is set to 64 (which can be adjusted according to actual needs). The parameters between each Expert module are not shared, so the output obtained according to the input training is also different. When the number of scenarios in the current business is n, the number of Gate modules can be determined to be n. Each Gate module is composed of a fully connected layer. The Gate module can output the probability of each Expert module being selected when determining the predicted score of each test data corresponding to each scenario. The output of the MMoE layer is the weighted sum of the probabilities of each Gate module being selected for each Expert module. The output of the MMoE layer can be input into each Tower layer. Each Tower layer is composed of several fully connected layers. The output layer is composed of fully connected layers. The output layer takes the output of the Tower layer as input and outputs the predicted scores of each test data corresponding to each scenario.
[0064] It should be noted that the multi-task learning model can also be other multi-task learning models such as PLE.
[0065] In an embodiment of the present invention, a multi-task learning model suitable for multi-scenario data drift detection is constructed according to the number of scenarios in the current business.
[0066] Step 240: Train the multi-task learning model based on the training data set to obtain the prediction model.
[0067] In one implementation, the user data includes user features, scene tags, and time tags. Accordingly, step 240 may specifically include:
[0068] Input each of the training data in the training data set into the multi-task learning model so that the multi-task learning model determines the prediction score of each of the training data corresponding to each of the scenarios; determine the loss function of each of the scenarios according to the time label and the prediction score of each of the training data corresponding to each of the scenarios; obtain the prediction model when it is determined that the loss function of each of the scenarios has converged, the number of training times is greater than the number threshold, or the model index is greater than the index threshold.
[0069] Specifically, each training data in the training data set is input into the multi-task learning model. The input layer can input each training data into the MMoE layer. For each Expert module in the MMoE layer, the weight and bias are initialized using a uniform distribution initializer, and l2 regularization is set to prevent overfitting. The activation function of each Expert module is set to the Relu function f(x) = max{0, x}, where x represents the training data, so that the model calculation speed and convergence speed are faster. The output of each Expert module is f i (x) = max{0, w i x+b}, where i=1,2,3…N, represents the ith Expert module, w i is the weight of the ith Expert module, and b is the bias of the Expert module. For each Gate module, the activation function is set to the Softmax function x is the output of the previous hidden layer of the activation function, with a dimension of N. The activation function ultimately converts the output of the previous layer into a number y between 0 and 1 i ,and This result represents the probability of the current Gate module selecting the output of each Expert module. Therefore, the probability of each Gate module selecting the output of each Expert module can be expressed as g n (x) i =softmax(w gn x+b), where n represents the nth scenario, i represents the i-th Expert module, and w gn represents the weight of the nth Gate module, and b is the deviation of each Gate. Then, the output of the MMoE layer can be determined as It means that the weights determined by each Gate module are weighted summed with the weights of each Expert module. The input of each Tower layer is the output of the MMoE layer. The initial weight and bias of each Tower layer are determined by the normal distribution initializer. Its activation function is the Relu function f(x) = max{0, x}. Finally, the output of each Tower layer can be determined as h n =max{0,wf n (x)+b}, where w is the weight of the Tower layer, b is the bias of the Tower layer, the output of each Tower layer is the input of the output layer, and the input layer uses the sigmoid activation function to output the predicted scores of each training data corresponding to each scenario.
[0070] When training a multi-task learning model, it is necessary to set a corresponding loss function for each scenario. The loss function is binary cross entropy. Among them, y i is the time label of the i-th training data in the scene, is the prediction score of the i-th training data in the scene. During the training process, it is necessary to identify the training data corresponding to each scene, and determine the loss of each scene training process based on its corresponding time label and prediction score.
[0071] In addition, the optimizer of the model can use the Adam optimizer, the number of training rounds can be set to 100, and the drop_out parameter (between 0 and 1) can be set to prevent overfitting of the model. The number threshold and the index threshold can be set according to actual needs. The batch size of each round of batch training is 1024. When the number of training times is less than or equal to the number threshold, the loss function of each scene converges, and the model when each loss function converges is determined as the prediction model. After each round of training, the model is verified with the validation data set to determine the AUC corresponding to each scene. AUC is a common classification model evaluation indicator, and its range is generally distributed between 0 and 1. The larger the value, the better the model effect. When the AUC corresponding to each scene is greater than the index threshold, the model when the AUC corresponding to each scene is greater than the index threshold can be determined as the prediction model. After the number of training times is greater than the number threshold, the loss function of any scene has not converged and the AUC corresponding to any scene is less than or equal to the index threshold, the model when the AUC of most scenes is greater than the index threshold can be determined as the prediction model, and the model training is implemented to obtain the prediction model.
[0072] It should be noted that model optimization can also be performed based on optimizers such as stochastic gradient descent and Momentum.
[0073] In an embodiment of the present invention, a multi-task learning model is trained based on training data including user features, scene labels and time labels to obtain a prediction model, so that the prediction model can be used to determine the prediction score of each training data corresponding to each scene.
[0074] Step 250: Input each test data in the test data set into a pre-trained prediction model, so that the prediction model determines the prediction score of each test data corresponding to each scenario.
[0075] The prediction model is obtained by training the training data set.
[0076] As in the first embodiment, each test data in the test data set is input into a pre-trained prediction model, and a prediction score of each test data corresponding to each scenario can be determined according to the prediction model. The prediction score is any value between 0 and 1.
[0077] The prediction score of the test data corresponding to the scenario indicates the probability that the test data is a sample obtained in the second time period.
[0078] In the embodiment of the present invention, based on the prediction model obtained by training the training data set, the prediction score of each test data corresponding to each scenario is determined to determine the probability that the test data is a sample obtained in the second time period.
[0079] Step 260: For each of the scenarios, determine the model index corresponding to the scenario based on the predicted scores of the test data corresponding to the scenario, and determine that data drift occurs in the user data corresponding to the scenario when it is determined that the model index is greater than an index threshold.
[0080] In one implementation, determining the model index corresponding to the scenario according to the prediction score of each of the test data corresponding to the scenario includes:
[0081] The model index corresponding to the scenario is determined according to the prediction score and the time label of each test data corresponding to the scenario.
[0082] Among them, the model index of the scene can be used as the data drift index of the corresponding scene. The larger the model index is, the greater the change in the test data corresponding to the scene obtained in the second time period compared to the test data obtained in the first time period. When the model index is greater than the index threshold, it can be determined that the test data corresponding to the scene has data drift, which further indicates that the user data corresponding to the scene has data drift. The index threshold can be set according to actual needs, for example, it can be set to 0.9.
[0083] Specifically, for each scenario, the model index corresponding to the scenario can be determined based on the prediction scores and time tags of each test data corresponding to the scenario, and specifically the AUC of the scenario can be determined. Then, the AUC of the scenario and the index threshold can be compared. When the AUC of the scenario is greater than the index threshold, it can be determined that the test data corresponding to the scenario has data drift, and then it can be determined that the user data corresponding to the scenario has data drift.
[0084] In an embodiment of the present invention, a model index corresponding to a scenario is determined based on the predicted scores and time labels of each test data corresponding to the scenario. When it is determined that the model index is greater than the index threshold, it is determined that data drift has occurred in the test data corresponding to the scenario. Furthermore, it can be determined that data drift has occurred in the user data corresponding to the scenario, thereby achieving simultaneous drift detection for user data of multiple scenarios in the current business.
[0085] Step 270: For the user data corresponding to the scenario, sort the user data obtained in the first time period based on the predicted scores of the user data, and perform equal-frequency binning according to the sorting results; determine a score threshold according to the binning results, and determine the user data with the predicted scores greater than the score threshold as the target user data corresponding to the scenario.
[0086] Specifically, after determining that the user data corresponding to the scene has data drift, it can be determined that the user data obtained in the second time period in the user data corresponding to the scene has a greater change than the user data obtained in the first time period. In order to reduce the attenuation of the actual model effect, it is necessary to filter the user data corresponding to the scene where data drift occurs to filter out the data in the user data corresponding to the scene that is closer to the distribution of the historical user data.
[0087] Specifically, for user data corresponding to scenarios where data drift occurs, data with a sample label of 1 are all timely user data, and data with a sample label of 0 are all historical user data. Therefore, the user data with a sample label of 0 can be sorted, and equal-frequency binning can be performed based on the sorting results. The score threshold is determined based on the binning results, and user data with a predicted score greater than the score threshold is determined as the target user data corresponding to the scenario. Compared with user data with a predicted score less than the score threshold, the distribution of target user data is closer to that of timely user data. The target user data and timely user data can be used as training data to train the actual model.
[0088] In addition, the score threshold can be determined according to the amount of training data required for actual model training.
[0089] In the embodiment of the present invention, the user data corresponding to the scenario where data drift occurs is screened to obtain the target user data and time-sensitive user data that can be used to train the actual model.
[0090] The data drift detection method provided by the embodiment of the present invention includes: obtaining user data in multiple scenarios of the current business, dividing the user data into a training data set and a test data set according to a preset ratio; determining the statistical value of each user feature in the user data based on each training data in the training data set; normalizing the user data according to the statistical value of each user feature to obtain normalized user data; building a multi-task learning model according to the number of scenarios in the current business; training the multi-task learning model based on the training data set to obtain the prediction model; inputting each test data in the test data set into the pre-trained prediction model respectively, so that the prediction model determines the prediction score of each of the test data corresponding to each of the scenarios; for each of the scenarios, the model index corresponding to the scenario is determined according to the prediction score of each of the test data corresponding to the scenario, and when it is determined that the model index is greater than the index threshold, it is determined that data drift has occurred in the user data corresponding to the scenario; for the user data corresponding to the scenario, the user data acquired in the first time period is sorted based on the prediction score of the user data, and equal-frequency binning is performed according to the sorting result; a score threshold is determined according to the binning result, and the user data with the prediction score greater than the score threshold is determined as the target user data corresponding to the scenario. The above technical solution can firstly obtain user data belonging to multiple scenarios in the current business, and divide the user data into a training data set, a verification data set and a test data set. Secondly, after determining the statistical value of each user feature of the user data according to the training data set, the user data is normalized according to the statistical value of each user feature to obtain normalized user data. A multi-task learning model suitable for multi-scenario data drift detection can also be constructed according to the number of scenarios in the current business. The multi-task learning model is trained based on the training data set to obtain a prediction model, so that the prediction model can be used to determine the prediction score of each test data corresponding to each scenario, and the test data set is used as the prediction score of each test data. Each test data is respectively input into the prediction model trained based on the training data set. The prediction model can determine the probability that the test data is a sample obtained in the second time period in the scenario to which it belongs. Furthermore, for each scenario, the model index corresponding to the scenario can be determined according to the prediction score and time label of each test data corresponding to the scenario. The larger the model index of the scenario, the greater the change in the test data obtained in the second time period in the test data corresponding to the scenario compared to the test data obtained in the first time period, that is, the test data corresponding to the scenario has data drift, and further, it can be determined that the user data corresponding to the scenario has data drift, so as to realize drift detection for user data of multiple scenarios in the current business at the same time.
[0091] In addition, after determining that the user data corresponding to the scenario has drifted, the user data corresponding to the scenario where the data drift has occurred can also be screened to obtain target user data and time-sensitive user data that can be used to train the actual model.
[0092] Figure 4 A schematic diagram of the structure of a data drift detection device provided by an embodiment of the present invention. The device and the data drift detection methods of the above embodiments belong to the same inventive concept, and the details not described in detail in the embodiments of the data drift detection device can refer to the embodiments of the above data drift detection methods.
[0093] The specific structure of the data drift detection device is as follows: Figure 4 As shown, including:
[0094] An acquisition module 410 is used to acquire user data in multiple scenarios of the current business, and divide the user data into a training data set and a test data set according to a preset ratio;
[0095] A prediction module 420, configured to input each test data in the test data set into a pre-trained prediction model, so that the prediction model determines a prediction score of each test data corresponding to each scenario, wherein the prediction model is trained by the training data set;
[0096] The execution module 430 is used to determine, for each of the scenarios, the model index corresponding to the scenario according to the predicted scores of the test data corresponding to the scenario, and determine that data drift occurs in the user data corresponding to the scenario when it is determined that the model index is greater than an index threshold.
[0097] In one implementation, the user data includes user data acquired in a first time period and user data acquired in a second time period, and the first time period is earlier than the second time period.
[0098] Based on the above embodiment, the device further includes:
[0099] A processing module is used to determine the statistical value of each user feature in the user data based on each training data in the training data set, wherein the user feature includes the user's behavioral characteristics; and normalize the user data according to the statistical value of each user feature to obtain normalized user data.
[0100] Based on the above embodiment, the device further includes:
[0101] A building module is used to build a multi-task learning model according to the number of scenarios in the current business.
[0102] Based on the above embodiment, the device further includes:
[0103] A training module is used to train the multi-task learning model based on the training data set to obtain the prediction model.
[0104] In one implementation, the user data includes user features, scene labels and time labels. Accordingly, the training module is specifically used to:
[0105] Input each of the training data in the training data set into the multi-task learning model so that the multi-task learning model determines the prediction score of each of the training data corresponding to each of the scenarios; determine the loss function of each of the scenarios according to the time label and the prediction score of each of the training data corresponding to each of the scenarios; obtain the prediction model when it is determined that the loss function of each of the scenarios has converged, the number of training times is greater than the number threshold, or the model index is greater than the index threshold.
[0106] Based on the above embodiment, the execution module 430 is specifically used for:
[0107] For each of the scenarios, the model index corresponding to the scenario is determined based on the predicted scores of each of the test data corresponding to the scenario and the time labels, and when it is determined that the model index is greater than an index threshold, it is determined that data drift has occurred in the user data corresponding to the scenario.
[0108] Based on the above embodiment, the device further includes:
[0109] The optimization module is used to, after determining that data drift occurs in the user data corresponding to the scenario, sort the user data obtained in the first time period based on the predicted scores of the user data corresponding to the scenario, and perform equal-frequency binning according to the sorting results; determine a score threshold according to the binning results, and determine the user data with the predicted score greater than the score threshold as the target user data corresponding to the scenario.
[0110] The data drift detection device provided in the embodiment of the present invention can execute the data drift detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the data drift detection method.
[0111] It is worth noting that in the embodiment of the above-mentioned data drift detection device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0112] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 5 A block diagram of an exemplary computer device 5 suitable for use in implementing embodiments of the present invention is shown. Figure 5 The computer device 5 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0113] like Figure 5 As shown, the computer device 5 is in the form of a general-purpose computing electronic device. The components of the computer device 5 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0114] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. By way of example, these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0115] The computer device 5 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 5, including volatile and non-volatile media, removable and non-removable media.
[0116] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 5 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 5 not shown, usually called a "hard drive"). Although Figure 5 Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The system memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.
[0117] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0118] The computer device 5 may also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with the computer device 5, and / or any device that enables the computer device 5 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed through an input / output (I / O) interface 22. Furthermore, the computer device 5 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 20. Figure 5 As shown, the network adapter 20 communicates with other modules of the computer device 5 via the bus 18. It should be understood that although Figure 5 Not shown, other hardware and / or software modules may be used in conjunction with computer device 5, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0119] The processing unit 16 executes various functional applications and page displays by running the programs stored in the system memory 28, for example, implementing the data drift detection method provided in the embodiment of the present invention, which includes:
[0120] Acquire user data in multiple scenarios of the current business, and divide the user data into a training data set and a test data set according to a preset ratio;
[0121] Inputting each test data in the test data set into a pre-trained prediction model, respectively, so that the prediction model determines the prediction score of each test data corresponding to each scenario, wherein the prediction model is trained by the training data set;
[0122] For each of the scenarios, the model index corresponding to the scenario is determined according to the predicted score of each of the test data corresponding to the scenario, and when it is determined that the model index is greater than an index threshold, it is determined that data drift has occurred in the user data corresponding to the scenario.
[0123] Of course, those skilled in the art can understand that the processor can also implement the technical solution of the data drift detection method provided by any embodiment of the present invention.
[0124] An embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, a data drift detection method provided by an embodiment of the present invention is implemented, for example, and the method includes:
[0125] Acquire user data in multiple scenarios of the current business, and divide the user data into a training data set and a test data set according to a preset ratio;
[0126] Inputting each test data in the test data set into a pre-trained prediction model, respectively, so that the prediction model determines the prediction score of each test data corresponding to each scenario, wherein the prediction model is trained by the training data set;
[0127] For each of the scenarios, a model index corresponding to the scenario is determined according to the predicted score of each of the test data corresponding to the scenario, and when it is determined that the model index is greater than an index threshold, it is determined that data drift has occurred in the user data corresponding to the scenario.
[0128] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0129] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0130] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0131] Computer program code for performing the operations of the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0132] It should be understood by those skilled in the art that the modules or steps of the present invention described above can be implemented by a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, optionally, they can be implemented by a program code executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0133] It should be noted that in the technical solution of the present invention, the collection, use, storage, sharing and transfer of user personal information involved are in compliance with the provisions of relevant laws and regulations, and it is necessary to inform the user and obtain the user's consent or authorization. When applicable, the user's personal information is de-identified and / or anonymized and / or encrypted.
[0134] For example: After collecting user behavior data and user portraits, we will de-identify the data through technical means.
[0135] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A data drift detection method, It is characterized in that include: Acquire user data in multiple scenarios of the current business, and divide the user data into a training data set and a test data set according to a preset ratio; Inputting each test data in the test data set into a pre-trained prediction model, respectively, so that the prediction model determines the prediction score of each test data corresponding to each scenario, wherein the prediction model is trained by the training data set; For each of the scenarios, the model index corresponding to the scenario is determined according to the predicted score of each of the test data corresponding to the scenario, and when it is determined that the model index is greater than an index threshold, it is determined that data drift has occurred in the user data corresponding to the scenario.
2. The data drift detection method according to claim 1, It is characterized in that The user data includes user data acquired in a first time period and user data acquired in a second time period, and the first time period is earlier than the second time period.
3. The data drift detection method according to claim 1, It is characterized in that After dividing the user data into a training data set and a test data set according to a preset ratio, the method further includes: Determining a statistical value of each user feature in the user data based on each of the training data in the training data set, wherein the user feature includes a behavioral feature of the user; The user data is normalized according to the statistical value of each of the user characteristics to obtain normalized user data.
4. The data drift detection method according to claim 1, It is characterized in that Before inputting each test data in the test data set into the pre-trained prediction model respectively, the method further includes: A multi-task learning model is constructed according to the number of scenarios in the current business.
5. The data drift detection method according to claim 4, It is characterized in that Before inputting each test data in the test data set into the pre-trained prediction model respectively, the method further includes: The multi-task learning model is trained based on the training data set to obtain the prediction model.
6. The data drift detection method according to claim 5, It is characterized in that The user data includes user features, scene labels and time labels. Accordingly, the multi-task learning model is trained based on the training data set to obtain the prediction model, including: Inputting each of the training data in the training data set into the multi-task learning model, so that the multi-task learning model determines the prediction score of each of the training data corresponding to each of the scenarios; Determining a loss function for each of the scenarios according to the time label and the prediction score of each of the training data corresponding to each of the scenarios; When it is determined that the loss functions of each of the scenarios have converged, the number of training times is greater than a number threshold, or the model index is greater than an index threshold, the prediction model is obtained.
7. The data drift detection method according to claim 6, It is characterized in that Determining a model indicator corresponding to the scenario according to the prediction score of each of the test data corresponding to the scenario includes: Determine the model metrics corresponding to the scenario based on the prediction scores and the time tags of the test data corresponding to the scenario.
8. The data drift detection method according to claim 2, wherein, after determining that the user data corresponding to the scenario has data drift, it further includes: For the user data corresponding to the scenario, sort the user data obtained in the first time period based on the prediction scores of the user data, and perform equal-frequency binning according to the sorting result; Determine a score threshold according to the binning result, and determine the user data with the prediction score greater than the score threshold as the target user data corresponding to the scenario.
9. A data drift detection device, wherein, it includes: An acquisition module, configured to acquire user data in multiple scenarios of the current service, and divide the user data into a training data set and a test data set according to a preset ratio; A prediction module, configured to input each test data in the test data set into a pre-trained prediction model respectively, so that the prediction model determines the prediction scores of each test data corresponding to each scenario, wherein the prediction model is trained by the training data set; An execution module, configured to, for each scenario, determine the model metrics corresponding to the scenario according to the prediction scores of each test data corresponding to the scenario, and determine that the user data corresponding to the scenario has data drift when it is determined that the model metrics are greater than the metric threshold.
10. A computer device, wherein, the computer device includes: One or more processors; A storage device, configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the data drift detection method according to any one of claims 1-8.
11. A storage medium containing computer-executable instructions, wherein, the computer-executable instructions are used to execute the data drift detection method according to any one of claims 1-8 when executed by a computer processor.