Drift detection method and system for AI data sets
By using the least squares density difference algorithm and the constraints of the Gaussian kernel model in the drift detection of AI data sets, the problem of insufficient accuracy when detecting image data sets compressed is solved, and higher detection accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202410777312.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-06-17
AI Technical Summary
The existing AI dataset drift detection methods are not accurate enough when detecting compressed image data sets, resulting in inaccurate detection results, affecting the evaluation of the data set.
A drift detector containing constraints is constructed using the least squares density difference algorithm. The estimation accuracy of the density difference function is improved by using the constraints whose properties difference between the Gaussian kernel model and the real density difference function whose set threshold value is less than the constraints.
By setting constraints, the density difference function is estimated more accurately, thereby improving the accuracy of drift detection, reducing misjudgment, and improving the reliability of detection results.
Smart Images

Figure CN118797750B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer security technology, and in particular to a drift detection method and system for an AI data set. Background Art
[0002] In the process of data collection and processing for machine learning and deep learning, due to improper data collection methods, inconsistent measurement specifications, and unintentional / malicious human introduction, AI data sets are doped with abnormal data or malicious adversarial sample data, which in turn leads to model training errors. In the process of generating abnormal data, the phenomenon that the statistical properties of the target data set change in an arbitrary manner over time is usually called concept drift. The detection of drift mainly uses the hypothesis testing method, which determines whether to accept / reject the original hypothesis by calculating the value of the test statistic and manually setting or dynamically adjusting the threshold.
[0003] However, in the actual detection process (taking the detection of image datasets as an example here), for specific compressed images, the statistical properties of the image dataset have obviously changed. However, existing detectors often find it difficult to detect drift due to inaccurate initial parameter settings or detection function settings, which leads to inaccurate detection results and affects the evaluation of the dataset. Summary of the invention
[0004] In view of the above problems, the present invention provides a drift detection method and system for an AI data set, which can effectively perform drift detection on various types of AI data sets.
[0005] To achieve the purpose of the invention, the technical solution of the present invention includes the following contents.
[0006] A drift detection method for an AI data set, the method comprising:
[0007] Based on the least squares density difference algorithm, a drift detector including constraint conditions is constructed; wherein the constraint condition is that the property difference between the Gaussian kernel model and the true density difference function is less than a set threshold;
[0008] Encode the data to be tested and the reference data;
[0009] Based on the encoding results of the data to be detected and the reference data, the drift detector is used to calculate the value L of the least square density difference between the data set to be detected and the reference data set. original ;
[0010] By randomly replacing the labels of the reference data and the data to be detected, a plurality of replacement samples are generated, and the value L of the least squares density difference between each replacement sample and the reference data set is calculated based on the drift detector. i , i is a positive integer;
[0011] Based on the value L of the least squares density difference original and the value of the least squares density difference L i , and obtain the drift detection result of the data set to be detected.
[0012] Furthermore, when encoding the constraint condition, the constraint condition is implemented by setting the mean value of each component of the high-dimensional parameter tensor in the Gaussian kernel model to 0.
[0013] Furthermore, the value L based on the least squares density difference original and the value of the least squares density difference L i , and obtain the drift detection results of the data set to be detected, including:
[0014] According to the value of the least squares density difference L i , and obtain the permutation distribution;
[0015] By calculating the value of the least squares density difference L original The percentile in the permutation distribution gives the p-value;
[0016] The p-value is compared with the pre-set significance level to obtain the drift detection result of the data set to be detected.
[0017] Furthermore, the value L based on the least squares density difference original and the value of the least squares density difference L i , and obtain the drift detection results of the data set to be detected, including:
[0018] The value of the least squares density difference L i Sort and get a LSDD list;
[0019] Determining an index position in the LSDD list according to a preset significance level;
[0020] Get the value of the corresponding index position in the LSDD list as the drift detection threshold;
[0021] Compare the value L of the least squares density difference original and the drift detection threshold to obtain a drift detection result of the data set to be detected.
[0022] Further, the data to be detected includes: image data;
[0023] The encoding of the data to be detected comprises:
[0024] An encoder is defined, wherein the network structure of the encoder includes a plurality of convolutional layers and fully connected layers;
[0025] The image data is encoded into a low-dimensional representation based on the encoder.
[0026] Furthermore, the data set to be detected includes: text data;
[0027] The encoding of the data to be detected comprises:
[0028] Use the pre-trained BERT model to obtain text data for word segmentation;
[0029] The word segmentation results are fed into the Transformer-based neural network model to obtain the encoding results of the text data.
[0030] A drift detection system for an AI data set, the system comprising:
[0031] A construction module, for constructing a drift detector including a constraint condition based on a least squares density difference algorithm; wherein the constraint condition is that the property difference between the Gaussian kernel model and the true density difference function is less than a set threshold;
[0032] A preprocessing module, used for encoding the data to be detected and the reference data;
[0033] A detection module is used to calculate the value L of the least square density difference between the data set to be detected and the reference data set based on the encoding results of the data to be detected and the reference data set using the drift detector original ; Generate multiple replacement samples by randomly replacing the labels of the reference data and the data to be detected, and calculate the value L of the least squares density difference between each replacement sample and the reference data set based on the drift detector i , i is a positive integer; based on the value L of the least squares density difference original and the value of the least squares density difference L i , and obtain the drift detection result of the data set to be detected.
[0034] An electronic device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the drift detection method of the AI data set described in any one of the above items is implemented.
[0035] A computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement any of the above-mentioned methods for detecting drift of an AI data set.
[0036] A computer program product, when the computer program product is run on a computer device, enables the computer device to execute any of the above-mentioned methods for detecting drift of an AI data set.
[0037] Compared with the prior art, the present invention can estimate the density difference function more accurately by setting constraint conditions, thereby improving the accuracy of drift detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Flowchart of the drift detection method for AI datasets. DETAILED DESCRIPTION
[0039] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present invention is described and illustrated below in conjunction with the accompanying drawings.
[0040] The drift detection process of the present invention is as follows: Figure 1 As shown, it can be divided into three steps: constructing a drift detector, preprocessing input data, and generating drift detection results.
[0041] Step 1: Construct a drift detector.
[0042] 1) Theoretical background of least squares density difference (LSDD).
[0043] In order to better illustrate the drift detector constructed in the present invention, the principle of the existing least squares density difference method is firstly explained.
[0044] Assume there are two sets of independent and identically distributed samples and samples From the probability distributions with densities p(x) and p′(x), we can get (x∈R d ),Right now:
[0045]
[0046] The goal of LSDD is to estimate the sample and The difference f(x) between p(x) and p′(x) in
[0047] f(x): = p(x)-p′(x).
[0048] In LSDD, the density difference model g(x) is fitted to the true density difference function f(x) under squared loss:
[0049]
[0050] Use the following linear parameter model as g(x):
[0051]
[0052] Where b represents the number of basis functions, ψ(x)=(ψ1(x),...,ψ b (x) Tis a b-dimensional linearly independent basis function vector, θ = (θ1, ..., θ b ) T is a b-dimensional parameter vector.
[0053] In practice, the following Gaussian kernel model is usually used as the linear parameter model g(x):
[0054]
[0055] Among them, (c1, ..., c n , c n,1 , ..., c n+n′ ): =(x1, ...x n , x′1,...,x′ n′ ) is the center of the Gaussian kernel. If n+n′ is large, only x1, ...x n , x′1,...,x′ n′ A subset of is taken as the Gaussian kernel center.
[0056] For the linear parameter model (1), the optimal parameters can be given by
[0057]
[0058] Where H is a b×b matrix and h is a b-dimensional vector. The definitions are as follows
[0059] H: =∫ψ(x)ψ(x) T dx,
[0060] h: =∫ψ(x)p(x)dx-∫ψ(x′)p′(x′)dx′.
[0061] Note that for the Gaussian kernel model (2), the integral in H can be calculated analytically as
[0062]
[0063] Where d represents the dimension of x. Using the empirical estimator instead of the expected value and adding an l2 regularizer to the objective function, we get the following optimization problem:
[0064]
[0065] where λ(>0) is the regularization parameter, is a b-dimensional vector. The definitions are as follows
[0066]
[0067] Taking the derivative of the above objective function and setting it to 0, we can get the analytical solution
[0068]
[0069] Among them, H λ =H+λI b , I b represents the b-order identity matrix.
[0070] Finally, the density difference estimator is given for:
[0071]
[0072] Then LSDD can be estimated as
[0073]
[0074] This is called the least squares density difference (LSDD) estimator. The LSDD estimator has good theoretical properties: it has an asymptotic order of The asymptotic normality of is considered to be the optimal convergence rate in the parameter setting. In addition, the nonparametric LSDD estimator also achieves the optimal convergence rate.
[0075] 2) Implementation of drift detector
[0076] Since the initial form of LSDD is shown in formula (3), which is just a simple integral form, its model estimation effect is limited.
[0077] LSDD(p,q)=∫(p(x)-q(x)) 2 dx (3)
[0078] Therefore, the present invention adds constraints to make the model have a better estimation effect. Specifically, according to the definition, the true density difference function satisfies
[0079] ∫f(x)dx=∫p(x)dx-∫p′(x)dx=1-1=0
[0080] In the present invention, it is necessary to use g(x) to approximate the true density difference function, that is, to make the properties of g(x) as close to f(x) as possible, so it is best to make the integral of g(x) equal to 0. Under this constraint, the approximation of f(x) by g(x) must be more accurate than the case without restriction.
[0081] Step 2: Input data preprocessing.
[0082] The purpose of data preprocessing is to convert the image data and text data in the AI dataset into the data structure expected by the drift detector.
[0083] In one embodiment, the preprocessing of the image data includes the following steps 210 - 211 .
[0084] Step 211: Define the encoder.
[0085] A simple convolutional neural network encoder is defined using PyTorch's nn.Module. The encoder includes several convolutional layers and fully connected layers to encode the input image into a low-dimensional representation. The output dimension of the encoder is 32.
[0086] Step 212: Complete image data preprocessing through a defined preprocessing function.
[0087] The present invention first defines a preprocessing function preprocess_fn, which uses the preprocess_drift function, which preprocesses the input image after encoding it through the encoder. In this way, the input data can be converted into a low-dimensional representation of the encoder and can meet the needs of drift detection.
[0088] In one embodiment, the preprocessing of text data includes the following steps 221 to 224 .
[0089] Step 221: Load the pre-trained model and word segmenter.
[0090] The present invention loads the pre-trained BERT model's tokenizer through the AutoTokenizer.from_pretrained method, which is used to convert the input text into the input format required by the model.
[0091] Step 222: Define the embedder.
[0092] The present invention defines an embedder using the TransformerEmbedding class, which encodes the input text into the hidden state of the pre-trained Transformer model.
[0093] Step 223: Define a simple neural network model.
[0094] The neural network model includes an embedder, a linear layer, and an activation function. Its main function is to adjust the output dimension of the embedder to match the dimension required by the drift detector for subsequent data processing and analysis.
[0095] Step 224: Complete text data preprocessing through the defined preprocessing function.
[0096] The present invention defines a preprocessing function preprocess_fn, which encodes the input text as the input of the model through the embedder after the input text is segmented by the tokenizer, and performs other necessary preprocessing operations, such as truncating the text to a maximum length. The output of the preprocessing function is a data structure that matches the input format expected by the model.
[0097] Step 3: Drift detection result generation.
[0098] After the drift detector is constructed and the data set is preprocessed, the process of completing drift detection based on the drift detector of the present invention includes the following steps 301 to 305.
[0099] Step 301: Initialize the LSDD drift detector.
[0100] Step 302: The LSDD drift detector class declares the initial parameters of the detector, initializes the method of the Gaussian kernel model for density difference estimation and calculates the estimated value of LSDD, and returns the estimated value of LSDD, the p-value for evaluating the significance of the degree of data drift, and the drift detection threshold for determining whether drift has occurred (compared with the LSDD estimated value).
[0101] Step 303: Calculate the Gaussian kernel matrix H of the test data and the reference data (i.e., the matrix H mentioned when solving the optimal parameters). Apply the constraint conditions of the present invention to the parameter vector of the Gaussian kernel model. In the actual process of writing the code, the parameter vector here is replaced by a high-dimensional parameter tensor (actually a higher-dimensional data structure representation), so this constraint condition can be converted to set the mean of each component of the tensor to 0. Then call the permed_lsdds function to calculate the original LSDD (Least-squares density difference) value between the reference data set and the data set to be tested.
[0102] Step 304: Obtain the LSDD value of the replaced sample by randomly replacing the labels of the reference data and the test data.
[0103] permed_lsdds generates multiple permuted samples by randomly permuting the labels of reference data and test data, and calculates the LSDD value for each permuted sample.
[0104] Step 305: Based on the LSDD value of the replaced sample and the original LSDD value, determine whether data drift occurs.
[0105] Method 1: Compare the original LSDD value with the value in the permuted distribution, calculate the percentile of the original LSDD value in the permuted distribution, and obtain the p-value. Compare the p-value with the pre-set significance level to determine whether the drift is significant. The principle of this permutation is that if the original data set does have drift, the original LSDD value will usually be larger than most permuted LSDD values; if the original data set does not have significant drift, the original LSDD value should be similar to the distribution of the permuted LSDD value.
[0106] Method 2: Sort the permuted LSDD values to get an ordered list. According to the pre-set significance level p_val (for example, 0.05), determine the index position in the sorted list. This index position indicates the percentage of permutation experiment results that the LSDD value exceeds. Get the value of the corresponding index position in the sorted LSDD list as the drift detection threshold. If the LSDD estimate is greater than this threshold, the data has drifted significantly.
[0107] In summary, the detector finally returns the results including LSDD estimation, p-value and threshold for marking drift. These values can be used to evaluate whether the input data is significantly different from the reference data distribution, thereby determining whether data drift has occurred.
[0108] Next, a specific experiment is used to verify the drift detection method of the AI data set provided by the present invention.
[0109] This experiment compares the drift detector with constraints and the drift detector without constraints, and evaluates the two prediction results by calculating the accuracy, precision, recall and F1 score. Accuracy is the proportion of correctly predicted samples to the total samples, precision is the proportion of samples predicted to drift that actually drift, recall is the proportion of samples that actually drift that are correctly predicted to drift, and F1 score is the harmonic mean of precision and recall.
[0110] The input of this experiment is images of different corruption levels (including compressed images) in the CIFAR10 image dataset. The results are shown in Table 1. It can be seen that for the specified input, the drift detector with constraints achieves better technical results than the drift detector without constraints.
[0111] Evaluation Metrics Drift detector with constraints Drift detector without constraints Accuracy 0.947 0.923 Accuracy 1.000 1.000 Recall 0.944 0.917 F1 score 0.971 0.957
[0112] Table 1
[0113] The above implementation is only used to illustrate the technical solution of the present invention rather than to limit it. Ordinary technicians in this field can modify or replace the technical solution of the present invention with equivalents without departing from the scope of the present invention. The protection scope of the present invention shall be based on the claims.
Claims
1. A drift detection method for an AI data set, characterized in that: The method comprises: Based on the least squares density difference algorithm, a drift detector including constraint conditions is constructed; wherein the constraint condition is that the property difference between the Gaussian kernel model and the true density difference function is less than a set threshold; Encode the data to be tested and the reference data; Based on the encoding results of the data to be detected and the reference data, the drift detector is used to calculate the value L of the least square density difference between the data set to be detected and the reference data set. original ; By randomly replacing the labels of the reference data and the data to be detected, a plurality of replacement samples are generated, and the value L of the least squares density difference between each replacement sample and the reference data set is calculated based on the drift detector. i , i is a positive integer; Based on the value L of the least squares density difference original and the value of the least squares density difference L i , and obtain the drift detection result of the data set to be detected.
2. The method according to claim 1, characterized in that When encoding the constraint condition, the constraint condition is implemented by setting the mean of each component of the high-dimensional parameter tensor in the Gaussian kernel model to be zero.
3. The method according to claim 1, characterized in that The value L based on the least squares density difference original and the value of the least squares density difference L i , and obtain the drift detection results of the data set to be detected, including: According to the value of the least squares density difference L i , and obtain the permutation distribution; By calculating the value of the least squares density difference L original The percentile in the permutation distribution gives the p-value; The p-value is compared with the pre-set significance level to obtain the drift detection result of the data set to be detected.
4. The method according to claim 1, characterized in that The value L based on the least squares density difference original and the value of the least squares density difference L i , and obtain the drift detection results of the data set to be detected, including: The value of the least squares density difference L i Sort and get a LSDD list; Determining an index position in the LSDD list according to a preset significance level; Get the value of the corresponding index position in the LSDD list as the drift detection threshold; Compare the value L of the least squares density difference original and the drift detection threshold to obtain a drift detection result of the data set to be detected.
5. The method according to claim 1, characterized in that The data to be detected includes: image data; The encoding of the data to be detected comprises: An encoder is defined, wherein the network structure of the encoder includes a plurality of convolutional layers and fully connected layers; The image data is encoded into a low-dimensional representation based on the encoder.
6. The method according to claim 1, characterized in that The data set to be detected includes: text data; The encoding of the data to be detected comprises: Use the pre-trained BERT model to obtain text data for word segmentation; The word segmentation results are fed into the Transformer-based neural network model to obtain the encoding results of the text data.
7. A drift detection system for AI data sets, characterized in that: The system comprises: A construction module, for constructing a drift detector including a constraint condition based on a least squares density difference algorithm; wherein the constraint condition is that the property difference between the Gaussian kernel model and the true density difference function is less than a set threshold; A preprocessing module, used for encoding the data to be detected and the reference data; A detection module is used to calculate the value L of the least square density difference between the data set to be detected and the reference data set based on the encoding results of the data to be detected and the reference data set using the drift detector original ; Generate multiple replacement samples by randomly replacing the labels of the reference data and the data to be detected, and calculate the value L of the least squares density difference between each replacement sample and the reference data set based on the drift detector i , i is a positive integer; based on the value L of the least squares density difference original and the value of the least squares density difference L i , and obtain the drift detection result of the data set to be detected.
8. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the drift detection method for the AI data set as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the drift detection method for the AI data set according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that When the computer program product runs on a computer device, the computer device executes the drift detection method for an AI data set as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Conceptual drift-oriented adaptive interpretable industrial control system anomaly detection method
CN116991137A
Abnormal value interference resistant complex curved surface reconstruction method and system
CN118097025A