Concept drift detection and adaptation method, system and device based on deep learning model and storage medium

Through calibration of deep learning models and updating adaptive loss function, the problem of concept drift in network intrusion detection system is solved, efficient and automatic distribution adaptation and detection is achieved, and false alarm rate and marking overhead are reduced.

CN120378203APending Publication Date: 2025-07-25XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510711662.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When facing concept drift, the prior art is difficult to automatically and effectively detect and adapt to distribution changes in the network intrusion detection system, resulting in high false alarm rates, difficulty in updating and large marking overhead.

Method used

By calibrating the deep learning model, fitting the segmented function using isotonic regression, combining optimization problems and hypothesis testing, detecting distribution offsets, and updating model parameters through adaptive loss functions, reducing marking overhead, and achieving automatic adaptation.

Benefits of technology

It realizes efficient detection and adaptation to concept drift under unsupervised conditions, reduces manual intervention, improves the generalization ability and update efficiency of the model, and avoids catastrophic forgetting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378203A_ABST
    Figure CN120378203A_ABST
Patent Text Reader

Abstract

The invention relates to the field of network intrusion detection, in particular to a concept drift detection and adaptation method, system and device based on a deep learning model and a storage medium. Comprising the steps of calibrating a deep learning model, performing hypothesis testing on sample data distribution, solving an optimization problem and updating the deep learning model. According to the method, model output is converted in an unsupervised mode so as to better represent data distribution, and then detection offset is counted by testing the hypothesis of model output distribution; the problem of concept drift detection of the neural network model can be solved, and various concept drifts generated by a complex system are explained and adapted, so that the method has good research value and application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network intrusion detection, and particularly to a concept drift detection and adaptation method, system, device and storage medium based on a deep learning model. Background Art

[0002] In recent years, the adoption of deep learning has enabled anomaly detection to extract more complex features from massive data and detect unforeseen threats such as zero-day attacks by learning only from normal data (referred to as zero-positive learning). So far, researchers have applied deep learning-based anomaly detection to various security applications, such as detecting network intrusions, finding threats from system logs, tracking advanced persistent threats (APTs), etc., and have achieved remarkable results.

[0003] Currently, the performance of learning-based applications is based on the closed-world assumption of independent and identical distribution between training samples and test samples. Due to the difference between the incoming test distribution and the original distribution (referred to as concept drift), this assumption usually does not hold in an open-world setting. In the security field, concept drift is common because malicious patterns suddenly and sharply change over time in a hostile environment. If there are concept drift problems or system anomalies, huge losses will be brought. Moreover, it is inevitable for complex and large-scale systems to have concept drift that developers have not discovered. The concept drift of these systems has become a non-negligible security risk. Once discovered and maliciously exploited by attackers, it will bring unimaginable losses to service providers, users and even society.

[0004] Currently, there are two methods to solve the concept drift problem: The first is to retrain the model regularly in a dynamic environment without much consideration of whether, when and how the drift occurs. However, this method is not suitable in the security field because continuous training is labor-intensive for labeling. In addition, it is difficult to determine when to update the model. Delayed updates will expose the model to new threats or overwhelming false alarms. Moreover, security practitioners need to obtain evidence and explanations for the drift behind the black-box model.

[0005] The second is to identify and adapt to the drift. However, most current research by researchers explores concept drift in the context of supervised learning rather than concept drift in zero-positive anomaly detection. In addition, most focus on detection rather than adaptation, resulting in a high labeling overhead when preparing samples for adaptation. Moreover, previous research prefers to adopt a sample-level method by detecting whether a certain sample is unevenly distributed. This method cannot understand the drift at the distribution level, which leads to poor generalization ability of the drift after adaptation. Summary of the Invention

[0006] In view of the problems mentioned in the prior art, the present invention proposes a concept drift detection and adaptation method, system, device and storage medium based on a deep learning model, which can detect anomalies or concept drifts by processing the output of a neural network, and further be used for the defense against network intrusion and the adaptation to concept drift.

[0007] To achieve the above object, the present invention adopts the following technical solutions: The present invention provides a concept drift detection and adaptation method based on a deep learning model, comprising the following steps: S1. Calibrate the obtained deep learning model. First, define the expected confidence level, select a piecewise linear function for parameter estimation according to the expected confidence level, and use isotonic regression to fit the piecewise function to obtain the true confidence level; S2. Select sample data from the true confidence level, conduct a hypothesis test on the sample data distribution to determine whether two sample data follow the same distribution. If they follow the same distribution, end; if not, proceed to the next step; S3. Construct an optimization problem based on the sample data, solve the optimization problem to obtain samples of the new distribution, and perform artificial interpretation on the samples of the new distribution to obtain the final interpretation; S4. Input the final interpretation into the deep learning model to complete the adaptation of the deep learning model.

[0008] As a further improvement of the present invention, the specific process of step S1 is: Set the perfect calibration of the deep learning model as the detection threshold to calculate the expected confidence level in normal samples, and the expression is as follows:

[0009] In the formula: represents model calibration; is the set of all samples; is a sample in the set of all samples; is the output of the deep learning model; is the detection threshold; is the false positive rate in normal samples after setting itself as the detection threshold.

[0010] As a further improvement of the present invention, the specific process of step S2 is: Select the old samples and the new samples from the true confidence level, calculate the discrete distributions of the old samples and the new samples through the frequency histograms of K squares, and use the KL divergence as the hypothesis test statistic to test the new samples and the old samples Whether they follow the same distribution. If they do not follow the same distribution, there is concept drift or anomaly. If they follow the same distribution, there is no concept drift or anomaly. As a further improvement of the present invention, step S3 specifically includes the following steps: When the conditions for constructing the optimization problem are satisfied, construct the optimization problem. The expression of the optimization problem is as follows:

[0011] In the formula: is the precision loss; is the overhead loss; is the certainty loss; is vector of; is vector of; is the dimensional component of is an indicator for selecting whether; is the dimensional component of is an indicator for selecting whether; is the weighting parameter; is the weighting parameter; Use the gradient descent method to solve the optimization problem to obtain samples of the new distribution, and use manual interpretation to process the samples of the new distribution to obtain the final interpretation.

[0012] As a further improvement of the present invention, the expressions of the three loss functions in the optimization problem are as follows:

[0013] In the formula: is the new distribution after shifting; is the new distribution after shifting; is the distance between the true new distribution and the reconstructed distribution; Convert the calibration output into a vector of M relative frequencies in the histogram; is the number of original data samples; is the number of new data samples; is the (row-by-row) concatenation of two (row) vectors; is the (element-by-element) product.

[0014] As a further improvement of the present invention, the specific process of step S4 is: Input the final into the deep learning model, and update the deep learning model according to the adaptive new loss function of the deep learning model. The expression of the adaptive new loss function is as follows:

[0015] In the formula: is the new loss function of the deep learning model; is the original loss function of the deep learning model; is the weighting parameter; is the model parameter; is the importance weight.

[0016] As a further improvement of the present invention, different weights are assigned according to the importance of different samples in the sample data The assignment expression is as follows:

[0017] In the formula: is the squared norm of the deep learning output; is the importance weight; is the old sample; is vector of.

[0018] A concept drift detection and adaptation system based on a deep learning model, comprising: A calibration module that calibrates the obtained deep learning model. First, define the expected confidence level, select a piecewise linear function for parameter estimation according to the expected confidence level, and use isotonic regression to fit the piecewise function to obtain the true confidence level; A determination module that selects sample data from the true confidence level, conducts a hypothesis test on the sample data distribution, and determines whether two sample data follow the same distribution. If they follow the same distribution, the process ends; if they do not follow the same distribution, proceed to the next step; An interpretation module that constructs an optimization problem based on the sample data, solves the optimization problem to obtain samples of the new distribution, and conducts an artificial interpretation of the samples of the new distribution to obtain the final interpretation; An adaptation module that inputs the final interpretation into the deep learning model to complete the adaptation of the deep learning model.

[0019] A concept drift detection and adaptation device based on a deep learning model, comprising a processor and a memory. Among them, when the processor executes the computer program stored in the memory, it implements the concept drift detection and adaptation method based on the deep learning model as described above.

[0020] A computer-readable storage medium for storing a computer program, wherein when the computer program is executed by a processor, it implements the concept drift detection and adaptation method based on the deep learning model as described above.

[0021] The present invention has achieved the following technical effects compared with the prior art: The present invention transforms the model output in an unsupervised manner to better represent the data distribution, and statistically detects the drift through hypothesis testing of the model output distribution; being unsupervised is more adaptable to actual usage scenarios, can be applied in complex systems in practice, and can better combine with various neural networks.

[0022] The present invention constructs an optimization problem to find as few drifting samples as possible. In this way, only the most influential samples need to be labeled, thereby reducing the labeling overhead and manual intervention, so that the present invention can achieve relatively automatic drift interpretation and adaptation.

[0023] The present invention assigns different importance weights to model parameters to indicate whether they contain valid distributions; during the adaptation process, updating of important model parameters is restricted to avoid forgetting old valid distributions, while updating of unimportant parameters is supported to generalize to new distributions, ensuring the efficiency of updating while avoiding catastrophic forgetting. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is a schematic flow diagram of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings, rather than all the structures.

[0025] As Figure 1 shown, a concept drift detection and adaptation method based on a deep learning model of the present invention includes the following steps: S1. Calibrate the obtained deep learning model. First, define the expected confidence level, select a piecewise linear function for parameter estimation according to the expected confidence level, and use isotonic regression to fit the piecewise function to obtain the true confidence level; S2. Select the sample data in the true confidence level, conduct a hypothesis test on the sample data distribution to determine whether two sample data follow the same distribution. If they follow the same distribution, end; if not, proceed to the next step; S3. Based on the sample data, construct an optimization problem, solve the optimization problem to obtain the samples of the new distribution, and conduct an artificial interpretation on the samples of the new distribution to obtain the final interpretation; S4. Input the final interpretation into the deep learning model to complete the adaptation of the deep learning model.

[0026] The present invention will be further explained below with reference to the drawings and specific embodiments: Step 1. Calibration of Model Output Confidence calibration is a technique in the machine learning and deep learning communities that enables the calibrated model output to represent the true probability. Intuitively, calibration is to find a function that maps the original output to the desired confidence level.

[0027] However, existing research calibrates supervised classifiers, as fully labeled binary samples are required to calculate the true probability, such as accuracy ( ACC ), which cannot be calculated when only using normal samples for anomaly detection. Therefore, we propose an unsupervised calibration method that does not require any prior knowledge of anomalies.

[0028] The calibration method of the present invention has the following three requirements: 1) Non-linearity: Linear transformation is meaningless because it cannot change the density distribution of the original output.

[0029] 2) Legality: The calibrated output must be within the range of [0,1] to represent probability; formally, .

[0030] 3) Monotonicity: The calibrated output cannot change the original order, otherwise the performance will be changed; formally, .

[0031] In view of the above considerations, a novel calibrator is proposed that can convert the model output into the expected confidence level using only normal data. Without loss of generality, for a normal confidence model, we use the false positive rate (FPR) to define the expected normal confidence level after calibration.

[0032] Since FPR is the ratio of false positives ( ) to negatives ( ), no anomalies are required; anomalies are determined by comparing the model output and the detection threshold . Therefore, it is necessary to determine to calculate and . Taking as all normal data for calibration, we set to when calibrating , so the following definition (perfect calibration of the normal confidence anomaly detection model) is given. The perfect calibration of the normal confidence output is defined as the FPR in normal samples after setting itself as the detection threshold, and the expression is as follows:

[0033] In the formula: represents model calibration; is the set of all samples; is a sample in the set of all samples; is the output of the deep learning model; is the detection threshold; is the false positive rate in normal samples after setting itself as the detection threshold.

[0034] After defining the expected confidence, a parametric function is selected and parameter estimation is performed to reduce the error between the calibrated output and the expected confidence. In this embodiment, a piecewise linear function (PWLF) is selected as a more general basis function for security applications.

[0035] To meet the three requirements of calibration mentioned above, isotonic regression is used to fit the piecewise function. To meet the legality, we set the lower and upper bounds of the fitted function to 0 and 1; in other words, the prediction will be clipped to the nearest fitted interval endpoint. Briefly, isotonic regression solves the following quadratic programming given an observed sequence as follows:

[0036] where: is the true confidence; is the expected confidence.

[0037] The observations of the fitted calibrator are as follows. For each old feature for the uncalibrated output sort them in ascending order and regard them as ; here represents as a set of , and the corresponding is calculated as the expected confidence .

[0038] In this work, the detection of distribution shift is to compare the old and features for . Using the calibrator, the model output essentially contains probability information to facilitate the statistical detection of distribution shift. Therefore, we further transform the shift detection problem into comparing for .

[0039] Step 2. Determine the shift To compare distributions, usually considering the overhead of collection and labeling, is set to be the same as the size of the training data of the anomaly detection model in this study, ≤ . Note that should be normal because they have been labeled in the previous round. However, since it cannot be ensured thatAll are normal because they are newly sampled from the current environment. Therefore, further processes will involve manual investigation to filter out anomalies.

[0040] The discrete distribution of is calculated through a frequency histogram with bins.

[0041] The null hypothesis is that come from the same distribution (no shift), while the alternative hypothesis is the opposite (shift). The test statistic is the Kullback-Leibler (KL) divergence between the two distributions. For two discrete probability distributions and defined on the same probability space, the KL divergence is defined as .

[0042] Step 3. Explain the shift This embodiment proposes an interpreter in the distribution-level shift explanation method OWAD to find important samples that cause concept shifts, helping security practitioners understand and adapt to the shifts. To provide better explanation results, three aspects need to be considered in optimization: 1) Explanation accuracy: The samples selected for explanation should accurately represent the shifted distribution.

[0043] 2) Labeling overhead: Since manual labeling is required, it is expected that the number of samples selected from the new space is small. 3) Explanation determinism: The explanation should be deterministic (that is, not ambiguous for the selection).

[0044] The above three requirements must be considered simultaneously, and the first two are contradictory (for example, introducing more new samples will increase the explanation accuracy but also increase the labeling overhead). Therefore, they need to be considered simultaneously in the objective function to find the best compromise.

[0045] Introduce two vectors and corresponding to and . The i-th dimensional components of and represented by and are indicators of whether to select and (1 or 0). For the convenience of optimization, we will and The values are relaxed to the range [0, 1] and they are binaryized to 0 / 1 for sample selection. The problem is formulated as:

[0046] where: is the precision loss; is the overhead loss; is the certainty loss; is 's vector; is 's vector; is 's -th dimensional component which is an indicator for selection yes / no; is 's -th dimensional component which is an indicator for selection yes / no; is the weighting parameter; is the weighting parameter.

[0047] where:

[0048] where: is the new shifted distribution; is the new shifted distribution; is the distance between the true new distribution and the reconstructed distribution; converts the calibrated output to a vector of M relative frequencies in the histogram; is the number of original data samples; is the number of new data samples; is the (row-wise) concatenation of two (row) vectors; is the (element-wise) product.

[0049] The optimization objective includes the above three terms. For security operators with different requirements in different applications, they are weighted by two customizable hyperparameters and For example, when the marking ability is limited, the operator can increase , while for applications sensitive to errors, can be decreased.

[0050] In the precision term in Equation 4, where and . Intuitively, it is to "reconstruct" the new shifted distribution with the selected and , similar to shift detection, using To measure the distance between the true new distribution and the reconstructed distribution. Convert the calibration output into a vector of M relative frequencies in the histogram. Compared with K used for shift detection, it is necessary to ensure that M >> K to better interpret and reconstruct the fine-grained distribution.

[0051] In the overhead term in Equation 5 . Note that since has been marked as normal before, while manual investigation is required to filter out anomalies. Therefore, we choose as many and as few to reduce the labeling overhead. Specifically, we measure 1 - and of norm, and then divide it by the number of samples (i.e., + ).

[0052] In the deterministic term in Equation 6 . To avoid the ambiguity of sample selection after optimization, we expect and to be deterministic (i.e., close to 0 or close to 1), and measure the entropy of or (referred to as uncertainty).

[0053] Adopt the gradient descent method to solve the above optimization and randomly initialize / in [0, 1]. Note that although in the precision term is non-differentiable, we only use this operation once before optimization because or (calculated from or ) bins are fixed during optimization. And manually include the samples outside the boundary (compared with the old set) in the final interpretation, that is, select the samples in the processing set whose model output probabilities are lower than the minimum value or higher than the maximum value of the old set, and add these samples to the final interpretation.

[0054] Step 4. Shift Adaptation After filtering out anomalies through manual investigation and providing important samples, the operator can gain insights from our interpretation; if concept shift is confirmed, the anomaly detection model needs to be updated to adapt to the concept drift to avoid performance degradation.

[0055] This embodiment proposes an adaptive method to prevent catastrophic forgetting of effective knowledge while ensuring generalization to new distributions. By adding a special regularization term to the original loss function during the model update process. Different from Norm regularization assigns different importance weights to each model parameter represented by before the update, denoted by . The intuitive meaning of is to evaluate the importance of each parameter to knowledge. In this way, the update of important parameters can be restricted to prevent catastrophic forgetting, while the regularization of unimportant parameters can be relaxed to allow the model to adapt to the new distribution. Specifically, the new loss function

[0056] of the model adaptation is: where is the new loss function of the deep learning model; is the weighted parameter; is the model parameter; is the importance weight.

[0057] In the above formula is the non - negative gradient evaluated using the labeled samples of the old distribution . In this method, old samples are randomly selected and equally considered for evaluation . However, in the new distribution, some old samples are outdated. Considering these samples will prevent the model from forgetting the outdated samples, the data knowledge in the old distribution, and thus it cannot effectively adapt to the new distribution.

[0058] Therefore, it is necessary to assign importance weights to different samples, and the importance weights here are exactly obtained from the OWAD interpreter (before binarization); lower mask values in the old set indicate that the samples are not selected to represent the new distribution. Therefore, the intuition is that the model will forget the outdated samples with lower weights / masks, but remember the outdated samples with higher weights / masks. Specifically, the old features weighted by from the interpreter are used to evaluate RPMI. The expression is as follows:

[0059] where is the squared norm of the deep learning output; is the importance weight; is the old sample; is vector.

[0060] According to the new loss function of the model adaptation, and selected after shift interpretation and manual investigation are adopted.Update the model to achieve the adaptation of the deep learning model.

[0061] Based on the same inventive concept, an embodiment of the present invention further provides a concept drift detection and adaptation system based on a deep learning model. Since the principle of solving problems by this concept drift detection and adaptation system based on a deep learning model is similar to that of the aforementioned concept drift detection and adaptation method based on a deep learning model, the implementation of this concept drift detection and adaptation system based on a deep learning model can refer to the implementation of the concept drift detection and adaptation method based on a deep learning model, and the repeated parts will not be elaborated.

[0062] In specific implementation, the concept drift detection and adaptation system based on a deep learning model provided by an embodiment of the present invention specifically includes: A calibration module that calibrates the obtained deep learning model. First, define the expected confidence level, select a piecewise linear function for parameter estimation according to the expected confidence level, and use isotonic regression to fit the piecewise function to obtain the true confidence level; A determination module that selects sample data from the true confidence level, conducts a hypothesis test on the sample data distribution to determine whether two sample data follow the same distribution. If they follow the same distribution, end; if they do not follow the same distribution, proceed to the next step; An interpretation module that constructs an optimization problem based on the sample data, solves the optimization problem to obtain samples of the new distribution, and conducts artificial interpretation on the samples of the new distribution to obtain the final interpretation; An adaptation module that inputs the final interpretation into the deep learning model to complete the adaptation of the deep learning model.

[0063] Correspondingly, an embodiment of the present invention further provides a concept drift detection and adaptation device based on a deep learning model, including a processor and a memory. Among them, when the processor executes the computer program saved in the memory, it implements the concept drift detection and adaptation method based on a deep learning model provided by an embodiment of the present invention.

[0064] For a more specific process of the above method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.

[0065] Correspondingly, an embodiment of the present invention further provides a computer-readable storage medium for storing a computer program. Among them, when the computer program is executed by a processor, it implements the above-mentioned concept drift detection and adaptation method based on a deep learning model provided by an embodiment of the present invention.

[0066] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the systems, devices, and storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0067] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0068] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0069] Finally, it should also be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, the element defined by the statement "including an..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0070] The above has introduced in detail the concept drift detection and adaptation method, system, device and storage medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A concept drift detection and adaptation method based on a deep learning model, characterized in that, It includes the following steps: S1. Calibrate the obtained deep learning model. First, define the expected confidence level, select a piecewise linear function for parameter estimation according to the expected confidence level, and use isotonic regression to fit the piecewise function to obtain the true confidence level; S2. Select the sample data from the true confidence level, conduct a hypothesis test on the sample data distribution to determine whether the two sample data follow the same distribution. If they follow the same distribution, end; if not, proceed to the next step; S3. Construct an optimization problem based on the sample data, solve the optimization problem to obtain the sample of the new distribution, and perform manual interpretation on the sample of the new distribution to obtain the final interpretation; S4. Input the final interpretation into the deep learning model to complete the adaptation of the deep learning model.

2. The concept drift detection and adaptation method based on a deep learning model according to claim 1, wherein The specific process of step S1 is as follows: Set the perfect calibration of the deep learning model as the detection threshold to calculate the expected confidence level in the normal samples. The expression is as follows: In the formula: represents model calibration; is the set of all samples; is a sample in the set of all samples; is the output of the deep learning model; is the detection threshold; is the false positive rate in normal samples after setting itself as the detection threshold.

3. The concept drift detection and adaptation method based on a deep learning model according to claim 1, characterized in that The specific process of step S2 is as follows: Select old samples from the true confidence and new samples , calculate the discrete distributions of the old samples and new samples through the frequency histograms of K squares, use the KL divergence as the hypothesis test statistic, and determine whether the new samples and the old samples follow the same distribution. If they do not follow the same distribution, there is concept drift or anomaly. If they follow the same distribution, there is no concept drift or anomaly.

4. The concept drift detection and adaptation method based on a deep learning model according to claim 1, wherein, Step S3 specifically includes the following steps: When the conditions for constructing the optimization problem are met, construct the optimization problem. The expression of the optimization problem is as follows: Where: is the precision loss; is the overhead loss; is the certainty loss; is a vector of; is a vector of; is the -th dimensional component of is an indicator for selection whether; is the -th dimensional component of is an indicator for selection whether; is a weighting parameter; is a weighting parameter; Use the gradient descent method to solve the optimization problem to obtain the sample of the new distribution, and use manual interpretation to process the sample of the new distribution to obtain the final interpretation.

5. The concept drift detection and adaptation method based on a deep learning model according to claim 4, characterized in that, The expressions of the three loss functions in the optimization problem are as follows: Wherein: is the new distribution after shifting; is the new distribution after shifting; is the distance between the true new distribution and the reconstructed distribution; converts the calibration output into a vector of M relative frequencies in the histogram; is the number of original data samples; is the number of new data samples; is the (row-by-row) concatenation of two (row) vectors; is the (element-by-element) product.

6. A concept drift detection and adaptation method based on a deep learning model according to claim 1, characterized in that The specific process of step S4 is as follows: Input the final into the deep learning model, and update the deep learning model according to the adaptive new loss function of the deep learning model. The expression of the adaptive new loss function is as follows: In the formula: is the new loss function of the deep learning model; is the original loss function of the deep learning model; is the weighting parameter; is the model parameter; is the importance weight.

7. A concept drift detection and adaptation method based on a deep learning model according to claim 6, characterized in that, According to the importance of different samples in the sample data different weights are assigned, and the assignment expression is as follows: In the formula: is the squared norm of the deep learning output; is the importance weight; is the old sample; is vector of 8. A concept drift detection and adaptation system based on a deep learning model, characterized in that, It includes: A calibration module that calibrates the obtained deep learning model. First, define the expected confidence level, select a piecewise linear function for parameter estimation according to the expected confidence level, and use isotonic regression to fit the piecewise function to obtain the true confidence level; A determination module that selects the sample data from the true confidence level, conducts a hypothesis test on the sample data distribution to determine whether the two sample data follow the same distribution. If they follow the same distribution, end; if not, proceed to the next step; An interpretation module that constructs an optimization problem based on the sample data, solves the optimization problem to obtain the sample of the new distribution, and performs manual interpretation on the sample of the new distribution to obtain the final interpretation; An adaptation module that inputs the final interpretation into the deep learning model to complete the adaptation of the deep learning model.

9. A concept drift detection and adaptation device based on a deep learning model, characterized in that, It includes a processor and a memory. Among them, when the processor executes the computer program saved in the memory, it implements the concept drift detection and adaptation method based on the deep learning model according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It is used to store a computer program. Among them, when the computer program is executed by the processor, it implements the concept drift detection and adaptation method based on the deep learning model according to any one of claims 1 to 7.