Knowledge drift detection method based on interaction matching

By applying interactive matching knowledge drift detection method in large-scale data, analyzing the quantity and quality changes of IF-THEN rules, solving the shortcomings in efficiency and accuracy of the existing technology, and achieving effective detection and characterization of knowledge drift.

CN120069047AInactive Publication Date: 2025-05-30NANCHANG NORMAL UNIV OF APPLIED TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411953901.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When processing large-scale data, existing concept drift detection methods are difficult to take into account efficiency and accuracy, and cannot fully reflect the essential characteristics of knowledge implicitly in the data.

Method used

A knowledge drift detection method based on interaction matching is proposed. By analyzing the quantity and quality changes of IF-THEN rules in the data, and combining the convergence characteristics of the rules credibility under the sampling background, a drift detection algorithm based on rule interaction matching is designed to adapt to the complexity and diversity of knowledge drift in large-scale data.

Benefits of technology

It significantly improves the accuracy and reliability of knowledge drift detection, can fully reflect the basic nature of knowledge drift, adapt to the needs of complex decision-making scenarios, and reduces the computational complexity of large-scale data set processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069047A_ABST
    Figure CN120069047A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge drift detection method based on interaction matching, and belongs to the technical field of data mining and artificial intelligence. The method aims at solving the technical problems that an existing method is lack of systematicness in description of knowledge features and not comprehensive in description of drift behaviors. According to the technical scheme, the method comprises the following steps of S1, expanding a sampling interaction matching knowledge drift detection model S-RKDD; s2, depicting forward and reverse drifts through quantity and quality changes of the rule base; and S3, by setting a rule threshold and a relaxation range, the complexity and diversity of knowledge drift in large-scale data are adapted. The method has the beneficial effects that the quantity and quality change of the IF-THEN rules in the data are analyzed, so that the knowledge drift of a large-scale data set is effectively detected and depicted, and the method is particularly used for the behavior analysis of forward and reverse rule drift.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of data mining and artificial intelligence, and particularly to a knowledge drift detection method based on interactive matching for knowledge change analysis and detection in large-scale datasets. Background Art

[0002] Existing concept drift detection methods usually rely on external features of data, but cannot fully reflect the essential knowledge features hidden in the data. Especially when dealing with large-scale data, it is difficult to balance efficiency and accuracy. The detection of knowledge drift is a key issue in data-driven decision-making, and existing methods have the following deficiencies: 1) The description of knowledge features lacks systematicness; 2) The characterization of drift behaviors is not comprehensive enough; 3) It is difficult to meet the requirements of large-scale data detection. To solve the above problems, the present invention proposes a knowledge drift detection method based on interactive matching.

[0003] How to solve the above technical problems is the subject faced by the present invention. Summary of the Invention

[0004] The purpose of the present invention is to provide a knowledge drift detection method based on interactive matching. Combining the core problem of data decision-making, that is, the knowledge contained in the data, and the deficiencies of existing drift detection methods, the present invention analyzes the quantity and quality changes of IF-THEN rules in the data to achieve effective detection and characterization of knowledge drift in large-scale datasets, especially for the behavior analysis of positive and negative rule drifts.

[0005] The inventive concept of the present invention is as follows: The present invention aims to construct a knowledge drift detection method based on rule interactive matching. By proposing an interactive matching knowledge drift description model (RKDD) with IF-THEN rules as the knowledge carrier; in the context of sampling, combining the convergence characteristics of rule knowledge, an extended sampling interactive matching knowledge drift detection model (S-RKDD) is developed; a drift detection algorithm based on rule interactive matching is designed to characterize positive and negative drifts through the quantity and quality changes of the rule base; finally, by setting rule thresholds and relaxation ranges, it adapts to the complexity and diversity of knowledge drift in large-scale data.

[0006] To achieve the above invention purpose, the technical solution adopted by the present invention is specifically: A knowledge drift detection method based on interactive matching, wherein, Ω 1 and Ω 2 are two datasets with the same formal structure, including the following steps:

[0007] S1. Divide the dataset Ω into a reference data window Ω 1 and a data window to be tested Ω 2, and determine the rule threshold β ∈ [0, 1] and its relaxation δ ∈ [0, 1].

[0008] S2. Randomly sample the data windows Ω 1 and Ω 2 n times to obtain a sample data set with a capacity of m, and calculate for each sampling result forward rule drift and backward rule drift as well as forward trust drift and backward trust drift

[0009] S3. According to the results of the n - time sampling detections, obtain the forward rule drift and backward rule drift as well as forward trust drift and backward trust drift and calculate bidirectional rule drift and bidirectional trust drift The symbol descriptions involved in the present invention are shown in the following table.

[0010] Table 1 Symbol Definition Description

[0011]

[0012]

[0013] The specific steps of step S1 include the following steps:

[0014] S1.1. Divide the data set Ω according to time sequence, where Ω 1 is the reference data window, and Ω 2 represents the data window to be tested. It can represent either a single data window to be tested or a set of multiple data windows to be tested.

[0015] S1.2. Select appropriate rule thresholds (β) and relaxation amounts (δ) according to the characteristics of the large-scale dataset. The rule threshold is used to define the credibility of the rule, and the relaxation amount is used to determine the boundary of the rule base to avoid large changes in the drift metric results caused by small changes in β. This step provides the basic parameters for the subsequent rule base construction and drift detection. Two points need to be noted regarding the determination of the β value: 1) As the β value increases, the detection speed of the algorithm is greatly improved because the number of extracted rules decreases, resulting in a decrease in the number of rule objects for interactive matching tests, thus improving the detection speed; 2) The larger (smaller) the β value, the fewer (more) the number of rule knowledge. As a result, the average change rate of the unit basic rules corresponding to the same change in the number of rules (quality) (i.e., knowledge drift) increases with the increase of β. Therefore, the knowledge drift increases with the increase of β, and the detection result may show an "oversensitive" phenomenon. Therefore, the value of β should not be too large or too small.

[0016] The specific steps of step S2 include the following steps:

[0017] S2.1. According to the sampling-with-replacement mode, respectively obtain datasets with a capacity of m from Ω 1 and Ω 2 k = 1, 2, i = 1, 2,..., n, where i represents the sequence of obtaining datasets with a capacity of m randomly from Ω k In practical applications, special attention should be paid to the following points: 1) This algorithm is constructed and analyzed in the sense of probability (i.e., m → ∞), and the sample size m cannot be too small; 2) On the other hand, as the sample size m increases and both have the characteristic of gradually approaching 0, but the detection time gradually becomes larger, but the increase amplitude is completely within the applicable range. The main reason for the increase in calculation time is that the computational complexity of rule acquisition and interactive matching increases with the increase in the number of samples. Therefore, m is preferably not less than 2000; 3) The number of sampling test times n does not have to be too large. The increase in n will lead to a linear increase in computational complexity, and generally n ≥ 10 should be satisfied. S2.2. According to the given rule threshold (β), use the ID3 decision tree method to determine the rule base of the dataset ​​The specific steps are as follows: ① For the current node M, determine whether M is a type-I leaf node. If it is a type-I leaf node, label it as a leaf node and record the decision attribute values of the examples in this leaf node. If it is not a type-I leaf node, go to ②; ② Check whether M can be further branched. If it can be further branched, select an expansion attribute for branching. If it cannot be further branched, go to ③; ③ Check whether M is a type-II leaf node. If it is a type-II leaf node, label it as a leaf node and record the decision attribute values with a similarity rate not less than the threshold β and the corresponding similarity rate in this leaf node. If it is not a type-II leaf node, label it as a discarded node (or delete this node). Repeat the above execution steps until all nodes are branch nodes, leaf nodes, or discarded nodes.

[0018] S2.3. Calculate to of the forward rule drift and the reverse rule drift and can be calculated from Equations (1) and (2):

[0019]

[0020] In the equations, and can be calculated from Equations (9) and (10)

[0021]

[0022] where represents the number of rules in the β rule base of represents the number of rules existing in the β rule base of in represents the number of rules existing in the β rule base of in

[0023] and only add a relaxation amount to β, and the calculation process remains unchanged. To highlight the dominant position of the threshold β and the equivalent effect of relaxation, the weight combination W in (1) and (2) should satisfy w 1 = w 3 ≤ w 2 , and usually w 1 = w 3 = 0.25, w 2 = 0.5.

[0024] S2.4. Calculate to of forward trust drift and reverse trust drift and can be calculated by Equation 11 and Equation 12:

[0025]

[0026]

[0027] In the formula and can be calculated by Equation 13 and Equation 14

[0028]

[0029] where C(Ω, R) represents the credibility of the rule (i.e., C(Ω, R) = s / t, s represents the number of samples in Ω that match both the antecedent B and the consequent D of R, and t represents the number of samples in Ω that match the antecedent B of R) N(Ω, R) represents the matching degree of Ω to R (i.e., N(Ω 1 , R) = s / t, s represents the number of samples in Ω that can be matched with both the antecedent B and the consequent D of R, and t represents Ω 1 the number of samples in which can be matched with the antecedent B of R).

[0030] and only add a relaxation amount to β, and the calculation process remains unchanged.

[0031] The specific steps of step S3 include the following steps:

[0032] S3.1. After completing the detection of n times of sampling with replacement, according to the detection results, combine Equation 9 - 11 to calculate the 1 from Ω 2 of forward rule drift reverse rule drift and bidirectional rule drift

[0033]

[0034] S3.2. Combine Equation 12 - 14 to calculate the 1 from Ω 2 of forward trust drift reverse trust drift and bidirectional trust drift

[0035]

[0036] S3.3. Draw Ω according to the experimental results 2 With respect to the degree of drift of Ω 1 under a certain set threshold θ 1 , θ 2 ∈ [0, ∞), determine whether drift in the sense of early warning occurs through .

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0038] 1. The knowledge drift detection method (S-RKDD) based on interaction matching proposed by the present invention can significantly improve the accuracy and reliability of knowledge drift detection by comprehensively analyzing the quantity and quality changes of IF-THEN rules. Compared with traditional methods, the present invention has the following advantages: multi-perspectivity: characterizing knowledge drift from two dimensions of rule quantity and quality simultaneously, making up for the deficiency of the existing methods in the single characterization of drift behavior; high adaptability: being able to comprehensively reflect the basic behavior of knowledge drift through the definitions of forward and reverse drift, and adapting to the requirements of complex decision-making scenarios; computational efficiency.

[0039] 2. The present invention combines the convergence characteristics of rule credibility in the context of sampling, significantly reducing the computational complexity of processing large-scale data sets and adapting to the analysis of massive data streams; structural advantage: providing a clear and interpretable drift detection structure, which is convenient for application and improvement in combination with practical problems.

[0040] 3. The present invention has achieved breakthroughs in both the theory and practicality of knowledge drift detection, providing strong technical support for data-driven decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.

[0042] Figure 1 is the flow chart of the detection method of the present invention;

[0043] Figure 2 is the schematic diagram of extracting a rule tree using the ID3 algorithm in the embodiment of the present invention;

[0044] Figure 3 is the schematic diagram of the rule drift detection result in the embodiment of the present invention; Windows 2 to 8 in the figure are all windows to be detected, that is, Ω 2 ;

[0045] Figure 4It is a schematic diagram of the trust drift detection result of the embodiment of the present invention; in the figure, Window2 to Window8 are all windows to be detected, that is, Ω 2 . Detailed implementation manners

[0046] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0047] Embodiment 1

[0048] Refer to Figures 1 to 4 As shown, this embodiment provides a knowledge drift detection method based on interaction matching, including 6 steps: ① Divide the data set; ② Select the rule threshold β ∈ [0, 1] and its relaxation amount δ ∈ [0, 1]; ③ Adopt the sampling method with replacement to obtain a data set with a capacity of m from Ω 1 and Ω 2 ; ④ Use the ID3 decision tree method to determine the rule base and ⑤ Calculate the forward rule drift from Ω to Ω 1 , 2 the reverse rule drift, and the two-way rule drift; ⑥ Calculate the forward trust drift from Ω to Ω 1 , 2 the reverse trust drift, and the two-way trust drift. The specific process is as shown. Figure 1 shown.

[0049] The specific decomposition of each step is as follows:

[0050] Step 1, divide the data set Ω according to time sequence, where Ω 1 is the reference data window, and Ω 2 represents the window to be tested data, which can represent either a single window to be tested data or a set of multiple windows to be tested data.

[0051] Step 2: According to the characteristics of the large-scale dataset, select appropriate rule thresholds (β) and relaxation amounts (δ). The rule threshold is used to define the credibility of the rule, and the relaxation amount is used to determine the boundary of the rule base to avoid large changes in the drift metric results caused by small changes in β. This step provides the basic parameters for subsequent rule base construction and drift detection. Two points need to be noted regarding the determination of the β value: 1) As the β value increases, the detection speed of the algorithm is significantly improved because the number of rules extracted decreases, resulting in a decrease in the number of rule objects for interactive matching tests, thereby improving the detection speed; 2) The larger (smaller) the β value, the fewer (more) the number of rule knowledge. As a result, the average change rate of the unit basic rules corresponding to the same change in the number of rules (quality) (i.e., knowledge drift) increases with the increase in β. Therefore, the knowledge drift increases with the increase in β, and the detection results may exhibit an "oversensitive" phenomenon. Therefore, the value of β should not be too large or too small.

[0052] Step 3: In the replacement sampling mode, obtain datasets with a capacity of m from Ω 1 and Ω 2 respectively. k = 1, 2, i = 1, 2,..., n, where i represents the sequence of obtaining datasets with a capacity of m randomly from Ω k In practical applications, the following points should be particularly noted: 1) This algorithm is constructed and analyzed in the sense of probability (i.e., m → ∞), and the sample size m cannot be too small; 2) On the other hand, as the sample size m increases and both have the characteristic of gradually approaching 0, but the detection time gradually increases, and the increase amplitude is completely within the applicable range. The main reason for the increase in the calculation time is that the computational complexity of rule acquisition and interactive matching increases with the increase in the number of samples. Therefore, m is preferably not less than 2000; 3) The number of sampling tests n does not have to be too large. The increase in n will lead to a linear increase in the computational complexity, and generally, n ≥ 10 should be satisfied.

[0053] Step 4: According to the given rule threshold (β), use the ID3 decision tree method to determine the rule base The specific steps are as follows: ① For the current node M, determine whether M is a type-I leaf node. If it is a type-I leaf node, label it as a leaf node and record the decision attribute values of the examples in this leaf node. If it is not a type-I leaf node, go to ②; ② Check whether M can be further branched. If it can be further branched, select an expansion attribute for branching. If it cannot be further branched, go to ③; ③ Check whether M is a type-II leaf node. If it is a type-II leaf node, label it as a leaf node and record the decision attribute values with a value similarity rate not less than the threshold β and the corresponding similarity rate in this leaf node. If it is not a type-II leaf node, label it as a discarded node (or delete this node). Repeat the above execution steps until all nodes are branch nodes, leaf nodes, or discarded nodes. Figure 2 Figure 1 shows a schematic diagram of a rule tree for an embodiment.

[0054] Step 5, calculate to of forward rule drift and backward rule drift and can be calculated from Equations (7) and (8):

[0055]

[0056] In the equations, and can be calculated from Equations (9) and (10)

[0057]

[0058] where represents the number of rules in the β rule base of represents the number of rules existing in the β rule base of in represents the number of rules existing in the β rule base of in

[0059] and only add a relaxation amount to β, and the calculation process remains unchanged. To highlight the dominant position of the threshold β and the equivalent effect of relaxation, the W in (7) and (8) should satisfy w 1 = w 3 ≤ w 2 , and usually w 1 = w 3 = 0.25, w 2 = 0.5.

[0060] Step 6, Calculate arrive of Positive trust drift and Reverse Trust Drift and It can be calculated by equation 11 and equation 12:

[0061]

[0062] In the formula and It can be calculated by equation 13 and equation 14

[0063]

[0064] Among them, C(Ω,R) represents the rule The credibility of (i.e., C(Ω, R) = s / t, s represents the number of samples in Ω that match both the antecedent B and the consequent D of R, and t represents the number of samples in Ω that match the antecedent B of R) N(Ω, R) represents the matching degree of Ω to R (i.e., N(Ω 1 , R) = s / t, s represents the number of samples in Ω that can match both the antecedent B and the consequent D of R, and t represents Ω 1 The number of samples that can match the antecedent B of R).

[0065] and Only the relaxation is added to β, and the calculation process does not change.

[0066] Step 7: After completing n times of sampling with replacement, calculate Ω based on the test results and formula 1-3. 1 to Ω 2 of Positive rule drift Reverse rule drift as well as Bidirectional rule drift

[0067]

[0068] Combine equations 4-6 to calculate Ω 1 to Ω 2 of Positive trust drift Reverse Trust Drift as well as Two-way trust drift

[0069]

[0070] According to the experimental results, Ω can be plotted. 2 With respect to the degree of drift of Ω 1 , within a certain set threshold θ 1 , θ 2 ∈ [0, ∞), it is determined whether drift in the sense of early warning has occurred through . The rule drift detection results and trust drift detection results of this embodiment are as Figure 3 , Figure 4 shown.

[0071] The knowledge drift detection method based on interactive matching proposed in this embodiment analyzes the changes in the quantity and quality of IF-THEN rules in the data, realizes the effective detection and characterization of knowledge drift in large-scale data sets, especially for the behavior analysis of forward and reverse rule drifts, enriches the existing data drift detection theory, and has broad application prospects.

[0072] Embodiment 2

[0073] The data description regarding the embodiment is as described in Table 2.

[0074] Table 2 is the basic information of the embodiment data set

[0075]

[0076] As Figures 1 to 4 shown, this embodiment takes the data set U 1 as an example to provide a knowledge drift detection method based on interactive matching. The data set U 1 is a shared bicycle usage data, recording the bicycle rental data for two years from 2011 to 2012. The detection method includes 3 steps: ① Divide the data set Ω into a reference data window Ω 1 and a data window to be tested Ω 2 according to time series, and determine the rule threshold β ∈ [0, 1] and its relaxation amount δ ∈ [0, 1]. ② Conduct n random samplings on the data windows Ω 1 and Ω 2 to obtain a sample data set with a capacity of m, and calculate forward rule drift and reverse rule drift as well as forward trust drift and reverse trust drift for each sampling result. ③ According to the results of n sampling detections, obtain the forward rule drift and reverse rule drift and forward trust drift and reverse trust drift and calculate two-way rule drift and two-way trust drift The specific process is as Figure 1 shown

[0077] Example U 1 The specific decomposition of each step is as follows:

[0078] S1.1. Divide the data into 8 data windows quarterly, where the data in the first quarter of 2011 is the benchmark data window Ω 1 , and the seven data windows from the second quarter of 2011 to the fourth quarter of 2012 are all data windows Ω to be tested 2 .

[0079] S1.2. According to the characteristics of the large-scale data set, select appropriate rule thresholds (β) and relaxation amounts (δ). In Example U 1 , the rule threshold β is 0.7 and δ is 0.1

[0080] S2.1. According to the sampling-with-replacement mode, repeatedly obtain data sets with a capacity of 1500 from Ω 1 and Ω 2 respectively, and the number of sampling times n = 20 Sampling times n = 20

[0081] S2.2. According to the given rule threshold β = 0.7, use the ID3 decision tree method to determine the rule base Figure 2 shows a schematic diagram of the rule tree of Example U 1 . S2.3. Calculate to of forward rule drift and reverse rule drift and can be calculated by Equations 1 and 2:

[0082]

[0083] In the formula and can be calculated by Equations 9 and 10

[0084]

[0085] and Only a relaxation amount is added to β, and the calculation process remains unchanged. To highlight the dominant position of the threshold β and the equivalent effect of relaxation, w in (1) and (2) is taken as w 1 = w 3 = 0.25, w 2 = 0.5.

[0086] S2.4. Calculate to of forward trust drift and reverse trust drift and can be calculated from Equations 11 and 12:

[0087]

[0088]

[0089] In the equation, and can be calculated from Equations 13 and 14

[0090]

[0091] and Only a relaxation amount is added to β, and the calculation process remains unchanged.

[0092] S3.1. According to the detection results, calculate Ω by combining Equations 9 - 11 1 to Ω 2 of forward rule drift reverse rule drift and bidirectional rule drift

[0093]

[0094] S3.2. Combine Equations 12 - 14 to calculate Ω 1 to Ω 2 of forward trust drift reverse trust drift and bidirectional trust drift

[0095]

[0096] S3.3. According to the experimental results, Ω can be plotted 2 relative to Ω 1The degree of drift, Figure 3 and Figure 4 respectively show Ω 2 relative to Ω 1 of two-way rule drift degree and two-way trust drift degree

[0097] Then, in combination with four typical drift detection methods, namely KS test, ADWIN, SPRT, and EDMM, the performance comparison with the RKDD method will be carried out for four performance indicators: accuracy (Acc), recall (Rec), precision (Pre), and F1-Score. The test results are shown in Table 3-5.

[0098] Table 3 Performance comparison results for U 1

[0099]

[0100] Table 4 Performance comparison results for U 2

[0101]

[0102] Table 5 Performance comparison results for U 3

[0103]

[0104]

[0105] As can be seen from Tables 3-5: 1) On different embodiment datasets, the prediction performance of the RKDD algorithm is generally better than that of other comparison algorithms. Among them, the low prediction performance of all algorithms in Embodiment U 1 is due to the unbalanced data distribution. The RKDD algorithm shows good performance on Embodiments U 2 and U 3 , indicating that the RKDD algorithm has excellent detection performance in processing balanced distribution data; 2) In terms of running time, the running time of the RKDD algorithm is generally higher than that of ADWIN, SPRT, and EDDM. This is because the time cost of extracting the rule tree is much greater than the time cost of directly testing from the perspective of data distribution. However, the RKDD algorithm can measure the degree of drift, which cannot be achieved by other algorithms. Therefore, we believe that these time costs are worthwhile.

[0106] ​​​The knowledge drift detection method based on interactive matching proposed in this embodiment realizes the effective detection and characterization of knowledge drift in large-scale data sets by analyzing the quantity and quality changes of IF-THEN rules in the data. Especially for the behavior analysis of forward and reverse rule drifts, it enriches the existing data drift detection theory and has broad application prospects.

[0107] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A knowledge drift detection method based on interactive matching, characterized in that: in, Ω1 and Ω2 are two datasets with the same formal structure, including the following steps: S1, divide the data set Ω into the benchmark data window Ω1 and the test data window Ω2 according to the time series, and determine the rule threshold β∈[0,1] and its relaxation δ∈[0,1]; S2, randomly sample the data windows Ω1 and Ω2 n times to obtain a sample data set with a capacity of m, and calculate for each sampling result Positive rule drift and Reverse rule drift as well as Positive trust drift and Reverse Trust Drift S3. According to the results of n sampling tests, the average Positive rule drift and Reverse rule drift as well as Positive trust drift and Reverse Trust Drift And calculate Bidirectional rule drift and Two-way trust drift 2. The knowledge drift detection method based on interactive matching according to claim 1 is characterized in that: The step S1 comprises the following steps: S1.1, divide the data set Ω according to the time series, where Ω1 is the reference data window, and Ω2 represents the test data window, which represents both a single test data window and a collection of multiple test data windows; S1.

2. According to the characteristics of large-scale data sets, select appropriate rule threshold β and relaxation amount δ. The rule threshold is used to define the credibility of the rule, and the relaxation amount is used to determine the boundary of the rule base.

3. The knowledge drift detection method based on interactive matching according to claim 1 is characterized in that: The step S2 comprises the following steps: S2.

1. According to the sampling mode with replacement, obtain data sets with capacity m from Ω1 and Ω2 respectively. k=1,2,i=1,2,...,n,i represents the k The sequence of randomly obtaining a data set with a capacity of m; S2.3 Calculation arrive of Positive rule drift and Reverse rule drift and Calculated by equation 1 and equation 2: In the formula and Calculated from equations 9 and 10 in, express The number of rules in the β rule base, express The beta rule base is The number of rules in express The beta rule base is The number of rules present in ; and Only the relaxation is added to β, and the calculation process does not change. The weight combination W in (1) and (2) should satisfy w1 = w3 ≤ w2, and w1 = w3 = 0.25, w2 = 0.5; S2.

4. Calculation arrive of Positive trust drift and Reverse Trust Drift and Calculated by formula (11) and formula (12): In the formula and Calculated by equation 13 and equation 14 Among them, C(Ω,R) represents the rule The credibility of R is C(Ω,R)=s / t, where s represents the number of samples in Ω that match both the antecedent B and the consequent D of R, and t represents the number of samples in Ω that match the antecedent B of R. N(Ω,R) represents the matching degree of Ω to R, that is, N(Ω1,R)=s / t, where s represents the number of samples in Ω that match both the antecedent B and the consequent D of R, and t represents the number of samples in Ω1 that can match the antecedent B of R. and Only the relaxation is added to β, and the calculation process does not change.

4. The knowledge drift detection method based on interactive matching according to claim 1 is characterized in that: The step S3 comprises the following steps: S3.

1. After completing n times of sampling with replacement, calculate the value of Ω1 to Ω2 based on the test results by combining equations (9) to (11). Positive rule drift Reverse rule drift as well as Bidirectional rule drift S3.2, Combine equations (12) to (14) to calculate Ω1 to Ω2 Positive trust drift Reverse Trust Drift as well as Two-way trust drift S3.3, according to the experimental results, plot the drift degree of Ω2 relative to Ω1, under a certain set threshold θ1,θ2∈[0,∞), through To determine whether drift with warning significance has occurred.

5. The knowledge drift detection method based on interactive matching according to claim 1 is characterized in that: Step S2.2, according to the given rule threshold β, the data set is determined using the ID3 decision tree method Rule base The specific steps are: ① The current node M determines whether M is a type I leaf node: if it is a type I leaf node, it is marked as a leaf node. And record the decision attribute value of the example in the leaf node; if it is not a type I leaf node, go to ②; ② Check whether M can continue to branch: If it can continue to branch, select the extended attribute to branch; If you cannot continue branching, go to step ③; ③ Check whether M is a type II leaf node: if it is a type II leaf node, mark it as a leaf node, and record the decision attribute value and corresponding same rate of the leaf node whose value identical rate is not less than the threshold β; if it is not a type II leaf node, mark it as a discarded node or delete the node; ④ Repeat the above steps ① to ③ until all nodes are branch nodes, leaf nodes or abandoned nodes.