A Method for Discovering Differential Dependencies in Relational Databases for Incremental Scenarios

By constructing a differential dependency prefix tree and using incremental data to generate distance vectors, differential dependency is dynamically verified, and the problem of inefficiency in incremental scenarios in the prior art is solved, and efficient differential dependency discovery is achieved.

CN116303816BActive Publication Date: 2025-07-08XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310085395.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-02
Publication Date
2025-07-08
Estimated Expiration
2043-02-02

AI Technical Summary

Technical Problem

The existing differential dependency discovery method is inefficient in incremental scenarios and cannot effectively utilize the original differential dependency set information, resulting in computational complexity and waste of resources.

Method used

Construct the differential dependency prefix tree, use the original differential dependency set information, generate distance vectors through incremental data, dynamically verify the differential dependency, reduce the unnecessary verification process, and improve the discovery speed.

Benefits of technology

It realizes dynamic discovery of differential dependence in incremental scenarios, reduces calculation overhead, and improves discovery speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116303816B_ABST
    Figure CN116303816B_ABST
Patent Text Reader

Abstract

A method for discovering differential dependencies in a relational database for incremental scenarios includes the following steps: constructing a differential dependency prefix tree according to the original differential dependency set ∑ of the database; calculating newly formed distance vectors based on the incremental data Δr to obtain a new distance vector set V; constructing a position list index for each attribute according to the new distance vector set V; verifying the original differential dependencies and the generated new differential dependencies; traversing the differential dependency prefix tree to obtain a new differential dependency set ∑'. The present invention can achieve dynamic discovery of differential dependencies in a data mining or data analysis architecture, and ensure accuracy without a large number of verification processes, thereby reducing the memory overhead required for calculation, improving the calculation efficiency, and providing a basis for discovering new rules for the database after increment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data mining and data analysis, and particularly relates to a method for discovering differential dependencies in a relational database for incremental scenarios. Background Art

[0002] Data mining is the process of discovering unknown and potentially valuable information hidden in a large amount of data from a database, and effectively reasoning and exploring the laws of the physical world through this information. Among them, data dependency discovery is an important research direction in the field of data mining. Data dependency is a constraint relationship between attributes within a relation, belonging to the inherent nature of data. For example, the most commonly used data dependency - functional dependency describes the constraint relationship that records with equal values on certain attributes in a database must also have equal values on other attributes. With the advent of the big data era, the diversity and complexity of data have led to the inability of constraints based on equality relationships to fully describe the relationships between attributes. Therefore, more and more types of data dependencies have been proposed and studied. Among them, differential dependency is a data dependency relationship defined based on distance semantics and is an extension of the traditional functional dependency relationship. In a relation instance r, the establishment of a differential dependency means that for any two tuples in r, if the distance between them on the left-hand side attributes satisfies the differential constraint, then the distance between them on the right-hand side attributes must also satisfy the differential constraint. Compared with functional dependencies, the definition of differential dependencies covers more general dependency relationships between data and has a more powerful expressive ability. Differential dependencies play a crucial role in many fields such as violation detection, data partitioning, query optimization, and record linkage. Therefore, discovering differential dependency relationships in a database has always been a hot topic of concern for researchers.

[0003] However, many modern applications need to continuously integrate data from distributed, heterogeneous, and autonomous data sources, which will lead to an increasing amount of data in the database, causing the set of established differential dependencies to change accordingly. The increased data can lead to two possible changes in the set of differential dependencies: the invalidation of the original differential dependencies and the generation of new differential dependencies. Traditional non-incremental differential dependency discovery methods, due to their own computational complexity and the lack of utilization of the information in the original differential dependency set, have problems such as low efficiency and redundant processes when solving differential dependency discovery in incremental scenarios. Efficiently maintaining and discovering differential dependencies in incremental scenarios is a new problem in the field of dependency relationship discovery and has broad application prospects.

[0004] Existing differential dependency discovery methods mainly solve the problem of automatically discovering all or part of the differential dependencies in static data instances, including reduction-based methods, lattice search-based methods, clustering-based methods, association rule-based methods, etc. Among them, the reduction-based method (e.g., S. Song and L. Chen, “Differential dependencies: Reasoning and discovery,” ACM Trans. Database Syst., vol. 36, pp. 16:1–16:41, 2011) verifies the formed differential dependencies by fixing each right-hand side and traversing all candidate left-hand side sets in turn, and proposes pruning based on the properties of differential dependencies to improve efficiency. The lattice search-based method (e.g., Jixue Liu, Selasi Kwashie, and Jiuyong Li, “Discovery of Approximate Differential Dependencies,” CoRR abs / 1309.3733, 2013) verifies candidate differential dependencies by establishing groups for tuple pairs and traversing the lattice formed by left-hand side sets layer by layer from top to bottom. The clustering-based method (e.g., S. Kwashie, J. Liu, J. Li, and F. Ye, “Mining differential dependencies: A subspace clustering approach,” in Proc. Australas. Database Conf., 2014, pp. 50–61) effectively discovers high-threshold differential dependencies using a distance-based subspace clustering model. The association rule-based method (e.g., S. Kwashie, J. Liu, J. Li, and F. Ye, “Efficient discovery of differential dependencies through association rules mining,” in Proc. Australas. Database Conf., 2015, pp. 3–15) discovers the relationship between association rules and differential dependencies and solves the differential dependency discovery problem by mining a class of non-redundant association rules.

[0005] In the problem of differential dependency discovery in an incremental scenario, existing differential dependency discovery methods can only maintain the differential dependency set by re-executing the algorithm every time the dataset changes. This approach does not utilize the established situation of differential dependencies before the data change, resulting in the repeated execution of some verification processes and wasting time and space overhead. In scenarios where data is updated frequently, due to the complexity of the calculation process, existing non-incremental discovery methods cannot update the differential dependency set within a reasonable time. Summary of the Invention

[0006] To overcome the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a method for discovering differential dependencies in a relational database for an incremental scenario, aiming to achieve dynamic discovery of differential dependencies in a data mining or data analysis architecture, and ensure accuracy without the need for a large number of verification processes, thereby reducing the memory overhead required for calculation, improving calculation efficiency, and providing a basis for the decision-making target of the data after increment.

[0007] To achieve the above purpose, the technical solution adopted by the present invention is:

[0008] A method for discovering differential dependencies in a relational database for an incremental scenario, comprising the following steps:

[0009] Step 1, construct a differential dependency prefix tree according to the original differential dependency set ∑ of the database;

[0010] Step 2, calculate the newly formed distance vectors according to the incremental data Δr to obtain a new set of distance vectors V;

[0011] Step 3, construct a position list index for each attribute according to the new set of distance vectors V;

[0012] Step 4, verify the original differential dependencies and the newly generated differential dependencies;

[0013] Step 5, traverse the differential dependency prefix tree to obtain a new differential dependency set ∑′.

[0014] Compared with the prior art, the beneficial effects of the present invention are:

[0015] Utilize the information of the original differential dependency set, reduce unnecessary verification processes, and achieve dynamic discovery of differential dependencies.

[0016] A method for discovering differential dependencies in an incremental scenario provided by the present invention utilizes the existing differential dependency set and dynamically verifies the differential dependencies through the distance vectors generated by the incremental data, thereby improving the discovery speed of differential dependencies in the incremental scenario. Brief Description of the Drawings

[0017] Figure 1 It is a schematic flowchart of the present invention.

[0018] Figure 2 It is a schematic diagram of relationship instances and incremental data in an embodiment of the present invention.

[0019] Figure 3 It is a schematic diagram of the constructed differential dependency prefix tree in an embodiment of the present invention. Detailed implementation manners

[0020] The implementation manners of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0021] As described above, in the prior art, in an incremental scenario, when discovering differential dependency relationships, the existing differential dependency relationships cannot be utilized, and only the entire algorithm can be repeatedly executed. Obviously, this method seriously occupies time and space resources and has low efficiency.

[0022] To solve this problem, the present invention considers the establishment of differential dependencies before data change, thereby omitting the execution of some unnecessary verification processes, and thus greatly improving the discovery speed of differential dependency relationships in an incremental scenario.

[0023] The knowledge basis on which the present invention is based is described as follows:

[0024] In a relationship instance r, the relationship is expressed as R = (A1,…,A i ,…,A m ), A i represents the i-th attribute in the relationship R, A i ∈R, and m represents the number of attributes in R. A distance function is defined on each attribute to measure the similarity of the values of different tuples on this attribute. For example, the distance function on A i is denoted as represents the distance between two tuples t1 and t2 in the relationship instance r. The choice of the distance function can be determined by the user according to actual needs. For example, the Euclidean distance can be selected for numerical attributes, and the edit distance can be selected for string attributes, etc. The differential interval is a number of intervals of interest divided by the user on each attribute. For example, a number of intervals of interest divided by the user on the attribute A i can be expressed as: These intervals are the differential intervals, where it is necessary to satisfy n i represents the number of intervals of interest of the user on the attribute A i . In ascending order, these differential intervals are numbered starting from 1, that is represents the first differential interval on the attribute A i , represents the j-th i on the attribute A iA differential interval, and so on.

[0025] The differential dependency is represented in the form of . Among them, X and Y are attribute sets in the relationship R, and X and Y have no intersection. represents the left - hand side set LHS, represents the right - hand side set RHS; A i <j i > and A k <j k > represent distance constraints. Two tuples t1 and t2 in the relationship instance r satisfy the distance constraint A i <j i > indicating that the distance between t1 and t2 on A i is in the j i th differential interval, that is Similarly, satisfying the distance constraint A k <j k > means that the distance between t1 and t2 on A k is within the j k th differential interval. The differential dependency holds, indicating that for any two tuples in the relationship instance r, if for any A i ∈X, all satisfy A i <j i >, then for each A k ∈Y, it must also satisfy A k <j k . The set of differential dependencies maintained by the present invention is a minimum cover set. In the minimum cover set, it is necessary to satisfy:

[0026] (1) The right - hand side set RHS is a single distance constraint, that is, it satisfies the form of .

[0027] (2) There does not exist any such that holds, where is called a generalization of .

[0028] Refer to Figure 1 , the method for discovering differential dependencies in a relational database for an incremental scenario of the present invention includes the following steps:

[0029] Step 1, construct a differential - dependency prefix tree according to the original set of differential dependencies ∑ of the database.

[0030] The structure of the differential - dependency prefix tree described in the present invention is as follows:

[0031] The nodes of the differential dependency prefix tree include a root node, an attribute node, and an interval node. The top layer is the only root node. The attribute nodes are the children of the root node, and the interval nodes are the children of the attribute nodes. There is only one root node, and both the attribute nodes and the interval nodes have multiple layers. Among them, the first-layer attribute nodes are the children of the root node, the first-layer interval nodes are the children of the first-layer attribute nodes, the second-layer attribute nodes are the children of the first-layer interval nodes, and the second-layer interval nodes are the children of the second-layer attribute nodes, and so on. However, the leaf nodes must be interval nodes. The parent and child nodes are connected by bidirectional pointers. An attribute node and its child interval nodes represent a distance constraint, and the path from the root node to an interval node represents an LHS.

[0032] The attribute node stores the unique identifier of the attribute. The set of attributes in all distance constraints of an LHS is called the attribute set corresponding to this LHS. The LHS attribute index maintains a node list for each attribute set corresponding to an LHS. This node list records the last attribute node of all LHSs in the differential dependency prefix tree corresponding to this attribute set.

[0033] The interval node records the interval number where its parent attribute node is located in the distance constraint, and stores a bit set and an array. The bit set and the array are used to represent the RHS situation corresponding to the differential dependency with the current path as the LHS. Each bit in the bit set and each position in the array correspond to an attribute. When case (1): the LHS represented by the current path can make this attribute satisfy a certain distance constraint as the RHS; or case (2): the LHS with the current path as the prefix can make this attribute satisfy a certain distance constraint as the RHS, the bit set sets the bit corresponding to this attribute to 1, otherwise it sets it to 0. Among them, when case (1) is satisfied, the interval number of the distance constraint is stored in the corresponding position in the array corresponding to this attribute. In other cases, the value in the array defaults to 0. When there is a non-zero value in the array, this interval node represents a differential constraint, its LHS is the path from the root node to this interval node, and its RHS is the distance constraint composed of the attribute with a non-zero value in the array and the corresponding interval.

[0034] The construction process of the differential dependency prefix tree is as follows:

[0035] (1) Initialize the prefix tree: Create an empty node as the root node.

[0036] (2) Traverse all differential dependencies in the original differential dependency set ∑, calculate the number of times each attribute appears in the LHS, and sort the attributes in descending order according to the number of times.

[0037] (3) For each differential dependency dd, sort the distance constraints in its corresponding left-hand side set dd.LHS according to the order of attributes. The steps to add the differential dependency dd to the prefix tree are as follows:

[0038] (3.1) Find the longest prefix path of dd.LHS in the current prefix tree, that is, find a path starting from the root node, whose represented LHS is the first n distance constraints of dd.LHS, such that n is the largest. Return the last interval node in the longest prefix path. If there is no prefix, return the root node. Set the returned node as node, and set i as the number of distance constraints included in the longest prefix path + 1.

[0039] (3.2) If i is greater than the number of distance constraints in dd.LHS, directly go to the next step. Otherwise, determine whether the child nodes of node include the attribute node of the i-th distance constraint in dd.LHS. If not, create a new attribute node for the i-th distance constraint in dd.LHS as a child node of node, add a pointer to this attribute node in the node list corresponding to the LHS attribute in the LHS attribute index, create a new interval node as a child node of the attribute node. If it exists, only create a new interval node and use it as a child node of node. Set the new interval node as node, increment the i value by 1, and repeat step (3.2).

[0040] (3.3) Set the bits in the bit sets of each interval node on the dd.LHS path to 1 on the right-hand side set dd.RHS attribute corresponding to the differential dependency dd, and set the value on the dd.RHS attribute in the array of the last interval node in dd.LHS to the corresponding interval number.

[0041] Step 2, calculate the newly formed distance vectors according to the incremental data Δr to obtain a new set of distance vectors V, including the following steps:

[0042] (1) Calculate the set of distance vectors V1 between the incremental data and the original data set:

[0043]

[0044] Denote the distance between tuple t1 and tuple t2 on the m-th attribute A m ; Δr refers to the relational instance of the incremental data, and r refers to the relational instance of the original data set.

[0045] (2) Calculate the set of distance vectors V2 between the incremental data:

[0046]

[0047] (3) Replace the specific distance values with the differential interval numbers where the elements in the distance vectors are located. If not in all intervals, replace with 0 to obtain the distance vector sets V1′, V2′.

[0048] (4) A new set of distance vectors V = V1' ∪ V2', and a unique number is defined for each distance vector in V.

[0049] Step 3: According to the new set of distance vectors V, construct a positional list index for each attribute, and the process is as follows:

[0050] (1) Initialize the index: Create an empty set for each attribute in all difference intervals, and each empty set serves as a group.

[0051] (2) Traverse the new set of distance vectors V, and according to the interval value of each distance vector on each attribute, add the unique number of this distance vector to the group corresponding to each attribute.

[0052] Step 4: Verify the original differential dependencies and the generated new differential dependencies. Verify the differential dependencies in the original differential dependency set ∑ in the order of the number of attributes in the LHS, and the verification process is as follows:

[0053] Step (1): Let i = 1;

[0054] Step (2): Traverse the attribute nodes in the node list corresponding to the attribute set of size i in the LHS attribute index in turn, and verify the differential dependencies represented by the child nodes of all attribute nodes in the node list at the same time.

[0055] In the present invention, for the attribute set A of size i l1 …A li , the verification steps are as follows:

[0056] Step (2.1) Traverse all A l1 …A li The child nodes of the attribute nodes in the j l1 group in the LHS attribute index. Suppose a differential dependency formed by an interval node is A l1 <j l1 >…A li <j li >→A r1 <j r1 >…A rk <j rk >, create a left-hand side set array to store j l1 …j li , and create a right-hand side set array to store j r1 …j rk , create a hash table, and add the left-hand side set array as the key and the right-hand side set array as the value to the hash table;

[0057] Step (2.2) In the positional list index of A l1 , find j l1The corresponding group is traversed through all distance vectors in the group, and the distance vectors are in A l1 …A li The value of...A is mapped in the hash table in step (2.1); if the corresponding left-hand side set array and the distance vector can be found in the hash table and the value of...A l1 …A li match, then the values in the corresponding right-hand side set array are compared with the values of the corresponding attributes in the distance vector, and the values in all non-matching right-hand side set arrays are set to 0. If all the values corresponding to the attributes in the right-hand side set array are 0, then the corresponding key-value pair is deleted from the hash table; when the hash table is empty or all the distance vectors in the group have been traversed, a verification process ends;

[0058] Step (2.3), generate the nodes to be verified in the next layer. For each attribute whose right-hand side set array is set to 0 in step (2.2), add a distance constraint of a new attribute that does not belong to the LHS and RHS to its corresponding LHS to form the differential dependency to be verified in the next layer, that is, the new differential dependency. If the new differential dependency is not included in the differential dependency prefix tree, or the differential dependency prefix tree does not include a generalization of the new differential dependency, then the new differential dependency is added to the differential dependency prefix tree.

[0059] Step (3), increment the value of i. If there are still unverified nodes in the differential dependency prefix tree, jump to step (2), otherwise go to the next step.

[0060] Step 5, traverse the differential dependency prefix tree to obtain the new differential dependency set ∑'.

[0061] In the present invention, for the incremental data Δr, the differential dependencies that hold in the database after the increment are obtained. Taking the right-hand side set attributes of the differential dependencies as the decision targets and the left-hand side set attributes as the decision variables, distance constraints on the decision targets are formulated for the incremental data according to the distance constraints on the decision variables, and these are used as the new rules that hold in the database. Based on this new rule, the decision targets can be adjusted accordingly according to the decision variables.

[0062] To implement the discovery method of the present invention, the present invention also provides a relational database differential dependency discovery system for an incremental scenario, including a storage unit, a differential dependency prefix tree construction unit, a distance vector calculation unit, a position list index construction unit, a verification unit, and an output unit.

[0063] A storage unit stores the original differential dependency set ∑ of the database, and the original differential dependency set ∑ can be generated according to the data content of the database. The differential dependency prefix tree construction unit generates a differential dependency prefix tree by calling the original differential dependency set ∑ using the aforementioned method. The distance vector calculation unit uses the aforementioned method to calculate the newly formed distance vector when new incremental data Δr enters the database; the position list index construction unit constructs a position list index for each attribute according to the newly formed distance vector; the verification unit verifies the original differential dependencies and the newly generated differential dependencies, and traverses the differential dependency prefix tree to obtain the new differential dependency set ∑'. In the output unit, according to the new differential dependencies, one attribute can be used as the decision target, and other attributes can be used as decision variables, and a decision target can be formulated for the data after the increment according to the decision variables.

[0064] In one embodiment of the present invention, it is necessary to extract the main factors affecting the salary differences among different employees from the employee information table. Figure 2 Figure (a) shows an example of an employee information table, which contains four attributes: age (A), education level (B), work experience (C), and salary (D), and a total of 5 groups of data, that is, 5 tuples. In this embodiment, the salary attribute is used as the decision target, and other attributes are used as decision variables, aiming to objectively give a salary range guidance for the new data according to the decision variables. The data types of the data in each attribute of the employee information table can be regarded as numerical values, and the Euclidean distance is used as the distance function. For example, the data in the age attribute can directly take the age value, the data in the work experience attribute can directly take the working years value, the data in the salary attribute can directly take the salary value, and in the education level attribute, junior college can be recorded as 1, undergraduate can be recorded as 2, master can be recorded as 3, and doctor can be recorded as 4. Assume that the differential intervals of interest on attributes A, B, and C are all [0,0] (equal) and [1,2] (similar), and the differential intervals of interest on attribute D are [0,1000) and [1000,2000], and the two intervals of each attribute are numbered 1, 2.

[0065] The differential dependencies that hold on the original data are default known and can be discovered through existing non-incremental differential dependency discovery methods or manually, etc. Assume that the differential dependency set that holds on the original relation instance is

[0066] ∑ = {A<2> → C<2>D<2>, B<1> → A<1>C<2>, C<1> → B<2>, D<1> → B<2>C<1>, A<1>B<2> → C<1>D<2>, A<1>C<2> → B<1>, A<1>D<2> → B<2>C<1>, B<2>D<1> → A<1>, C<2>D<2> → A<2}}.

[0067] Among them, A<2>→C<2>D<2> represents two differential dependencies, A<2>→C<2> and A<2>→D<2>, and the other expressions are similar. From this set of differential dependencies, users can obtain some valid information. For the problem of the main factors affecting the salary differences among different employees, it is necessary to extract the differential dependencies related to attribute D (i.e., the right-hand side set is D). For example, A<2>→D<2> reflects that the salary difference between employees with an age difference of 1 - 2 years is 1000 - 2000 yuan. A<1>B<2>→D<2> reflects that the salary difference between employees of the same age with an educational level difference of 1 - 2 levels is 1000 - 2000 yuan. Through this information, the company can judge the rationality of salary setting and make targeted decisions. Suppose new employees (employee numbers 6 and 7, i.e., tuple 6 and tuple 7) join at this time, and the incremental data is as shown in Figure 2 in (b) below, and the updated differential dependencies can be obtained through the method of the present invention. Refer to Figure 3 , and the complete process can be described as follows:

[0068] 1. Construct a differential dependency prefix tree

[0069] (1) Create an empty node as the root node root.

[0070] (2) Calculate the number of times each attribute appears in the LHS of all differential dependencies in ∑. A, B, C, and D appear 4 times, 3 times, 3 times, and 4 times respectively in the left-hand side set of ∑, that is: A: 4, B: 3, C: 3, D: 4. The sorting of attributes determined according to the number of times is ADBC.

[0071] (3) Add all differential dependencies in ∑ to the differential dependency prefix tree in turn. The following takes A<1>B<2>→C<1>D<2> as an example to illustrate the process of adding a differential dependency to the prefix tree:

[0072] (3.1) There is no prefix of A<1>B<2> in the current tree, return to the root node.

[0073] (3.2) Create an attribute node A as the child node of the root node, create an interval node 1 as the child node of the attribute node A, and represent it with the letter M in Figure 3 . Create an attribute node B as the child node of the interval node 1, create an interval node 2 as the child node of the attribute node B, and represent it with the letter N in Figure 3 .

[0074] (3.3) The initial bit set is defaulted to all 0s. The right part set is C<1>D<2> corresponding to the attributes CD. Therefore, the values of the CD bits are set to 1. That is, the bit sets of interval node 1 and interval node 2 are changed to 0011. The value of the array of interval node 2 for attribute C is set to 1, indicating that attribute C is within the differential interval 1. The value of the array for attribute D is set to 2, indicating that attribute D is within the differential interval 2.

[0075] 2. Calculate the incremental distance vector set.

[0076] (1) Calculate the incremental distance vectors pairwise between the incremental data and the original data set. For example Figure 2 tuple 1 shown in (a) and Figure 2 tuple 6 shown in (b) calculate the Euclidean distance (i.e., the absolute value of the difference) on each attribute, obtaining the distance vector (1, 1, 2, 2000). Tuple 2 and tuple 6 obtain the distance vector (1, 1, 1, 1000), and so on. Finally, the incremental distance vector set V1 is obtained:

[0077] {(1, 1, 2, 2000), (1, 1, 1, 1000), (2, 0, 2, 2000), (0, 2, 3, 1000), (0, 3, 3, 0), (2, 0, 1, 1000), (1, 1, 0, 0), (1, 2, 0, 1000)}

[0078] It should be noted that duplicate removal is performed here. For example, (2, 0, 2, 2000) appears twice, and only one needs to be retained.

[0079] (2) Calculate the distance vectors pairwise between the incremental data. Tuple 6 and tuple 7 calculate the Euclidean distance on each attribute to obtain the distance vector (1, 1, 3, 1000), obtaining the incremental distance vector set

[0080] V2 = {(1, 1, 3, 1000)}

[0081] (3) Replace the specific distance values with interval numbers and perform duplicate removal. For example, the distance value of attribute A in (1, 2, 0, 1000) is in the interval [1, 2] numbered 2, and is replaced with 2. The distance value of attribute B is in the interval [1, 2] numbered 2, and is replaced with 2. The distance value of attribute C is in the interval [0, 0] numbered 1, and is replaced with 1. The distance value of attribute D is in the interval [1000, 2000] numbered 2, and is replaced with 2. Therefore, (1, 2, 0, 1000) is replaced with (2, 2, 1, 2). Replace all distance vectors to obtain

[0082] V1' = {(2, 2, 2, 2), (2, 1, 2, 2), (1, 2, 0, 2), (1, 0, 0, 1), (2, 2, 1, 1), (2, 2, 1, 2)}

[0083] V2′ = {(2, 2, 0, 2)}

[0084] Similarly, after replacing V1 with V1′, some duplicate distance vectors will appear, and duplicate removal is also performed here.

[0085] (4) The final set of incremental distance vectors is V = V′1 ∪ V′2, that is

[0086] {(2, 2, 2, 2), (2, 1, 2, 2), (1, 2, 0, 2), (1, 0, 0, 1), (2, 2, 1, 1), (2, 2, 1, 2), (2, 2, 0, 2)}

[0087] These vectors are numbered 1 - 7 in sequence.

[0088] 3. Construct a position list index PLI for each attribute

[0089] (1) Initialize the index:

[0090] (2) Traverse the distance vectors to construct the position list index:

[0091] The first distance vector is (2, 2, 2, 2), and its value on attribute A is 2. Create a group for interval 2 in PLI(A) and add the number of the distance vector to the group, obtaining PLI(A) = {2: {1}}, where 2 represents the interval number and 1 represents the number of the distance vector. Similarly, according to the values of distance vector 1 on B, C, and D, PLI(B) = {2: {1}}, PLI(C) = {2: {1}}, and PLI(D) = {2: {1}} are obtained.

[0092] The second distance vector is (2, 1, 2, 2), and its value on attribute A is 2. Add it to the group for interval 2 in PLI(A), obtaining PLI(A) = {2: {1, 2}}. Its value on attribute B is 1. Create a group for interval 1 in PLI(B) and add the number of the distance vector to the group, obtaining PLI(B) = {1: {2}, 2: {1}}. Similarly, according to the values of distance vector 2 on C and D, PLI(C) = {2: {1, 2}} and PLI(D) = {2: {1, 2}} are obtained.

[0093] And so on, finally obtaining

[0094] PLI(A) = {1: {3, 4}, 2: {1, 2, 5, 6, 7}}

[0095] PLI(B) = {1: {2}, 2: {1, 3, 5, 6, 7}}

[0096] PLI(C) = {1: {5, 6}, 2: {1, 2}}

[0097] PLI(D) = {1: {4, 5}, 2: {1, 2, 3, 6, 7}}

[0098] 4. Verify the original differential dependency

[0099] Verify the differential dependency with the number of LHS attributes being 1. Taking A as an example:

[0100] (1) According to the LHS attribute index, find the attribute node corresponding to A in the prefix tree. The differential dependency represented by the interval node is A<2> → C<2>D<2>. Establish a hash table and add A = 2 as the key and C = 2, D = 2 as the values to the hash table.

[0101] (2) Find the group where interval 2 is located in PLI(A): {1, 2, 5, 6, 7}. Traverse all distance vectors in the partition.

[0102] 1: (2, 2, 2, 2) matches A = 2, and the values on CD are 2, 2, which are equal to C = 2, D = 2.

[0103] 2: (2, 1, 2, 2) matches A = 2, and the values on CD are 2, 2, which are equal to C = 2, D = 2.

[0104] 5: (2, 2, 1, 1) matches A = 2, and the values on CD are 1, 1, which are not equal to C = 2, D = 2. Set both C and D to 0. Delete the item corresponding to the key A = 2 from the hash table. At this time, the hash table is empty, and the verification process ends.

[0105] (3) From Generate new candidate differential dependencies

[0106] A<2>B<1> → C<2> (from B<1> → C<2>, not minimal)

[0107] A<2>B<2> → C<2> (minimal, add to the prefix tree)

[0108] A<2>D<1> → C<2> (from D<1> → C<1>, not minimal)

[0109] A<2>D<2> → C<2< (minimal, add to the prefix tree)

[0110] A<2>B<1> → D<2> (minimal, add to the prefix tree)

[0111] A<2>B<2> → D<2> (minimal, add to the prefix tree)

[0112] A<2>C<1> → D<2> (minimal, add to the prefix tree)

[0113] A<2>C<2> → D<2> (minimum, added to the prefix tree)

[0114] For the differential dependencies B<1> → A<1>C<2> and C<1> → B<2> with the number of LHS attributes being 1, perform the same verification as above. After the first layer of verification, the set of differential dependencies in the prefix tree is

[0115] B<1> → C<2>, C<1> → B<2>, A<1>B<2> → C<1>D<2>, A<1>C<2> → B<1>, A<1>D<2> → B<2>C<1>, B<2>D<1> → A<1>C<1>, C<2>D<2> → A<2>, A<2>B<2> → C<2>D<2>, A<2>D<2> → C<2>, A<2>B<1> → D<2>,

[0116] A<2>C<1> → D<2>, A<2>C<2> → D<2>, B<1>C<1> → A<2>, B<1>C<2> → A<2>, B<1>D<1> → A<2>, B<1>D<2> → A<2>, A<1>D<1> → B<2>, A<2>D<1> → B<2>, C<2>D<1> → B<2>, A<1>D<1> → C<1>, A<2>D<1> → C<1>

[0117] Continue with the differential dependencies having the number of LHS attributes equal to 2. Taking AB as an example:

[0118] (1) Find the attribute nodes corresponding to AB in the prefix tree according to the LHS attribute index. The differential dependencies represented by the interval nodes are A<1>B<2> → C<1>D<2>, A<2>B<2> → C<2>D<2>, A<2>B<1> → D<2>. Create a hash table and add the key (A = 1, B = 2) value (C = 1, D = 2); key (A = 2, B = 2) value (C = 2, D = 2); key (A = 2, B = 1) value (D = 2) to the hash table.

[0119] (2) Find the group where interval 1 is located in PLI(A): {3, 4} and the group where interval 2 is located: {1, 2, 5, 6, 7}, and traverse all the distance vectors in the two partitions.

[0120] 3: (1, 2, 0, 2) matches A = 1, B = 2. The values on CD are 0, 2, which are not equal to C = 1 and equal to D = 2. Set the value of C to 0.

[0121] 4: (1, 0, 0, 1) does not match any key.

[0122] 1: (2, 2, 2, 2) matches A = 2, B = 2, the values on CD are 2, 2, which are equal to C = 2, D = 2.

[0123] 2: (2, 1, 2, 2) matches A = 2, B = 1, the value on D is 2, which is equal to D = 2.

[0124] 5: (2, 2, 1, 1) matches A = 2, B = 2, the values on CD are 1, 1, which are not equal to C = 2, D = 2. Set both C and D to 0. Delete the item corresponding to the key A = 2, B = 2 from the hash table.

[0125] 6: (2, 2, 1, 2) does not match any key.

[0126] 7: (2, 2, 0, 2) does not match any key.

[0127] The traversal is completed and the verification ends.

[0128] Finally, it is obtained that A<1>B<2> → D<2> and A<2>B<1> → D<2> hold, while A<1>B<2> → C<1> and A<2>B<2> → C<2>D<2> no longer hold.

[0129] The method of generating new differential dependencies is the same as above, so it will not be elaborated here.

[0130] Perform the same verification as above for other differential dependencies with the number of LHS attributes being 2.

[0131] Verify candidate differential dependencies with a larger number of LHS attributes according to the above steps until all differential dependencies are verified.

[0132] 5. Traverse the prefix tree to obtain the differential dependency set ∑′ as

[0133] B<1> → C<2>, C<1> → B<2>, A<1>B<2> → D<2>, A<2>B<1> → D<2>,

[0134] A<2>C<2> → D<2>, A<1>D<2> → B<2>, A<2>D<1> → B<2>C<1>,

[0135] B<1>C<2> → A<2>, B<1>D<2> → A<2>, B<2>D<1> → C<1>,

[0136] C<2>D<2> → A<2>, A<2>B<2>C<2> → D<2>

[0137] As can be seen from the results, some originally established differential dependencies have failed. For example, A<2>→C<2>, but A<2>C<2>→D<2> holds. This indicates that among all the existing employees, not all employees with an age difference within differential interval 2 (a difference of 1 - 2 years) have a salary difference within differential interval 2 (a difference of 1000 - 2000 yuan). However, if the condition that the work experience difference is within differential interval 2 (a difference of 1 - 2 years) is added, then the salary difference between them can still meet differential interval 2, that is, remain within 1000 - 2000 yuan. Through these results, the company can make decisions on the rationality of salary setting and thus make timely adjustments.

Claims

1. A method for discovering differential dependencies in a relational database for incremental scenarios, characterized in that It includes the following steps: Step 1, construct a differential dependency prefix tree according to the original differential dependency set ∑ of the database; Step 2, calculate the newly formed distance vectors according to the incremental data Δr to obtain a new distance vector set V; Step 3, construct a position list index for each attribute according to the new distance vector set V; Step 4, verify the original differential dependencies and the generated new differential dependencies; Step 5, traverse the differential dependency prefix tree to obtain a new differential dependency set ∑'; Among them, in the said step 1, differential dependencies are represented in the form of , where X and Y are attribute sets in relation R, and R = (A1, …, A i , …, A m ), A i represents the i-th attribute, and m represents the number of attributes in R; represents the left-hand side set LHS, represents the right-hand side set RHS; where j i and j k respectively represent the i-th and k-th differential intervals of attributes A i and A k ; A i <j i > and A k <j k > represent distance constraints; two tuples t1 and t2 in relation instance r satisfy the distance constraint A i <j i >(A k <j k ) indicating that the distance between t1 and t2 on A i (A k ) is within the j i (j k )-th differential interval; ∧ represents simultaneously satisfying distance constraints on multiple attributes, that is, for tuple pairs that satisfy , for any A i ∈ X (A k ∈ Y), it is necessary to satisfy A i <j i >(A k <j k >); the differential intervals are several intervals of interest divided by the user on each attribute, and in ascending order, the differential intervals are numbered starting from 1; when the differential dependency holds, it means that for any two tuples in relation instance r, if for any A i ∈ X, it satisfies A i <j i , then for each A k ∈ Y, it must also satisfy A k <j k >; The differential dependency set is a minimal cover set. In the minimal cover set, it satisfies: (1) The right-hand side set RHS is a single distance constraint, that is, it satisfies in the form of; (2) There does not exist any such that holds, where is said to be a generalization of 2. The method for discovering differential dependencies in a relational database for an incremental scenario according to claim 1, wherein The nodes of the differential dependency prefix tree are root nodes, attribute nodes, and interval nodes; there is only one root node, and both the attribute nodes and interval nodes have multiple layers. Among them, the first-layer attribute nodes are the children of the root node, the first-layer interval nodes are the children of the first-layer attribute nodes, the second-layer attribute nodes are the children of the first-layer interval nodes, the second-layer interval nodes are the children of the second-layer attribute nodes, and so on; the parent and child nodes are connected by two-way pointers. An attribute node and its child interval nodes represent a distance constraint, and the path from the root node to an interval node represents an LHS; The unique identifier of the attribute is stored in the attribute node. The set of attributes in all distance constraints of an LHS is called the attribute set corresponding to this LHS. The LHS attribute index maintains a node list for the attribute set corresponding to each LHS, and the last attribute node of all LHSs corresponding to this attribute set in the differential dependency prefix tree is recorded in this node list; The interval node records the interval number where its parent attribute node is located in the distance constraint, and stores a bit set and an array; the bit set and the array are used to represent the RHS situation corresponding to the differential dependency with the current path as the LHS. Each bit in the bit set and each position in the array correspond to an attribute; when case (1): the LHS represented by the current path can make this attribute satisfy a certain distance constraint as the RHS; or case (2): the LHS with the current path as the prefix can make this attribute satisfy a certain distance constraint as the RHS, the bit in the bit set is set to 1 at the corresponding attribute bit, otherwise it is set to 0; among them, when case (1) is satisfied, the interval number of the distance constraint is stored in the corresponding position in the array for this attribute, and in other cases, the value in the array defaults to 0; when there is a non-zero value in the array, this interval node represents a differential constraint, its LHS is the path from the root node to this interval node, and the RHS is the distance constraint composed of the attributes with non-zero values in the array and the corresponding intervals; 3. The method for discovering differential dependencies in a relational database for an incremental scenario according to claim 2, wherein The construction process of the differential dependency prefix tree is as follows: (1) Initialize the prefix tree: establish an empty node as the root node; (2) Traverse all differential dependencies in the original differential dependency set ∑, calculate the number of times each attribute appears in the LHS, and sort the attributes in descending order according to the number of times; (3) For each differential dependency dd, sort the distance constraints in its corresponding left-hand side set dd.LHS according to the order of attributes. The steps to add the differential dependency dd to the prefix tree are as follows: (3.1) Find the longest prefix path of dd.LHS in the current prefix tree, that is, find a path starting from the root node, whose represented LHS is the first n distance constraints of dd.LHS, such that n is the largest; return the last interval node in the longest prefix path, if there is no prefix, return the root node; set the returned node as node, and set i as the number of distance constraints included in the longest prefix path + 1; (3.2) If i is greater than the number of distance constraints in dd.LHS, directly go to the next step; Otherwise, determine whether the child nodes of node contain the attribute node of the i-th distance constraint in dd.LHS; if not, create a new attribute node for the i-th distance constraint in dd.LHS as a child node of node, add a pointer to this attribute node in the node list corresponding to the LHS attribute in the LHS attribute index, and create a new interval node as a child node of the attribute node; If it exists, only create a new interval node and use it as a child node of node; set the new interval node as node, increment the i value by 1, and repeat step (3.2); (3.3) Set the bits in the bit sets of each interval node on the path of dd.LHS to 1 on the right-hand side set dd.RHS attribute corresponding to the differential dependency dd, and set the value on the dd.RHS attribute in the array of the last interval node in dd.LHS to the corresponding interval number.

4. The method for discovering differential dependencies in a relational database for an incremental scenario according to claim 1, wherein In step 2, the newly formed distance vector is calculated through the following process: (1) Calculate the distance vector set V1 between the incremental data and the original data set: Indicates the distance between tuple t1 and tuple t2 on the m-th attribute A m ; Δr refers to the relational instance of the incremental data, and r refers to the relational instance of the original data set (2) Calculate the distance vector set V2 between the incremental data: (3) Replace the specific distance value with the differential interval number where the element in the distance vector is located, if it is not in all intervals, replace it with 0, to obtain the distance vector sets V1' and V2'; the differential interval is several intervals of interest divided by the user on each attribute, in ascending order, the differential intervals are numbered starting from 1; (4) The new distance vector set V = V1' ∪ V2', and define a unique number for each distance vector in V.

5. The method for discovering differential dependencies of a relational database for an incremental scenario according to claim 1, wherein In step 3, the process of constructing the position list index for each attribute is as follows: (1) Initialize the index: create an empty set for each attribute in all differential intervals, and each empty set is used as a group; (2) Traverse the new distance vector set V, and add the unique number of this distance vector to the corresponding group of each attribute according to the interval value of each distance vector on each attribute.

6. The method for discovering differential dependencies in a relational database for an incremental scenario according to claim 1, wherein In step 4, verify the differential dependencies in the original differential dependency set ∑ in the order of the number of LHS attributes, and the verification process is as follows: Step (1), set i = 1; Step (2), sequentially traverse the attribute nodes in the node list corresponding to the attribute set of size i in the LHS attribute index, and verify the differential dependencies represented by the child nodes of all attribute nodes in the node list at the same time; Step (3), increment the i value by 1, if there are still un-verified nodes in the differential dependency prefix tree, jump to step (2), otherwise go to the next step.

7. The method for discovering differential dependencies in a relational database for an incremental scenario according to claim 6, characterized in that, In the said step (2), for the attribute set A of size i l1 …A li , the verification steps are as follows: Step (2.1) traverses all As l1 …A li where the LHS attribute index is at j l1 child nodes of the attribute nodes in the group; The differential dependency formed by an interval node is A l1 <j l1 >…A li <j li >→A r1 <j r1 >…A rk <j rk > to create a left set array to store j l1 …j li and a right set array to store j r1 …j rk Create a hash table and add the left set array as the key and the right set array as the value to the hash table; Step (2.2) finds j in the position list index of A l1 corresponding grouping, traverses all distance vectors in the grouping, and maps through the distance vectors in the values of A l1 …A l1 in the hash table in step (2.1); if the corresponding left part set array and the values of the distance vectors in A li …A l1 …A li are found to match in the hash table, then the values in the corresponding right part set array are compared with the values on the corresponding attributes in the distance vector, and the values in all non-matching right part set arrays are set to 0. If all the values corresponding to the attributes in the right part set array are 0, then the corresponding key-value pair is deleted from the hash table; when the hash table is empty or all the distance vectors in the grouping have been traversed, a verification process ends. Step (2.3): For each attribute whose right-hand set array is set to 0 in step (2.2), add a distance constraint of a new attribute that does not belong to the LHS and RHS to its corresponding LHS, forming the next-layer differential dependency to be verified, i.e., the new differential dependency. If the new differential dependency is not included in the differential dependency prefix tree, or the differential dependency prefix tree does not contain a generalization of the new differential dependency, add the new differential dependency to the differential dependency prefix tree.

8. The method for discovering differential dependencies in a relational database for an incremental scenario according to claim 1, wherein In step 5, for the incremental data Δr, obtain the differential dependencies that hold in the database after the increment. Using the attributes of the right-hand set of the differential dependencies as the decision targets and the attributes of the left-hand set as the decision variables, formulate the distance constraints on the decision targets for the data after the increment according to the distance constraints on the decision variables, and use this as the new rule that holds in the database.

Citation Information

Patent Citations

  • Differential privacy trajectory data protection method based on prefix tree

    CN110727958A

  • Sliding window-based frequent item set parallel incremental mining method

    CN114691749A