A method and an electronic device for data quality repair that satisfy negative constraints

By constructing an index based on a negative constraint set, only detecting tuple pairs between relational instances and incremental relational instances, the problem of inefficient repair of incremental data in the prior art is solved, and fast indexing and efficient repair are achieved.

CN113901050BActive Publication Date: 2025-05-30SHANGHAI FUJIA INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111178933.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-08
Publication Date
2025-05-30
Estimated Expiration
2041-10-08

AI Technical Summary

Technical Problem

The prior art uses low rapid repair of incremental data when processing large-scale consistent data sets, resulting in reduced efficiency of conflicting data and repairing data.

Method used

By building an index based on a negative constraint set, only tuple pairs between the relational instance and the incremental relational instance are detected, tuple pairs that violate the negative constraints are found, and the index is updated after incremental repairs, for fast indexing and efficient repairs.

Benefits of technology

Improve the repair efficiency of incremental data, narrow the detection range, and find tuple pairs that violate negative constraints through indexes, achieving rapid finding and repairing conflict data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113901050B_ABST
    Figure CN113901050B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing, and discloses a method and an electronic device for data quality repair that satisfy negative constraints, including the following steps: Step 1, construct an index for the relational instance I based on the negative constraint set Σ, where the negative constraint set Σ and the relational instance I are non-empty; Step 2, detect the tuple pairs (s, t) between the relational instance I and the incremental relational instance ΔI, and find the tuple pairs (s, t) that violate the negative constraints, where s ∈ I and t ∈ ΔI; Step 3, detect the tuple pairs (s, t) generated within the incremental relational instance ΔI, and find the tuple pairs (s, t) that violate the negative constraints, where s ∈ ΔI and t ∈ ΔI; Step 4, repair the incremental relational instance ΔI according to the tuple pairs (s, t) that violate the negative constraints detected in Step 2 and Step 3 to obtain the incremental repair instance ΔI'; Step 5, update the index constructed in Step 1 based on the union of the relational instance I and the incremental repair instance ΔI' with respect to the negative constraint set Σ.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a data quality repair algorithm based on negative constraints. Background Art

[0002] As is well known, due to reasons such as errors, missing values, and conflicts, data in real life is often inconsistent, which makes it essential to explore efficient and effective data quality management. Many data repair technologies have emerged in the academic community. Data repair usually uses constraints, and common constraints include unique column combinations (UCCs), functional dependencies (FDs), conditional functional dependencies (CFDs), and negative constraints (DCs), etc. The method adopted is to given a relational instance I that violates the data dependency ∑, modify the relational instance I to obtain a new relational instance I' such that the new relational instance I' satisfies the data dependency ∑. Among them, negative constraints have strong expressiveness and are sufficient to include many other data dependencies. Therefore, in order to improve data quality, negative constraints have been well applied in data repair.

[0003] Since the existing consistent data sets may be very large in volume, the original full-scale data does not distinguish between consistent data and incremental data, and putting all the data together for data conflict detection greatly reduces the efficiency of finding and repairing conflict data. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and an electronic device for data quality repair that satisfy negative constraints to solve the problem of rapid repair of incremental data.

[0005] To solve the above technical problems, an embodiment of the present invention provides a method for data quality repair that satisfies negative constraints, including:

[0006] Step 1, constructing an index for the relational instance I based on the negative constraint set Σ, where the negative constraint set Σ and the relational instance I are non-empty, and the negative constraint set Σ is a set of negative constraints;

[0007] Step 2, detecting the tuple pairs (s, t) between the relational instance I and the incremental relational instance ΔI, and finding the tuple pairs (s, t) that violate the negative constraints, where s ∈ I and t ∈ ΔI;

[0008] Step 3, detecting the tuple pairs (s, t) generated inside the incremental relational instance ΔI, and finding the tuple pairs (s, t) that violate the negative constraints, where s ∈ ΔI and t ∈ ΔI;

[0009] Step 4, repairing the incremental relational instance ΔI according to the tuple pairs (s, t) that violate the negative constraints detected in Step 2 and Step 3 to obtain an incremental repair instance ΔI';

[0010] Step 5: Update the index constructed in Step 1 for the union of the relationship instance I and the incremental repair instance ΔI' based on the negative constraint set Σ.

[0011] An embodiment of the present invention further provides an electronic device, including: at least one processor; and,

[0012] a memory communicatively connected to the at least one processor; wherein,

[0013] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method for repairing data quality that satisfies negative constraints as described above.

[0014] In the embodiment of the present invention, first, an index is constructed for the relationship instance I based on the negative constraint set Σ. At the same time, the existing relationship instance I and the incremental relationship instance ΔI are processed separately. Only the tuple pairs (s, t) between the relationship instance I and the incremental relationship instance ΔI are detected, and the tuple pairs (s, t) that violate the negative constraints are found, without detecting and processing the relationship instance I. Therefore, the detection range is narrowed. Further, when finding the tuple pairs (s, t) that violate the negative constraints, the tuple pairs that violate the negative constraints are found by means of the index, which further improves the efficiency. Moreover, after the incremental relationship instance ΔI is repaired, the index is updated for the union of the relationship instance I and the incremental repair instance ΔI' based on the negative constraint set Σ, so that the index always conforms to the relationship instance, thereby enabling the conflict data to be quickly found and the conflict data to be quickly repaired. That is to say, through incremental data detection, taking advantage of the fact that the existing consistent data itself will no longer generate conflicts, a fast index can be established, and only the conflicts between the consistent data and the incremental data and the conflicts within the incremental data can be detected, thereby greatly improving the efficiency of repairing data. Also, because negative constraints have strong expressiveness and can include rules such as functional dependencies, conditional functional dependencies, and unique column combinations, various types of conflicts generated by rules can be solved. That is to say, through incremental data detection, taking advantage of the fact that the existing consistent data itself will no longer generate conflicts, a fast index can be established, and only the conflicts between the consistent data and the incremental data and the conflicts within the incremental data can be detected, thereby greatly improving the efficiency of repairing data.

[0015] Preferably, the index in Step 1 is constructed based on the relationship instance I through the equality predicates and inequality predicates in the rule φ, and the index is a subset of the predicates, where the rule φ ∈ {<, ≤, >, ≥, =, ≠}.

[0016] Preferably, the index in step 5 is an index constructed by the equality predicates and inequality predicates in the rule φ based on the union of the relational instance I and the incremental repair instance ΔI', and the index is a subset of the predicates.

[0017] That is, update the index based on the negative constraint set Σ for the union of the relational instance I and the incremental repair instance ΔI' so that the index always conforms to the relational instance.

[0018] Preferably, in step 2, the step of detecting the tuple pairs (s, t) that violate the negative constraints between the relational instance I and the incremental relational instance ΔI, where s ∈ I and t ∈ ΔI:

[0019] Step 21: Select a tuple t from the incremental relational instance ΔI, and filter the tuples s that conform to the index from the relational instance I according to the index constructed in step 1 as the detection candidates;

[0020] Step 22: Check whether the detection candidates conform to other predicates in the rule φ other than the index;

[0021] Step 23: If the tuple pair (s, t) satisfies all the predicates in the rule φ, then the tuple pair (s, t) is the tuple pair (s, t) that violates the negative constraints.

[0022] Preferably, in step 3, use the full-scale detection method to detect the tuple pairs (s, t) that violate the negative constraints inside the incremental relational instance ΔI, where s ∈ ΔI and t ∈ ΔI.

[0023] That is, when detecting the tuple pairs (s, t) that violate the negative constraints between the relational instance I and the incremental relational instance ΔI, the index constructed in step 1 can be used to filter out the tuples s that conform to the index from the relational instance I as the detection candidates, and then the tuple pairs (s, t) that violate the negative constraints can be quickly found from the detection candidates.

[0024] Preferably, in step 4, the step of repairing the incremental relational instance ΔI:

[0025] Step 41: Modify one of the attribute values in the tuple s or the tuple t through an arithmetic formula to obtain a new tuple s or a new tuple t, so that the new tuple pair (s, t) does not satisfy at least one of the predicates in the rule φ;

[0026] Step 42: Detect whether the new tuple s or the new tuple t obtained in step 41 conflicts with all other tuples. If there is a conflict, set the modified attribute value in the new tuple s or the new tuple t to a new value FV. If there is no conflict, retain the modified attribute value in step 41;

[0027] Step 43: If the new tuple s or the new tuple t that does not cause a conflict cannot be obtained in Steps 41 and 42, modify the attribute value of the tuple s or the tuple t to the new value FV.

[0028] Preferably, the new value FV is a value that does not satisfy any predicate in the rule φ or NULL.

[0029] Preferably, the initial value of the relationship instance I in Step 1 conforms to data consistency.

[0030] Preferably, the union of the relationship instance I and the incremental repair instance ΔI' in Step 5 conforms to the principle of minimum cost.

[0031] That is to say, the initial value of the relationship instance I and the union of the repaired relationship instance I and the incremental repair instance ΔI' conform to data consistency, and the union of the relationship instance I and the incremental repair instance ΔI' conforms to the principle of minimum cost.

[0032] In summary, first, an index is constructed for the relationship instance I based on the negative constraint set Σ. At the same time, the existing relationship instance I and the incremental relationship instance ΔI are processed separately. Only the tuple pairs (s, t) between the relationship instance I and the incremental relationship instance ΔI are detected, and the tuple pairs (s, t) that violate the negative constraints are found, without detecting and processing the relationship instance I. Therefore, the detection range is narrowed. Furthermore, when finding the tuple pairs (s, t) that violate the negative constraints, the tuple pairs that violate the negative constraints are found by means of the index, which further improves the efficiency. Moreover, after the incremental relationship instance ΔI is repaired, the index is updated for the union of the relationship instance I and the incremental repair instance ΔI' based on the negative constraint set Σ, so that the index always conforms to the relationship instance, thereby achieving the purpose of improving the data repair efficiency. Also, because the negative constraints have strong expressiveness and can include rules such as functional dependencies, conditional functional dependencies, and unique column combinations, various types of conflicts generated by rules can be solved. Description of the Drawings

[0033] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. In the drawings:

[0034] Figure 1 is the flowchart of the method for data quality repair that satisfies negative constraints according to the embodiment of the present invention;

[0035] Figure 2 is the architecture diagram of the method for data quality repair that satisfies negative constraints according to the embodiment of the present invention;

[0036] Figure 3aImpact graph of the number of repaired consistent data pairs added to the Tax dataset in the embodiments of the present invention;

[0037] Figure 3b Impact graph of the repair time of added consistent data pairs to the Tax dataset in the embodiments of the present invention;

[0038] Figure 3c Impact graph of the number of units to be repaired for added incorrect data pairs to the Tax dataset in the embodiments of the present invention;

[0039] Figure 3d Impact graph of the repair time of added incorrect data pairs to the Tax dataset in the embodiments of the present invention;

[0040] Figure 3e Impact graph of the number of units to be repaired for added incremental data pairs to the Tax dataset in the embodiments of the present invention;

[0041] Figure 3f Impact graph of the repair time of added incremental data pairs to the Tax dataset in the embodiments of the present invention. Detailed implementation manners

[0042] The present invention provides many applicable creative concepts, and these creative concepts can be embodied in a large number in specific contexts. The specific embodiments described in the following embodiments of the present invention are only illustrative descriptions of the specific implementation manners of the present invention, and do not constitute a limitation on the scope of the present invention.

[0043] Figure 1 It is a flowchart of an incremental data repair algorithm, and the specific steps are as follows:

[0044] Step 1: An electronic device, such as a computer device, constructs an index for a relationship instance I based on a set of negative constraints Σ. The set of negative constraints Σ and the relationship instance I are non-empty, and the set of negative constraints Σ is a set of negative constraints;

[0045] The relationship instance I is a relationship instance I generated based on the R(A, B,...) relationship schema. The relationship instance I is composed of a set of tuples t, and I(t, A) represents the attribute A of the tuple t.

[0046] The negative constraints, the rules φ of each negative constraint constitute the set of negative constraints Σ. The rules φ of the negative constraints are all negations connected by predicates and are defined in a universally quantified first-order logic form:

[0047]

[0048] where P i (i ∈ [1, m]) represents a predicate and is in the form of or where v1 , v 2 in the form of t x [A], where x ∈ {α, β}, A is an attribute of the tuple t, and c is a constant is a rule composed of predicates

[0049] It should be noted that those skilled in the art can understand that the negative constraint set involved in this embodiment can be a conventional and well-known rule (such as for census data, obtaining the negative constraint set through ID number information, etc.), which can be obtained from experts in the relevant field familiar with the data, or can be obtained through methods such as data mining. The negative constraint set of this embodiment can also be obtained through reasonable rules applicable to this embodiment.

[0050] Step 2: The computer device detects the tuple pairs (s, t) between the relationship instance I and the incremental relationship instance ΔI, and finds the tuple pairs (s, t) that violate the negative constraints, where s ∈ I and t ∈ ΔI;

[0051] For the incremental relationship instance ΔI, tuples are inserted into the relationship instance I, and the set of inserted tuples is the incremental relationship instance ΔI. The tuples of the relationship instance I and the tuples of the incremental relationship instance ΔI can form tuple pairs (s, t), where s ∈ I and t ∈ ΔI. Correspondingly, the tuple t in the incremental relationship instance ΔI is repaired, and the incremental repair instance ΔI' is obtained after repair.

[0052] Step 3: The computer device detects the tuple pairs (s, t) generated inside the incremental relationship instance ΔI, and finds the tuple pairs (s, t) that violate the negative constraints, where s ∈ ΔI and t ∈ ΔI;

[0053] Step 4: The computer device repairs the incremental relationship instance ΔI according to the tuple pairs (s, t) that violate the negative constraints detected in Step 2 and Step 3 to obtain the incremental repair instance ΔI';

[0054] For the repair of the conflict of the tuple pair, when a tuple pair (s, t) satisfies all the predicates in the rule φ, the tuple pair (s, t) violates the negative constraint, that is, the tuple pair (s, t) generates a conflict. In this case, only one attribute value in s or t needs to be modified, and the new tuple pair (s, t) after modification does not satisfy at least one predicate in the rule φ, so the new tuple pair (s, t) after modification is repaired.

[0055] Step 5: The computer device updates the index constructed in Step 1 based on the negative constraint set Σ for the union of the relationship instance I and the incremental repair instance ΔI'.

[0056] The incremental repair instance ΔI' and the union of the relationship instance I and the incremental repair instance ΔI' satisfy data consistency, and the repaired data satisfies the principle of minimum cost.

[0057] Regarding the data consistency, when a tuple pair (s, t) satisfies all the predicates in the rule φ, the tuple pair (s, t) violates the negative constraint, that is, the tuple pair (s, t) has a conflict. If any tuple pair (t, s) in the relation instance I does not violate the negative constraint, then the relation instance I is consistent with respect to the rule φ, denoted as I ╞ φ. For a set of negative constraints Σ, if and only if all φ ∈ Σ and I ╞ φ, then it is considered that I ╞ Σ, that is, I has data consistency with respect to Σ.

[0058] Regarding the principle of minimum cost, the repair with the smallest difference between all the relation instances I in the repair and the repair instance I':

[0059]

[0060] w(t, A) is the weight on the attribute A of the tuple t, and the distance function dist(I(t, A), I'(t, A)) is the difference in the value of the attribute t[A] between the relation instance I and the repair instance I'.

[0061] It should be noted that the method for repairing data quality that satisfies negative constraints in this embodiment is applicable to numerical and character relational data, and has a wide range of application scenarios, such as aircraft flight data, hospital information data, personal purchase records, etc. For example, for census data, the method for repairing data quality that satisfies negative constraints in this embodiment can repair data recording errors caused by human negligence of recorders.

[0062] Figure 2 It is the architecture diagram of the incremental data repair algorithm, and the relationship of the processing modules is as follows:

[0063] The input includes a set of negative constraints Σ, a relation instance I that is consistent with the set of negative constraints Σ, and a set of incremental relation instances ΔI.

[0064] The set of negative constraints Σ creates an index according to the relation instance I, and the relation instance I and the incremental relation instance ΔI quickly detect the conflicts between the relation instance I and the incremental relation instance ΔI based on the set of negative constraints Σ through the index.

[0065] The incremental relation instance ΔI detects the conflicts inside the incremental relation instance ΔI based on the set of negative constraints Σ.

[0066] Repair the detected conflicts between the relation instance I and the incremental relation instance ΔI and the detected conflicts inside the incremental relation instance ΔI to obtain an incremental repair instance ΔI' as the output.

[0067] The union of the incremental repair instance ΔI' and the relation instance I updates the index based on the set of negative constraints Σ.

[0068] In summary, first, an index is constructed for the relational instance I based on the negative constraint set Σ. At the same time, only the tuple pairs (s, t) that violate the negative constraints are found from the tuple pairs (s, t) between the relational instance I and the incremental relational instance ΔI, without detecting and processing the relational instance I, thus narrowing the detection scope. Furthermore, when finding the tuple pairs (s, t) that violate the negative constraints, the tuple pairs that violate the negative constraints are found by means of the index, which further improves the efficiency. Finally, the index is updated based on the negative constraint set Σ for the union of the relational instance I and the incremental repair instance ΔI', so that the index always conforms to the relational instance, thereby achieving the purpose of improving the efficiency of repairing data.

[0069] Figure 1 The construction method of the index in step 1 is as follows:

[0070] The index is constructed for the relational instance I based on the negative constraint set Σ in step 1. The rules φ in the negative constraint set Σ are divided into two categories. One category is for the equality predicates in the rules, i.e., {=}, and the other category is for the inequality predicates, i.e., {<, ≤, >, ≥, ≠}. The index is constructed according to the classification of the predicates, and the index is a subset of the predicates.

[0071] Index for equality predicates:

[0072] The index for equality predicates is an equi-index, and its form is t α [A]=t β [B]. Each equi-index is constructed for a single attribute A and is denoted as Index=(A). Specifically, the tuples are sorted on A, and the tuples with the same value on A are collected in an equivalence class. For a predicate P of the form t α [A]=t β [B] 1 If A≠B, Index=(A) and Index=(B) need to be created. Index=(P 1 ) represents the equi-index of the predicate P 1 . The retrieval efficiency of the equi-index is: for t∈ΔI, it takes O(log(|I|)) to identify s∈I such that s[A]=t[B] or t[A]=s[B], where |I| is the number of tuples in I.

[0073] Index for inequality predicates:

[0074] The index for inequality predicates has the form t α [A]θt β[B], where θ ∈ {<, ≤, >, ≥}. For a tuple t ∈ ΔI, if s[A] θ t[B] or t[A] θ s[B], a tuple s ∈ I needs to be obtained. Since a single inequality predicate usually has a low selectivity, it may result in a large result set of s. Therefore, an index is built by combining two predicates with inequality operators:

[0075] (P 1 = t α [A] θ 1 t β [B]) ^ (P 2 = t α [C] θ 2 t β [D]), where θ 1 , θ 2 ∈ {<, ≤, >, ≥}. Thus, the selectivity of the index is improved by the tuples s ∈ I that satisfy P 1 and P 2 . That is, sorted arrays are created on attributes A, B, C, and D respectively, and permutation arrays are constructed between attribute A and attribute B and between attribute C and attribute D. For t ∈ ΔI, it takes O(|I|) time to identify the tuple s ∈ I such that (t[A] θ 1 s[B]) ^ (t[C] θ 2 s[D]) or (s[A] θ 1 t[B]) ^ (s[C] θ 2 t[D]).

[0076] This paper regards the operator ≠ as the union of a pair of operators > and <. That is, the negative constraint

[0077] is violated if and only if or is violated. Therefore, indexes are built for and respectively, and the union of their results is used as the result of the rule φ.

[0078] Figure 1 The way to detect the conflict between I and ΔI in step 2 of

[0079] This paper uses index technology to speed up the detection of conflicts between s ∈ I and t ∈ ΔI. If s and t satisfy all the predicates in DCφ, then s and t violate φ. To find all the violating tuples s for a given t, this paper first identifies the candidate s by using the index on (some) predicates of φ, and then obtains the final set of s by checking the remaining predicates on the candidates. Building indexes helps to improve efficiency because the number of candidates is usually much smaller than the number of tuples in I.

[0080] Figure 1 The method for detecting conflicts generated inside ΔI in step 3 is as follows:

[0081] Use the full - scale detection method for detection.

[0082] Figure 1 The method for repairing the incremental relationship instance ΔI in step 4 is as follows:

[0083] Repairing the incremental relationship instance ΔI is heuristic, and the repair steps are as follows:

[0084] Step 41: Modify one attribute value in tuple s (or tuple t) through an arithmetic formula to obtain a new tuple s (or new tuple t), so that the new tuple pair (s, t) does not satisfy at least one predicate in rule φ;

[0085] That is, the tuple pair (s, t) violates rule φ because the tuple pair (s, t) satisfies all predicates in rule φ. Therefore, only one attribute value in tuple s or tuple t needs to be modified so that the tuple pair (s, t) does not satisfy at least one predicate in rule φ, then the tuple pair (s, t) does not violate rule φ. Modify the attribute value to another value through an arithmetic formula so that the tuple pair (s, t) does not satisfy at least one predicate in rule φ.

[0086] Step 42: Detect whether the new tuple s (or new tuple t) obtained in step 41 conflicts with all other tuples. If there is a conflict, set the modified attribute value in the new tuple s (or new tuple t) to the new value FV. If there is no conflict, retain the modified attribute value in step 41;

[0087] That is, although one attribute value is modified in step 41 to solve the current conflict, other conflicts may occur. To prevent other conflicts from occurring, when modifying the current attribute value, it is necessary to consider whether the current attribute value in the tuple and all other tuples will conflict. If a new conflict will occur, do not modify the attribute value, or set the attribute value to the new value FV. When no new conflict will occur, then modify the current attribute value. The repair heuristically identifies the maximum number of predicates that can be satisfied simultaneously from the modified new tuple s (or new tuple t) and combines them with the original predicates to form a new rule φ.

[0088] Step 43: If the new tuple s (or new tuple t) that does not generate conflicts cannot be obtained in steps 41 and 42, then modify the attribute value of tuple s (or tuple t) to the new value FV.

[0089] That is, if the new tuple s (or new tuple t) that does not generate conflicts cannot be obtained in steps 41 and 42, then set the attribute value to the new value FV.

[0090] Figures 3a to 3fThe experimental results of the Tax dataset are analyzed as follows:

[0091] In the experiment, the data is a synthetic dataset Tax generated by a Tax data generator, which contains character and numerical attributes. Data conflicts in the dataset are injected according to the error conflict rate ρ, where ρ ∈ [1% - 10%], indicating that among 100 units of incremental data, ρ units have conflicts. The conflicting attributes are modified to other values to ensure the randomness of the conflicting attributes.

[0092] This paper compares the incremental repair method of this paper (denoted as Inc) with the holistic method (denoted as Holistic) and VFree (denoted as VFree).

[0093] As Figure 3a and Figure 3b show the influence of the size of the change source dataset I on repair. Figure 3a and Figure 3b The abscissa in them is the size of the change source dataset I. As the amount of consistent data increases, the number of repairs remains almost unchanged, the repair time shows an upward trend, and adding more evidence has little effect on the repair result. Therefore, the method of this paper is superior to other methods in terms of both the number of repairs and the repair time.

[0094] As Figure 3c and Figure 3d show the influence of the size of the change error rate ρ on repair. Figure 3c and Figure 3d The abscissa in them is the size of the change error rate. As the amount of incorrect data increases, the number of units that need to be modified to repair the error shows an upward trend, and the repair time shows an upward trend. Therefore, the method of this paper is superior to other methods in terms of both the number of repairs and the repair time.

[0095] As Figure 3e and Figure 3f show the influence of the size of the change incremental dataset IncI on repair. Figure 3e and Figure 3f The abscissa in them is the size of the change incremental dataset IncI. As the amount of incremental data increases, the number of units that need to be modified to repair the error shows an upward trend, and the repair time shows an upward trend. Therefore, the method of this paper is superior to other methods in terms of both the number of repairs and the repair time.

[0096] Another embodiment of this application relates to an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for repairing data quality that satisfies negative constraints in the above embodiment.

[0097] Among them, the memory and the processor are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0098] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.

[0099] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by the processor, the above method embodiment is implemented.

[0100] That is, those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0101] Those of ordinary skill in the art can understand that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims.

Claims

1. A method for data quality repair that satisfies negative constraints, characterized in that, it is applied to numerical and character relational data, and the relational data is census data. The method includes the following steps: Step 1: Build an index for the relational instance I based on the negative constraint set Σ. The negative constraint set Σ and the relational instance I are non-empty, and the negative constraint set Σ is a set of negative constraints; wherein, the negative constraint set Σ is obtained through ID card number information; Step 2: Detect the tuple pairs (s, t) between the relational instance I and the incremental relational instance ΔI, and find the tuple pairs (s, t) that violate the negative constraints, where s ∈ I and t ∈ ΔI; Step 3: Detect the tuple pairs (s, t) generated within the incremental relational instance ΔI, and find the tuple pairs (s, t) that violate the negative constraints, where s ∈ ΔI and t ∈ ΔI; Step 4: Repair the incremental relational instance ΔI according to the tuple pairs (s, t) that violate the negative constraints detected in Step 2 and Step 3 to obtain an incremental repair instance ΔI'; Step 5: Update the index constructed in Step 1 based on the union of the relational instance I and the incremental repair instance ΔI' with the negative constraint set Σ.

2. The method according to claim 1, characterized in that: the index in Step 1 is constructed based on the relational instance I through the equality predicate and inequality predicate in the rule φ, and the index is a subset of the predicates, where the rule φ ∈ {<, ≤, >, ≥, =, ≠}.

3. The method according to claim 2, characterized in that: the index in Step 5 is an index constructed based on the union of the relational instance I and the incremental repair instance ΔI' through the equality predicate and inequality predicate in the rule φ, and the index is a subset of the predicates.

4. The method according to claim 1, characterized in that: in Step 2, the steps of detecting the tuple pairs (s, t) that violate the negative constraints between the relational instance I and the incremental relational instance ΔI, where s ∈ I and t ∈ ΔI: Step 21: Select a tuple t in the incremental relational instance ΔI, and screen the tuples s that meet the index from the relational instance I according to the index constructed in Step 1 as detection candidates; Step 22: Check whether the detection candidates meet other predicates in the rule φ other than the index; Step 23: If the tuple pair (s, t) satisfies all the predicates in the rule φ, then the tuple pair (s, t) is a tuple pair (s, t) that violates the negative constraints.

5. The method according to claim 1, characterized in that: in Step 3, a full-scale detection method is used to detect the tuple pairs (s, t) that violate the negative constraints within the incremental relational instance ΔI, where s ∈ ΔI and t ∈ ΔI.

6. The method according to claim 1, characterized in that: in Step 4, the steps of repairing the incremental relational instance ΔI: Step 41: Modify an attribute value in tuple s or tuple t through an arithmetic formula to obtain a new tuple s or a new tuple t, so that the new tuple pair (s, t) does not satisfy at least one predicate in the rule φ; Step 42: Detect whether the new tuple s or new tuple t obtained in Step 41 conflicts with all other tuples. If there is a conflict, set the modified attribute value in the new tuple s or new tuple t to the new value FV. If there is no conflict, retain the attribute value modified in Step 41; Step 43: If the new tuple s or new tuple t that does not cause a conflict cannot be obtained in Step 41 and Step 42, modify the attribute value of the tuple s or tuple t to the new value FV.

7. The method according to claim 6, wherein: The new value FV is a value that does not satisfy any predicate in the rule φ or NULL.

8. The method according to claim 1, wherein: The initial value of the relational instance I in Step 1 conforms to data consistency.

9. The method according to claim 1, wherein: The union of the relational instance I and the incremental repair instance ΔI' in Step 5 conforms to the principle of minimum cost.

10. An electronic device, wherein, comprising: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for data quality repair that satisfies negative constraints as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Mechanisms for merging index structures in MOLAP while preserving query consistency

    CN108140024A

  • Incremental time series data conflict detection method, device and storage medium

    CN113282616A