Error table data interpretable repair method and device, and electronic device

By obtaining and retrieving context-related tuples of table error data, sampling representative error tuples and generating repair thinking chains and rule examples using large language models, the problem of table error data repair in the prior art depends on expert knowledge and lack of interpretability, and efficient, accurate and interpretable table error data repair is achieved.

CN119761320BActive Publication Date: 2025-06-17ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510253217.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-17
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Existing methods for repairing table error data are highly dependent on expert knowledge, and machine learning-based methods are difficult to implement when data is scarce or labeled, and lack interpretability, making it difficult for repair processes and results to be understood and trusted by users.

Method used

By obtaining the error table data information matrix, the table data error mask matrix and the error table data semantic embedding matrix, the tuple with the strongest correlation with each error data is retrieved, representative error tuples are sampled for repair, and repair thinking chains and rule examples are generated using a large language model, and the final repair results are voted for to achieve interpretability of data repair.

Benefits of technology

It improves the accuracy and efficiency of table error data repair, reduces the "illusion" problem of data repair in specific fields of large language models, realizes the interpretability of data repair, improves accuracy by 42%, and improves efficiency by 2.8 times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119761320B_ABST
    Figure CN119761320B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for interpretable repair of error table data, and an electronic device, including: obtaining an error table data information matrix, a table data error mask matrix, and an error table data semantic embedding matrix; according to the above three matrices, retrieving a number of tuples with the strongest association with each error data; sampling a number of representative error tuples in the error table data information matrix for repair; according to the repair and retrieval results of the representative error data, using a large language model to generate a repair thought chain for each representative error data; according to the repair thought chain of the representative error data, using a large language model to generate multiple rule instances for each repair rule; applying all rule instances to generate multiple repair candidate results for each error data, and voting to select the final result for each error data; for the error data not repaired by the rules, according to its retrieval result, using a large language model to obtain the repair result and the repair thought chain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of table data repair, and specifically relates to an interpretable repair method and device for incorrect table data, and an electronic device. Background Art

[0002] Traditional methods for repairing incorrect table data mainly rely on preset data rules, such as functional dependencies, conditional functional dependencies, and negative constraints, etc. These rules are usually formulated by domain experts based on a deep understanding of data characteristics and business logic. However, such methods highly rely on expert knowledge, requiring users to not only fully understand the data to be repaired, but also be familiar with the configuration of the data repair system to reasonably set and adjust the rules. In recent years, with the development of machine learning technology, some emerging methods automatically complete the data repair task by training a model, avoiding the need for users to manually provide specific rules, simplifying the repair process and improving the automation level. However, these machine learning-based methods usually rely on large-scale training data and are difficult to implement in cases where data is scarce or annotation is difficult, and the models, especially deep learning models, lack sufficient interpretability, resulting in the repair process and results being difficult for users to understand and trust. In addition, these methods perform poorly when dealing with undefined or novel error patterns because the model may not be able to effectively identify and handle abnormal situations beyond the scope of the training data. Summary of the Invention

[0003] To overcome the above technical problems, an interpretable repair method and device for incorrect table data, and an electronic device are implemented and provided in this application to ensure the fairness of prediction.

[0004] According to the first aspect of the embodiments of this application, an interpretable repair method for incorrect table data is provided, including:

[0005] S1: Obtain an incorrect table data information matrix, a table data error mask matrix, and an incorrect table data semantic embedding matrix;

[0006] S2: According to the incorrect table data information matrix, the table data error mask matrix, and the incorrect table data semantic embedding matrix, retrieve several tuples with the strongest association with each incorrect data;

[0007] S3: In the incorrect table data information matrix, sample several representative incorrect tuples for repair;

[0008] S4: According to the repair and retrieval results of the representative incorrect data, use a large language model to generate a repair thought chain for each representative incorrect data;

[0009] S5: According to the repair thought chain of the representative incorrect data, use a large language model to generate multiple rule instances for each repair rule;

[0010] S6: Apply all rule instances to generate multiple repair candidates for each error data, and vote to select the final result for each error data;

[0011] S7: For the error data not repaired by the rules, according to its retrieval results, use a large language model to obtain the repair results and the repair thought chain.

[0012] According to the second aspect of the embodiments of the present application, there is provided an explainable repair device for error table data, including:

[0013] An acquisition module, configured to acquire an error table data information matrix, a table data error mask matrix, and an error table data semantic embedding matrix;

[0014] A retrieval module, configured to retrieve a number of tuples with the strongest association with each error data according to the error table data information matrix, the table data error mask matrix, and the error table data semantic embedding matrix;

[0015] A sampling repair module, configured to sample a number of representative error tuples in the error table data information matrix for repair;

[0016] A repair thought chain generation module, configured to use a large language model to generate a repair thought chain for each representative error data according to the repair results and retrieval results of the representative error data;

[0017] A repair rule generation module, configured to use a large language model to generate multiple rule instances for each repair rule according to the repair thought chain of the representative error data;

[0018] A rule repair module, configured to apply all rule instances to generate multiple repair candidates for each error data, and vote to select the final result for each error data;

[0019] A large language model repair module, configured to, for the error data not repaired by the rules, according to its retrieval results, use a large language model to obtain the repair results and the repair thought chain.

[0020] According to the third aspect of the embodiments of the present application, there is provided an electronic device, including:

[0021] One or more sensors;

[0022] One or more processors;

[0023] A memory, configured to store one or more programs;

[0024] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.

[0025] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:

[0026] The present invention designs a context-related tuple retrieval method for the task of repairing incorrect data in tables, which can effectively identify and extract context information closely related to the incorrect data, and improve the accuracy of data repair by large language models in a retrieval-enhanced manner; by sampling representative incorrect tuples and manually repairing them by users, combined with the repair thought chain generated by large language models, the large language model can learn the repair strategies for typical errors in the dataset, thereby reducing the "hallucination" problem of the large language model when facing data in a specific field and further improving the repair accuracy; by automatically generating data repair rules and abstracting some frequently used repair strategies into data rules, frequent calls to large language models are avoided, thereby improving the efficiency of data repair and reducing costs; the interpretability of data repair is achieved in two ways: on the one hand, the data repair rules are used to provide clear repair bases; on the other hand, for incorrect data not repaired by the rules, the repair reasons are elaborated in the form of a thought chain. The prediction accuracy of this method has been improved by 42% in terms of accuracy and 2.8 times in terms of efficiency compared with the current optimal algorithm for repairing incorrect data in tables. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0028] Figure 1 is a flowchart of an interpretable repair method for incorrect table data provided by an embodiment of the present application.

[0029] Figure 2 is a block diagram of an interpretable repair device for incorrect table data provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0031] The interpretable repair method for incorrect table data provided by the present application is applicable to the interpretable repair of incorrect table data in almost all fields, such as medical data and traffic data, etc. The following takes medical data as an example for detailed elaboration.

[0032] Figure 1 is a flowchart of a method for interpreting error table data provided by the present application according to an exemplary embodiment. As Figure 1 shown, the method may include the following steps:

[0033] S1: Obtain an error table data information matrix, a table data error mask matrix, and an error table data semantic embedding matrix;

[0034] S11: Obtain the table data information and table data error information collected by the sensor;

[0035] Specifically, a medical sensor can be used to collect patients' health data, an environmental sensor can be used to monitor environmental data, and an industrial sensor can be used to collect the operating status and fault information of equipment, etc. In this embodiment, taking the use of a sensor to collect patients' health data as an example for specific elaboration. The health data of patients can have multiple fields, such as multiple characteristic fields of a patient's blood pressure, blood sugar level, heart rate, weight, height, etc.; due to reasons such as doctors filling in or omitting information incorrectly, and measurement equipment failures, there are errors in some health data fields. S12: According to the table data information and table data error information, construct an error table data information matrix and a table data error mask matrix;

[0036] Specifically, for an error table containing N tuples, each tuple containing d characteristic fields, its error table data information matrix can be expressed as , where each tuple can be expressed as . The error information of these data will be stored in a table data error mask matrix, and the table data error mask matrix can be expressed as , where corresponds to the error status of a single tuple, has a value range of {0, 1}, and ; when the th dimension data of is in error, ; otherwise

[0037] S13: Through a semantic embedding method, perform semantic embedding on each element in the error table data information matrix to generate an error table data semantic embedding matrix.

[0038] Specifically, the semantic embedding model can adopt a pre-trained Sentence-BERT model to perform semantic embedding on the data of each cell in the table, convert it into a high-dimensional normalized vector, and obtain an error table data semantic embedding matrix, expressed as , where is a tuple semantic embedding. The model can extract the semantic features of cell data and map them into a unified vector space, thereby capturing the deep semantic information of the data.

[0039] S2: According to the error table data information matrix, the table data error mask matrix, and the error table data semantic embedding matrix, retrieve several tuples with the strongest association with each error data;

[0040] S21: For each column to be repaired, calculate the normalized mutual information value between it and the other columns, and consider the columns with the normalized mutual information value greater than or equal to the specified threshold as relevant columns;

[0041] Specifically, for each column to be repaired , the context-related tuple retrieval module calculates the normalized mutual information value with the other columns. For any column in , its normalized mutual information with

[0042]

[0043] is expressed as: is and 's joint probability distribution, and are respectively and 's marginal probability distributions. For any column , if 's value is greater than or equal to the specified threshold, such as 0.5, then it is selected as a relevant column. This application filters out the irrelevant columns in the table data through normalized mutual information, thereby reducing the computational complexity of the subsequent retrieval process and enhancing the accuracy of the retrieval.

[0044] S22: Based on the weighted cosine similarity in the relevant columns, measure the similarity between the other tuples and the tuple to be repaired, and preliminarily sort the tuples according to the similarity, and retrieve several tuples with the top rankings;

[0045] Specifically, for any tuple ( ) in , its weighted cosine similarity with

[0046]

[0047] is expressed as: and are respectively the tuples and the error status of the elements in column is the th column of the table, is the column and is the normalized mutual information value of and respectively are the semantic embeddings of the elements in the and th column of the tuples

[0048] Calculate the similarity between the tuple to be repaired and other tuples using the above weighted cosine similarity and sort them, and retrieve several tuples with the highest similarity. This application designs the weighted cosine similarity considering the data error status, the normalized mutual information value between the column where the data is located and the column to be repaired, and the semantic embedding of the data. Use the error status weight to reduce the impact of incorrect data on the retrieval, adjust the weight according to the mutual information value between columns to ensure the dominant position of important columns in the similarity calculation, capture more complex data associations through semantic embedding, accurately measure the similarity between tuples, so that each tuple is accurately sorted according to the similarity with the tuple to be repaired.

[0049] S23: Further sort the preliminarily sorted tuples according to the data quality to finally determine several tuples with the highest ranking, that is, the tuples most strongly associated with the incorrect data.

[0050] Specifically, re - sort the tuples with the same weighted cosine similarity, and select a few tuples with the highest ranking. The sorting basis is:

[0051] Judge whether the value of the column to be repaired is correct. If the value of the column to be repaired in this tuple is correct, then give priority to selecting this tuple;

[0052] If the values of the columns to be repaired in these tuples are all the same (that is, all correct or all incorrect), then give priority to selecting the tuple containing more correct data.

[0053] After the sorting is completed, retrieve several tuples with the highest ranking as the context - related tuples of an incorrect data. Re - sort the tuples with the same weighted cosine similarity according to the data quality, so as to improve the accuracy of the sorting and further help retrieve the context information most relevant to the tuple to be repaired.

[0054] S3: Sample several representative incorrect tuples in the incorrect table data information matrix for repair; this step may include the following sub - steps:

[0055] S31: Group the tuples using a clustering algorithm. The user specifies the number of clusters for clustering, and all tuples are divided into several clusters.

[0056] Specifically, the K-Means algorithm can be used to cluster all tuples, where the number of clusters for clustering is specified by the user, and all tuples are divided into several clusters.

[0057] S32: Use a greedy algorithm to select one tuple from each cluster as a representative error tuple, such that the selected representative error tuples can cover as many error columns as possible and contain more error data.

[0058] Specifically, use the greedy algorithm to select one representative tuple from each cluster. The selection priority is as follows:

[0059] Cover as many error columns as possible;

[0060] Cover as many error cells as possible.

[0061] This application designs a greedy algorithm such that the selected representative tuples can cover as many error columns as possible and contain as much error data as possible, so that the representative tuples can cover a more comprehensive range of error types, thereby helping to generate a more comprehensive error repair strategy and improve error repair accuracy.

[0062] S33: Hand over the selected representative error tuples to the user for manual repair to obtain the corresponding correct results.

[0063] Specifically, the finally selected representative error tuples will be manually corrected by the user to obtain the correct values of these tuples.

[0064] S4: According to the repair and retrieval results of the representative error data, use a large language model to generate a repair thought chain for each representative error data; this step may include the following sub-steps:

[0065] S41: Based on the representative error tuples repaired by the user, construct a prompt for generating a repair thought chain for each error column.

[0066] Specifically, the prompt for generating a repair thought chain mainly includes the following parts:

[0067] i. Task description: Introduce the goal of the repair thought chain generation task, which is to generate a thought chain for error repair.

[0068] ii. General examples: Generate examples by randomly introducing four common error types (such as missing values, spelling mistakes, format issues, and constraint violations) in a dataset, such as the Chicago Food dataset. These examples are used as context demonstrations to guide the large language model to automatically generate a repair thought chain.

[0069] iii. Representative tuple information: Provide detailed information about representative tuples, including error tuples and error values.

[0070] iv. Retrieved relevant tuples: Retrieve tuple information related to the representative tuples through a context - related tuple retriever to serve as a reference context for repair.

[0071] v. Correction results provided by the user: Record the correction information of the user for the error columns of the representative tuples, which is used to guide the generation of the repair thought chain.

[0072] S42: Generate prompting words using the above - mentioned repair thought chain, and call the large - language model to obtain the repair thought chain for each error data.

[0073] Specifically, by using the above - mentioned repair thought chain to generate prompting words, call the large - language model to automatically generate the repair thought chain for each error data. This application uses the large - language model to automatically generate the repair thought chain of representative error data, revealing the repair strategy of representative error data, reducing the cost and potential inconsistencies of manually generating the repair thought chain, and providing guidance for subsequent rule generation and error repair.

[0074] S5: According to the repair thought chain of the representative error data, use the large - language model to generate multiple rule instances for each repair rule; this step may include the following sub - steps:

[0075] S51: For each specified repair rule, use a prompting - word paraphrasing tool to generate multiple different versions of repair - rule generation prompting words according to some representative tuples and their repair thought chains;

[0076] Specifically, first randomly select half of the tuples from the representative tuple set of the dirty data as the training set for generating repair rules. Then, for each rule type in the repair - rule description set, use the prompting - word paraphrasing tool to reconstruct the specified task description to generate multiple different versions of repair - rule generation task descriptions to guide the large - language model to generate a candidate set of possible repair rules. On this basis, by combining the generated task descriptions with the representative tuple repair workflow, form the repair - rule generation prompting words.

[0077] S52: Use the above - mentioned repair - rule generation prompting words to call the large - language model to generate specific instances of the repair rules;

[0078] Specifically, by using the above - mentioned repair - rule generation prompting words, call the large - language model to automatically generate repair - rule instances.

[0079] S53: Verify each rule on all representative tuples and all correct tuples, and only retain the rule instances that pass the verification.

[0080] Specifically, all the labeled representative tuples and the detected correct tuples are selected as the validation set, and the repair rules that can correctly repair all the dirty data values without destroying the correct data are retained to form the final set of repair rules. In this application, a large language model is used to extract repair rules from the repair process of representative tuples, thereby reducing the number of calls to the large language model in the subsequent repair process, reducing the computational cost, and providing an explanation for the error repair result in the form of repair rules.

[0081] S6: Apply all rule instances to generate multiple repair candidates for each error data, and vote to select the final result for each error data; this step includes the following sub-steps:

[0082] S61: For each column to be repaired, first execute all the verified rule instances in parallel to generate multiple interpretable repair results for each error data;

[0083] Specifically, for each column to be repaired, first execute all the verified rule instances in parallel to generate multiple interpretable repair results for each error data.

[0084] S62: Among the multiple interpretable repair results, select the repair result supported by the majority of rules as the final repair result for each error data.

[0085] Specifically, among the multiple interpretable repair results, select the repair result supported by the majority of rules as the final repair result for each error data. The final repair result is selected through a voting strategy, effectively reducing the risk of misrepair caused by the deviation of a single rule and improving the reliability of the repair result.

[0086] S7: For the error data not repaired by the rules, according to its retrieval result, use the large language model to obtain the repair result and the repair thought chain. This step may include the following sub-steps:

[0087] S71: Construct an interpretable data repair prompt for each error data not repaired by the rules based on the representative error tuple and its repair thought chain;

[0088] Specifically, the interpretable data repair prompt includes the following five components:

[0089] Task description: Introduce the target task of generating the repair result and its detailed repair thought chain;

[0090] ii. General repair examples: Based on four dirty tuples with randomly injected errors in the Chicago food dataset, construct four general repair examples. Each example includes the dirty tuple, the dirty value, and the retrieved tuple, and outputs the repair result and the corresponding repair thought chain;

[0091] iii. Representative repair examples: For each representative tuple containing the target dirty column, construct an example with the representative tuple, the dirty value, and the retrieved tuple as input, and the repair result and the reasoning thought chain generated by the thought chain generation prompt as output;

[0092] iv. Error tuples: Provide the tuples containing dirty values in the target dirty column and their corresponding dirty values;

[0093] v. Context-related tuples: Select the k tuples most similar to the error tuple through the context-related tuple retrieval tool.

[0094] S72: Use the above interpretable data repair prompts to call the large language model to generate repair results and repair thought chains for each error data.

[0095] According to the interpretable data repair prompts, call the large language model to generate repair results and corresponding repair thought chains for all dirty tuples. The repair result is the predicted value of the dirty data, and the repair thought chain details the reasoning steps, considerations, and context information during the repair process. In this application, the large language model is used to generate repair results and repair thought chains for error data through methods such as retrieval augmentation, thought chain prompting, and few-shot prompting, reducing the "hallucination" problem during the repair process and improving the accuracy of the repair results.

[0096] The interpretable repair method for incorrect tabular data of the present invention is implemented and run on the Ubuntu 20.04 system of a server with an Intel core of 2.30 GHz and 512 GB of memory. The accuracy and efficiency of the system are tested under datasets containing data from different fields, different error rates, different error types, and different sizes. The dataset introduction is shown in Table 1. For the measurement of system accuracy, we use precision (P), recall (R), and F1 value (F1). The calculation formula for precision is , and the calculation formula for recall is , where is the number of errors correctly repaired by the system, is the number of errors mis-repaired by the system, is the number of errors missed by the system. The larger the values of these two metrics, the higher the repair accuracy of the system; for the measurement of system efficiency, the running time metric is used, and the smaller the value of this metric, the faster the running speed.

[0097] The performance of the proposed fair prediction method for missing tabular data (i.e., ZeroEC) of the present invention was analyzed through experimental tests, and the results are shown in Tables 2 and 3. It can be seen that ZeroEC has improved the accuracy by more than 42% compared with the existing error tabular data repair systems (Scare, Holoclean, Baran, FMs). Because ZeroEC samples representative error data and learns the repair strategies for these representative errors. In the way of few-shot learning, these repair strategies are generalized to other data to be repaired, overcoming the "hallucination" problem existing in large language models when repairing data in specific fields. In addition, the average running efficiency of ZeroEC is more than 2.8 times that of the existing optimal system. This is because ZeroEC can complete data repair without training and automatically generates data repair rules for quick repair, thus avoiding calling the large language model for each error data one by one for repair.

[0098] Table 1: Dataset Introduction

[0099]

[0100] Table 2: Accuracy Comparison of Repair Systems on Each Dataset

[0101]

[0102] Table 3: Efficiency Comparison of Repair Systems on Each Dataset (Time Unit: Seconds)

[0103]

[0104] Corresponding to the foregoing embodiments of the fair prediction method for missing tabular data, the present application also provides embodiments of the fair prediction method for missing tabular data.

[0105] Figure 2 It is a block diagram of the device for the fair prediction method of missing tabular data of the present application. Referring to Figure 2 , the device includes:

[0106] An acquisition module, configured to acquire an error tabular data information matrix, a tabular data error mask matrix, and an error tabular data semantic embedding matrix;

[0107] A retrieval module, configured to retrieve a plurality of tuples with the strongest association with each error data according to the error tabular data information matrix, the tabular data error mask matrix, and the error tabular data semantic embedding matrix;

[0108] A sampling and repair module, configured to sample a plurality of representative error tuples in the error tabular data information matrix for repair;

[0109] A repair thought chain generation module, configured to use a large language model to generate a repair thought chain for each representative error data according to the repair result and retrieval result of the representative error data;

[0110] A repair rule generation module, configured to use a large language model to generate multiple rule instances for each repair rule according to the repair thought chain of the representative error data;

[0111] A rule repair module, configured to apply all rule instances to generate multiple repair candidate results for each error data, and vote to select the final result for each error data;

[0112] A large language model repair module, configured to, for the error data not repaired by the rules, use the large language model to obtain a repair result and a repair thought chain according to its retrieval result.

[0113] For the apparatus embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial descriptions of the method embodiments. The apparatus embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0114] Correspondingly, this application also provides an electronic device, including: one or more sensors; one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect

[0115] Correspondingly, this application also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the missing table data fairness prediction method as described above is implemented.

[0116] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily think of other implementation schemes of this application. This application is intended to cover any variations, uses, or adaptive changes of this application, and these variations, uses, or adaptive changes follow the general principles of this application and include the common general knowledge or conventional technical means in the technical field not disclosed in this application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of this application are pointed out by the claims.

[0117] It should be understood that the present application is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for interpretable repair of erroneous table data, characterized in that: include: S1: Obtain the error table data information matrix, the table data error mask matrix and the error table data semantic embedding matrix; S2: Retrieving a number of tuples with the strongest association with each error data according to the error table data information matrix, the table data error mask matrix and the error table data semantic embedding matrix; S3: In the error table data information matrix, a number of representative error tuples are sampled for repair; S4: according to the repair and retrieval results of the representative error tuples, a repair thought chain of each representative error tuple is generated using a large language model; S5: generating multiple rule instances for each repair rule using a large language model according to the repair thought chain of the representative error tuple; S6: Apply all rule instances, generate multiple repair candidate results for each erroneous data, and vote for each erroneous data to select the final result; S7: For the erroneous data that is not repaired by the rules, the large language model is used to obtain the repair results and the repair thinking chain according to its retrieval results; S5 specifically includes: S51: for each designated repair rule, using a prompt word paraphrase tool, based on some representative tuples and their repair thought chains, generate multiple different versions of repair rule generation prompt words; S52: Generate prompt words using the above repair rules, and call the large language model to generate specific instances of the repair rules; S53: Verify each rule on all representative tuples and all correct tuples, and only retain the rule instances that pass the verification.

2. The method according to claim 1, characterized in that S1 specifically includes: S11: Acquire table data information and table data error information collected by the sensor; S12: constructing an error table data information matrix and a table data error mask matrix according to the table data information and the table data error information; S13: semantically embedding each element in the error table data information matrix through a semantic embedding method, thereby generating an error table data semantic embedding matrix.

3. The method according to claim 1, characterized in that S2 specifically includes: S21: for each column to be repaired, calculate the normalized mutual information value between it and the remaining columns, and regard the columns whose normalized mutual information values ​​are greater than or equal to the specified threshold as relevant columns; S22: Based on the weighted cosine similarity in the relevant column, measure the similarity between the remaining tuples and the tuple to be repaired, and preliminarily sort the tuples according to the similarity, and retrieve the top several tuples; S23: The preliminarily sorted tuples are further sorted according to data quality to ultimately determine the tuples with the highest ranking, that is, the tuples with the strongest correlation with the erroneous data.

4. The method according to claim 1, characterized in that S3 specifically includes: S31: using clustering algorithm to group tuples, the user specifies the number of clusters, and divides all tuples into several clusters; S32: adopt a greedy algorithm to select a tuple from each cluster as a representative error tuple, so that the selected representative error tuple can cover as many error columns as possible and contain more error data; S33: The selected representative error tuple is handed over to the user for manual repair to obtain the corresponding correct result.

5. The method according to claim 1, characterized in that S4 specifically includes: S41: Based on the representative error tuples repaired by the user, a repair thought chain is constructed for each error column to generate prompt words; S42: Generate prompt words using the above-mentioned repair thought chain, and call the large language model to obtain the repair thought chain for each error data.

6. The method according to claim 1, characterized in that S6 specifically includes: S61: For each column to be repaired, firstly execute all verified rule instances in parallel to generate multiple interpretable repair results for each erroneous data; S62: Among the multiple interpretable repair results, select a repair result supported by most rules as the final result of each erroneous data.

7. The method according to claim 1, characterized in that S7 specifically includes: S71: Construct an interpretable data repair prompt word for each error data that is not repaired by the rule based on the representative error tuple and its repair thinking chain; S72: Use the above-mentioned explainable data repair prompt words and call the large language model to generate a repair result and a repair thought chain for each erroneous data.

8. An interpretable repair device for erroneous table data, characterized in that: include: An acquisition module, used for acquiring an error table data information matrix, a table data error mask matrix and an error table data semantic embedding matrix; A retrieval module, used to retrieve a number of tuples with the strongest association with each error data according to the error table data information matrix, the table data error mask matrix and the error table data semantic embedding matrix; A sampling and repairing module, used for sampling a number of representative error tuples in the error table data information matrix for repair; A repair thought chain generation module, used to generate a repair thought chain for each representative error tuple using a large language model according to the repair result and retrieval result of the representative error tuple; A repair rule generation module, used for generating multiple rule instances for each repair rule using a large language model according to the repair thought chain of the representative error tuple; The rule repair module is used to apply all rule instances, generate multiple repair candidate results for each erroneous data, and vote for each erroneous data to select the final result; The large language model repair module is used to obtain the repair results and repair thought chains using the large language model for the erroneous data that has not been repaired by the rules according to its retrieval results; According to the repair thought chain of the representative error tuple, a large language model is used to generate multiple rule instances for each repair rule, specifically including: For each specified repair rule, a prompt word paraphrase tool is used to generate multiple different versions of repair rule generation prompt words based on some representative tuples and their repair thought chains; Use the above repair rules to generate prompt words, and call the large language model to generate specific instances of the repair rules; Each rule is verified on all representative tuples and all correct tuples, and only the rule instances that pass the verification are retained.

9. An electronic device, characterized in that: include: one or more sensors; one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Downstream analysis feedback-oriented data cleaning method and system and electronic equipment

    CN119149903A

  • Source code patch generation with retrieval-augmented transformer

    WO2024081075A1