Complex table data error detection method and device and electronic equipment

Through the grouping module and joint characterization model based on the distance correlation coefficient, error detection is performed on the table data, and combined with special cases to sample and inverse tendency score correction classifiers, the problems of detection accuracy and inefficiency in the prior art are solved, and efficient and accurate error detection is achieved.

CN119988082AActive Publication Date: 2025-05-13ZHEJIANG UNIV BINJIANG RES INST +1

Patent Information

Application Number
CN202510467265.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing table data error detection algorithm has low detection accuracy, low time efficiency, and a bias in selecting sampling data, resulting in insufficient recognition capabilities for low-frequency category samples.

Method used

The grouping module based on the distance correlation coefficient is used to logically group the attribute columns, combine the sentence-transformer and wordvec models for joint characterization, use the intra-cluster square sum function and the bce-rerank model for special samples, and eliminate the selection bias in the sampling process through the inverse tendency score correction classifier.

Benefits of technology

With a small number of labeled examples, the error detection accuracy is significantly improved, the detection accuracy is improved by more than 27%, and the detection efficiency is improved by more than 13%, overcoming the problems of accuracy and inefficiency in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988082A_ABST
    Figure CN119988082A_ABST
Patent Text Reader

Abstract

The invention discloses a complex table data error detection method. The method comprises the following steps: acquiring a complex table data sample; constructing a logic grouping module, a joint representation module, a special pair sampling module and an inverse tendency score correction classifier module; the logic grouping module performs logic grouping on the attribute columns based on the distance correlation coefficient, and divides the attribute columns most likely to have the context semantic relationship into one group; the joint characterization module strengthens the characterization capability of the context logic relation of the feature vectors in each group; a special pair sampling module based on an intra-cluster quadratic sum function and a bce-rank model can accurately position a special case data pair with the most information amount under the condition of a small number of labeled examples; the classifier module based on inverse tendency score correction aims to eliminate the selection deviation problem in the sampling process; and inputting the table data annotation sample into the error detection model for processing to obtain a final error detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of complex table data detection, and more particularly to a complex table data error detection method and device, and electronic equipment. Background Art

[0002] With the rapid development of the Internet, massive amounts of data continue to emerge and accumulate. By analyzing massive amounts of data, we can obtain potential data value. Data cleaning is the most basic but most challenging step in various relational data analysis tasks (such as entity matching, data integration or analysis). This is because data errors will seriously affect data quality, mislead downstream decision-making analysis, and even cause huge economic losses to enterprises. Relevant studies have shown that in the real world, about 5% of the data units in a data set have data errors due to various reasons.

[0003] Data error detection usually includes two stages: error detection (ED) to identify erroneous values ​​and error correction (EC) to fix erroneous values. Currently, error detection methods can be mainly divided into the following categories: (1) outlier detection, which detects erroneous values ​​that obviously deviate from the overall distribution of data based on statistical methods; (2) duplicate value detection, which removes duplicate values ​​by identifying tuples pointing to the same entity; (3) format error detection, which detects format errors through pre-defined patterns (such as the date format 2000 / 01 / 01); (4) rule violation detection, which detects data errors that violate the rules through given data rules; (5) external knowledge violation detection, which uses external knowledge (such as master data, knowledge base) to detect data errors.

[0004] There are a large number of error detection algorithms in the existing technology. However, in order to achieve a higher prediction score when sampling a small number of labeled instances, the sampling strategy in the existing error detection algorithms usually prioritizes data with a large amount of information and a high frequency of occurrence, while the sampling probability of data with a small amount of information and a low frequency of occurrence is very low. This leads to an uneven distribution of error patterns in the sampled data, and the number of sampled data under different error patterns varies greatly, that is, there is a selection bias problem in the sampled data. The selection bias problem of sampled data may cause the binary classification model to overfit high-frequency category samples during training, while the recognition ability of low-frequency category samples is insufficient. This bias affects the generalization ability of the model. Summary of the invention

[0005] In view of the deficiencies in the prior art, the present invention aims to provide a complex table data error detection method and device, and electronic equipment to improve the accuracy and efficiency of table data error detection.

[0006] To achieve the above object, the present invention provides the following technical solution: a complex table data error detection method, comprising the following steps: Step 1: Obtain tabular data samples and construct tabular data matrix and tabular data label matrix; Step 2: Build a complex table data error detection model, which includes: A grouping module, which logically groups attribute columns based on distance correlation coefficients, and groups attribute columns that are most likely to have contextual semantic relationships into one group; The joint representation module uses sentence-transformer and wordvec to represent the data in the logical group, thereby enhancing the representation ability of the contextual logical relationship of the feature vectors in each group; Exception pair sampling module: This module uses the intra-cluster square sum function and the BCE-rerank model to accurately locate the exception data pairs in the same cluster, thereby screening out the most informative exception data pairs; The classifier module adopts the inverse propensity score correction classifier to eliminate the influence of the selection bias problem in the sampling process on the classifier; Step three, input the tabular data samples into the trained complex tabular data error detection model, and obtain the error detection results after analysis and processing by the grouping module, joint representation module, special case sampling module and classifier module.

[0007] As a further improvement of the present invention, the specific steps of obtaining the table data sample in step 1 and constructing the table data matrix are as follows: Step 11, obtaining a standardized table data sample from a financial website through an automated crawler tool or a database query interface; Step 1 and 2: Construct a table data matrix with n data and m attribute columns based on multi-view missing face image data samples , No. Expressed as At the same time, a tabular data label matrix with k data and m attribute columns is constructed based on the multi-view missing face image label samples { }, Y Expressed as ; In the process of constructing the tabular data matrix, the error in the tabular data is defined as any cell value D[i, j] that is different from its corresponding ground truth value D*[i, j], and the cell value is assigned to the positive class when D[i, j] = D*[i, j], or assigned to the negative class when D[i, j] ̸= D*[i, j], then the cell value is considered to be correctly classified; if the cell value is assigned to the negative class when D[i, j] = D*[i, j], or assigned to the positive class when D[i, j] ̸= D*[i, j], then the cell value is considered to be misclassified.

[0008] As a further improvement of the present invention, the analysis and processing steps of the grouping module in the complex table data error detection model in step 2 are as follows: Step 21: For two groups of samples and , calculate the Eubian distance matrix between them respectively: ; Step 22, then double-center the distance matrix, that is, subtract the average value of its row and column from each element, and then add back the average value of the entire matrix: and, , , Respectively The row, column and overall averages of Step 23: Based on the dual-centered distance matrix obtained in step 22, further calculate the distance covariance and distance variance: ; Step 24: Calculate the distance correlation coefficient based on the distance covariance and distance variance , the distance correlation coefficient , the value range is [0,1], 0 means complete independence, 1 means there is a deterministic relationship: Step 25: Using the dCor metric, the DS strategy is used for each dirty column. Select columns ,in Exceeding the threshold , forming a related column set ,Right now: .

[0009] As a further improvement of the present invention, the analysis and processing steps of the joint characterization module in step 2 are as follows: First, Sentence-Transformer is used to obtain the sentence-level representation of the data in each logical group, and then WordVec is used to extract word-level features. The two are concatenated or weighted summed to form a richer feature vector. Based on the fused features, the self-attention mechanism is applied to assign different weights to different parts according to the context to further optimize the feature representation.

[0010] As a further improvement of the present invention, the analysis and processing steps of the sampling module in the special case of step 2 are as follows: First, the grouped data is converted into multiple sets of embedding vectors using the joint representation module, and then the embedded vectors are clustered using the DBSCAN algorithm to obtain If there are k labeled instances in a cluster, k / 2 samplings are performed. The sampling strategy is divided into two stages: the first stage is group sampling, and the second stage is special pair sampling, in order to screen out the most informative special case data pairs.

[0011] As a further improvement of the present invention, the analysis and processing steps of the classifier module in step 2 are as follows: The concept of IPS score is introduced into the classifier module. By calculating the propensity score and constructing the weight matrix, the weight matrix is ​​then applied to the standard cross entropy loss function of the binary classifier to obtain a binary classifier corrected based on the IPS score. The binary classifier corrected based on the IPS score is then used to eliminate the influence of the selection bias problem in the sampling process on the classifier.

[0012] As a further improvement of the present invention, in step 3, before the table data sample is input into the trained complex table data error detection model, the table data matrix and the table data label matrix are first spliced ​​to obtain the table data splicing matrix C, which is as follows: C ={ } in, Indicates the value of the mth attribute column of the ith data. Indicates the value of the annotation label corresponding to the mth attribute column of the i-th data; Then the table data concatenation matrix C is input into the complex table data error detection model for training to obtain a trained table data error detection model.

[0013] Another aspect of the present invention provides a complex table data error detection device, which is used to execute the complex table data error detection method, comprising: An acquisition module (1) is used to acquire tabular data samples and construct a tabular data matrix and a tabular data label matrix; A modeling module (2) for constructing a complex table data error detection model, the complex table data error detection model comprising a grouping module based on a distance correlation coefficient, a joint representation module based on a WordVec model and a Sentence-transformer model, a special case pair sampling module based on a cluster sum of squares function and a bce-rerank model, and a classifier module based on inverse propensity score correction; A splicing module (3) is used to splice the table data matrix and the table data label matrix to obtain a table data splicing matrix; A training module (4) is used to input the table data concatenation matrix into the complex table data error detection model for training, thereby obtaining a trained table data error detection model; The error detection module (5) is used to input the table data matrix into the trained complex table data error detection model for error detection to obtain the final error detection result.

[0014] Another aspect of the present invention provides an electronic device, comprising: one or more processors; a memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to: execute the above-mentioned complex table data error detection method.

[0015] Another aspect of the present invention provides a computer-readable storage medium, wherein the computer storage medium is used to store computer instructions, and when the computer-readable storage medium is run on a computer, the computer is enabled to execute the complex table data error detection method.

[0016] The beneficial effects of the present invention are as follows: compared with the prior art, the present method uses a distance correlation coefficient to logically group attribute columns, and divides the attribute columns that are most likely to have contextual semantic relationships into a group; the sentence-transformer model and the wordvec model are jointly used to represent the data after logical grouping, thereby enhancing the representation ability of the contextual semantics of the feature vectors in each group, so that a high error detection accuracy can be achieved in the case of a small number of labeled instances; the intra-cluster square sum function and the bce-rerank model are used to accurately locate special case data pairs, thereby accurately screening out the most informative special case data; the inverse tendency score correction classifier is used to eliminate the selection bias problem in the sampling process; the present method overcomes the problems of low detection accuracy and low time efficiency of the existing table data error detection model in the case of a small number of labeled instances, and the present method improves the detection accuracy by more than 27% and the detection efficiency by more than 13% compared with the most advanced complex table data error detection method. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flow chart of the complex table data error detection method of the present invention; Figure 2 A schematic diagram of a complex tabular data error detection model; Figure 3 The block diagram of the complex table data error detection device. DETAILED DESCRIPTION

[0018] The present invention will be further described below in detail with reference to the embodiments shown in the accompanying drawings.

[0019] Among the five error types mentioned in the background technology, the most complex one is the violation of data context semantic relationship. The other four error types are easy to find, but the data context semantic relationship is not easy to find. In a tabular data set, there is often a data context semantic relationship between some attribute columns, but no context semantic relationship with other attribute columns. At this time, if they can be divided into different groups, the resulting embedded vector representation can more accurately represent the context semantic relationship.

[0020] Based on the above concept, this embodiment provides a complex table data error detection method, referring to Figure 1 As shown, it mainly includes the following steps: S1: Obtain tabular data samples and construct tabular data matrix and tabular data label matrix; this step includes the following sub-steps: S11: Collect standard table data samples from financial websites through automated crawler tools or database query interfaces.

[0021] S12: Construct a tabular data matrix.

[0022] Construct a table data matrix of n data and m attribute columns according to the multi-view missing face image data samples . No. It can be expressed as ; Construct a tabular data label matrix of k data and m attribute columns according to the multi-view missing face image label samples { }. It can be expressed as .

[0023] In this example, an error is defined as any cell value D[i, j] that is different from its corresponding ground truth value D*[i, j]. The goal is to determine the most likely positive or negative class assignment for each cell value D[i, j] in a given dataset D. A cell value is considered correctly classified if it is assigned to the positive class when D[i, j]= D*[i, j] or to the negative class when D[i, j]̸= D*[i, j]; a cell value is considered misclassified if it is assigned to the negative class when D[i, j]= D*[i, j] or to the positive class when D[i, j]̸= D*[i, j].

[0024] S2: Construct a complex table data error detection model, including a grouping module based on distance correlation coefficient, a joint representation module based on WordVec model and Sentence-transformer model, a special case pair sampling module based on the intra-cluster square sum function and bce-rerank model, and a classifier module based on inverse propensity score correction; refer to Figure 2 , building the model includes the following modules: Grouping module: Logically group attribute columns based on distance correlation coefficients, and group attribute columns that are most likely to have contextual semantic relationships into one group; Joint representation module: Use sentence-transformer and wordvec to represent the data in the logical group, thereby enhancing the representation ability of the contextual logical relationship of the feature vectors in each group; Special case pair sampling module: It uses the intra-cluster square sum function and the bce-rerank model to accurately locate the special case data pairs in the same cluster, thereby accurately screening out the most informative special case data pairs; Classifier module: The inverse propensity score correction classifier is used to eliminate the impact of selection bias in the sampling process on the classifier.

[0025] Next, the above four modules are further explained: S21: The grouping module logically groups the attribute columns based on the distance correlation coefficient. The grouping module groups the attribute columns that are most likely to have contextual semantic relationships into one group.

[0026] Specifically, this embodiment introduces a column selection strategy DS based on distance correlation coefficient, which uses the distance correlation coefficient (DCOR) to evaluate the correlation between columns and filter out irrelevant columns to facilitate data retrieval. Specifically, the DS strategy uses the DCOR metric to evaluate the correlation between columns. Relative to the target dirty column m, the column n with a higher DCOR value often contains more actual correlation information with the m column, which is crucial for effective ED.

[0027] The first step of the grouping module is to construct a distance matrix: First, for the two groups of samples and , calculate the Eubian distance matrix between them respectively: The distance matrix is ​​then double-centered, which means that the average value of its row and column is subtracted from each element, and then the average value of the entire matrix is ​​added back. This process is intended to eliminate the offset effect of the distance matrix, and the formula is as follows: and, , , Respectively The row, column and overall averages of .

[0028] Based on the dual-centered distance matrix, the distance covariance and distance variance are further calculated. Distance covariance is a key indicator to measure the distance similarity between two vectors, while distance variance is the quantification of the distance change within a single vector. The calculation formula is as follows: Finally, the distance correlation coefficient is calculated based on the distance covariance and distance variance , which is a standardized indicator to measure the dependence of two random vectors, with a value range of [0,1], where 0 indicates complete independence and 1 indicates a deterministic relationship: Using the dCor metric, the DS strategy for each dirty column Select columns ,in Exceeding the threshold , forming a related column set ,Right now: In this embodiment, this embodiment will It is set to 0.5 to balance the coverage of relevant columns with the need for high information content, ensuring that key features are retained while minimizing information loss. This strategy ensures that the most relevant columns with higher dCor values ​​are selected to construct each grouping of the data, so that the data columns within each group are closely related, and the embedding vector model built on this basis can contain contextual semantic information.

[0029] S22: The joint representation module uses sentence-transformer and wordvec to represent the data in the logical group, thereby enhancing the representation ability of the contextual logical relationship of the feature vectors in each group; Specifically, Sentence-Transformer is first used to obtain the sentence-level representation of the data in each logical group, and then WordVec is used to extract word-level features. The two are concatenated or weighted summed to form a richer feature vector. Based on the fused features, the self-attention mechanism is applied to assign different weights to different parts according to the context to further optimize the feature representation. Through the above method, not only can local semantic information be captured, but also the global contextual logical relationship can be better understood, ultimately improving the quality and expression ability of the feature vectors in each group.

[0030] The joint representation module first processes the tabular data: Contains m rows and n columns, where the first k columns are text columns and the last nk columns are numeric columns. The specific representation is as follows: in Represents the data in the i-th row and j-th column.

[0031] Next, use SentenceTransformer to embed and extract the text column from the tabular data: Then use SentenceTransformer to encode the text column: in is the output dimension of SentenceTransformer. Then use WordVec to encode the text column: in is the output dimension of WordVec. Concatenate the embedding results of SentenceTransformer and WordVec: Extract numeric columns: Concatenate the joint embedding of the text column with the features of the numeric column: Define the query, key, and value matrices: Calculate the attention score: in is the size of the key dimension, usually = + + . Apply the Softmax function to get the attention weight distribution: Finally, the weighted sum is calculated to get the output: S23: The special case pair sampling module uses the intra-cluster sum of squares function and the bce-rerank model to accurately locate the special case data pairs in the highest scoring cluster, thereby accurately screening out the most informative tabular data.

[0032] Specifically, the grouped data is first converted into multiple sets of embedding vectors using the joint representation module, and then the embedded vectors are clustered using the DBSCAN algorithm to obtain If there are k labeled instances, k / 2 samplings are performed. The sampling strategy is divided into two stages: the first stage is group sampling, and the second stage is special pair sampling.

[0033] First stage sampling: The first stage of sampling is The cluster with the highest score is found among the clusters based on the cluster score function. The cluster score function is as follows: in: Indicates the number of data items that have been labeled in the cluster. Represents the Within-Cluster Sum of Squares, which is used to measure the compactness or variability of data points within a cluster.

[0034] The cluster score function evaluates the quality of each cluster by combining the above three variables. In order to evaluate the homogeneity of the data set, this embodiment introduces the concept of within-cluster sum of squares (WCSS) into the cluster score function. It is an effective measure of the variability or compactness of a data set and represents the sum of the squared differences between each data point and the center point of the data set. Specifically, The calculation formula is: in, represents the number of data points in the cluster, represents the number of features, represents the jth eigenvalue of the ith data point, Represents the jth eigenvalue of the cluster center.

[0035] The cluster score function takes into account the size of the cluster, the number of labeled data points, and the compactness of the data points within the cluster to assign a score to each cluster. The higher the score, the more representative the cluster is and the more suitable it is to be selected as the target cluster for sampling.

[0036] The second stage of sampling: find the most relevant data and outlier data based on the bce-rerank model. Using the bce-rerank model, find the data that is most relevant to all the data in the cluster and the outlier data in the cluster with the highest score. First, calculate the similarity. For each data point , use the bce-rerank model to calculate the similarity scores between data The specific formula is: Then re-sort the data: sort the data from high to low by similarity score. The data with the highest similarity score is the most relevant data. The data with the lowest similarity score is the outlier data.

[0037] Form special case pairs: The most relevant data and the outlier data are combined into special case pairs, and a sampling process is completed to obtain a special case pair containing two sampled data.

[0038] Using the bce-rerank model can significantly improve the retrieval effect. Through the above two-stage sampling strategy, the most informative table data can be efficiently screened, thereby improving the accuracy of error detection.

[0039] S23: The classifier module uses an inverse propensity score correction classifier to eliminate the impact of selection bias in the sampling process on the classifier.

[0040] Specifically, this embodiment proposes to introduce the concept of IPS score into the classifier module. By calculating the propensity score, constructing the weight matrix, and then applying the weight matrix to the standard cross entropy loss function of the binary classifier, a binary classifier corrected based on the IPS score is obtained. The corrected binary classifier can reduce excessive reliance on popular cluster data while increasing learning on unpopular cluster data, thereby achieving more balanced and comprehensive error detection. In order to prove this more rigorously, this embodiment designs an expected loss function. Assume that the true label of each sample and predicted probability are independent and identically distributed, and the observation probability of each sample is .

[0041] Under ideal balanced sampling conditions, the expectation of the standard cross entropy loss function for a binary classifier is: because and is independent and identically distributed, the expectation can be split into: In the case of selection bias in the sampled data, the probability of observing each sample is is not uniformly distributed. Assume that the sample The observation probability is , then the expected loss function of all data is: However, due to the observation probability is not uniform, and the number of samples actually observed may be biased. Therefore, the expected loss function of all data can be expressed as: In order to prove that the expected sum of all loss functions in the biased case is not equal to the expected sum of loss functions in the balanced case, this example compares and : Obviously, unless For all ,otherwise In order to eliminate the deviation, this embodiment uses the IPS (Importance Sampling) weight: The expected loss function of all data after IPS weighting is: After simplification, we get: This is exactly what you would expect from a standard cross entropy loss function. .

[0042] Through the above analysis, this embodiment can draw the following conclusions: In the case of selection bias, the expected loss function of all data is Compared with the standard cross entropy loss function Different. The expected loss function of all data weighted by IPS score is After transformation, it is consistent with the standard cross entropy loss function same.

[0043] In other words, by introducing appropriate weights (such as IPS weights), the influence of selection bias can be eliminated, making the weighted expected loss function consistent with the standard expected loss function, solving the sampling bias problem caused by the pursuit of high-information labeled instances in the error detection algorithm, and enabling the weighted corrected binary classifier to learn each error pattern in a balanced manner.

[0044] S3: concatenate the table data matrix and the table data label matrix to obtain a table data concatenation matrix C; Specifically, the table data matrix X and the table data label matrix Y are concatenated, which can be expressed as the table data concatenation matrix C: C ={ } in, Indicates the value of the mth attribute column of the ith data. Indicates the value of the annotation label corresponding to the mth attribute column of the i-th data.

[0045] S4: inputting the table data concatenation matrix C into the complex table data error detection model for training to obtain a trained table data error detection model; Specifically, the tabular data concatenation matrix C is input into the complex tabular data error detection model. Through the sampling strategy based on the cosine function and the decision tree model, the specified number of labeled tabular data items are sampled multiple times. Through multiple sampling training, the model can be better generalized in practical applications and the overfitting of specific training samples can be reduced, thereby enhancing the robustness of the training model.

[0046] S5: input the table data concatenation matrix C into the trained complex table data error detection model to perform error detection and obtain the final error detection result; Corresponding to the aforementioned embodiment of the complex table data error detection method, the present application also provides an embodiment of the complex table data error detection device.

[0047] Based on the complex table data error detection method described above, this embodiment further specifically provides a corresponding execution device, referring to Figure 3 As shown, the device comprises: Acquisition module 1, used to acquire tabular data samples and construct tabular data matrix and tabular data label matrix; Modeling module 2 is used to build a complex table data error detection model, including a grouping module based on distance correlation coefficient, a joint representation module based on WordVec model and Sentence-transformer model, a special case pair sampling module based on the intra-cluster square sum function and bce-rerank model, and a classifier module based on inverse propensity score correction; A splicing module 3 is used to splice the table data matrix and the table data label matrix to obtain a table data splicing matrix; A training module 4 is used to input the table data splicing matrix into the complex table data error detection model for training to obtain a trained table data error detection model; The error detection module 5 is used to input the table data matrix into the trained complex table data error detection model for error detection to obtain the final error detection result.

[0048] The specific manner in which the above-mentioned modules perform operations has been described in detail in the above-mentioned embodiments of the method, and will not be described in detail here. At the same time, based on the device, corresponding electronic equipment and computer storage media are provided to cooperate with the execution of the above-mentioned method, so they will not be repeated.

[0049] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A complex table data error detection method, characterized by: The steps include: Step 1: Obtain tabular data samples and construct tabular data matrix and tabular data label matrix; Step 2: Build a complex table data error detection model, which includes: A grouping module, which logically groups attribute columns based on distance correlation coefficients, and groups attribute columns that are most likely to have contextual semantic relationships into one group; The joint representation module uses sentence-transformer and wordvec to represent the data in the logical group, thereby enhancing the representation ability of the contextual logical relationship of the feature vectors in each group; Exception pair sampling module: This module uses the intra-cluster square sum function and the BCE-rerank model to accurately locate the exception data pairs in the same cluster, thereby screening out the most informative exception data pairs; The classifier module adopts the inverse propensity score correction classifier to eliminate the influence of the selection bias problem in the sampling process on the classifier; Step three, input the tabular data samples into the trained complex tabular data error detection model, and obtain the error detection results after analysis and processing by the grouping module, joint representation module, special case sampling module and classifier module.

2. The complex table data error detection method according to claim 1, characterized in that: The specific steps of obtaining the table data sample in step 1 and constructing the table data matrix are as follows: Step 11, obtaining a standardized table data sample from a financial website through an automated crawler tool or a database query interface; Step 1 and 2: Construct a table data matrix with n data and m attribute columns based on multi-view missing face image data samples , No. Expressed as At the same time, a tabular data label matrix with k data and m attribute columns is constructed based on the multi-view missing face image label samples { }, Y Expressed as ; In the process of constructing the tabular data matrix, the error in the tabular data is defined as any cell value D[i, j] that is different from its corresponding ground truth value D*[i, j], and the cell value is considered to be correctly classified when D[i, j] = D*[i, j], or assigned to the positive class when D[i, j] ̸= D*[i, j]; if the cell value is assigned to the negative class when D[i, j] = D*[i, j], or assigned to the positive class when D[i, j] ̸= D*[i, j], then the cell value is considered to be misclassified.

3. The complex table data error detection method according to claim 1 or 2, characterized in that: The analysis and processing steps of the grouping module in the complex table data error detection model in step 2 are as follows: Step 21: For two groups of samples and , calculate the Eubian distance matrix between them respectively: ; Step 22, then double-center the distance matrix, that is, subtract the average value of its row and column from each element, and then add back the average value of the entire matrix: ; and, , , Respectively The row, column and overall average of Step 23: Based on the dual-centered distance matrix obtained in step 22, further calculate the distance covariance and distance variance: ; ; Step 24: Calculate the distance correlation coefficient based on the distance covariance and distance variance , the distance correlation coefficient , the value range is [0,1], 0 means complete independence, 1 means there is a deterministic relationship: ; Step 25: Using the dCor metric, the DS strategy is used for each dirty column. Select columns ,in Exceeding the threshold , forming a related column set ,Right now: 。 4. The complex table data error detection method according to claim 1 or 2, characterized in that: The analysis and processing steps of the joint characterization module in step 2 are as follows: First, Sentence-Transformer is used to obtain the sentence-level representation of the data in each logical group, and then WordVec is used to extract word-level features. The two are concatenated or weighted summed to form a richer feature vector. Based on the fused features, the self-attention mechanism is applied to assign different weights to different parts according to the context to further optimize the feature representation.

5. The complex table data error detection method according to claim 1 or 2, characterized in that: The analysis and processing steps for the sampling module in the special case of step 2 are as follows: First, the grouped data is converted into multiple sets of embedding vectors using the joint representation module, and then the embedded vectors are clustered using the DBSCAN algorithm to obtain If there are k labeled instances in a cluster, k / 2 samplings are performed. The sampling strategy is divided into two stages: the first stage is group sampling, and the second stage is special pair sampling, in order to screen out the most informative special case data pairs.

6. The complex table data error detection method according to claim 1 or 2, characterized in that: The analysis and processing steps of the classifier module in step 2 are as follows: The concept of IPS score is introduced into the classifier module. By calculating the propensity score and constructing the weight matrix, the weight matrix is ​​then applied to the standard cross entropy loss function of the binary classifier to obtain a binary classifier corrected based on the IPS score. The binary classifier corrected based on the IPS score is then used to eliminate the influence of the selection bias problem in the sampling process on the classifier.

7. The complex table data error detection method according to claim 1 or 2, characterized in that: In the step 3, before the table data sample is input into the trained complex table data error detection model, the table data matrix and the table data label matrix are first concatenated to obtain the table data concatenation matrix C, which is as follows: C ={ }; in, Indicates the value of the mth attribute column of the ith data. Indicates the value of the annotation label corresponding to the mth attribute column of the i-th data; Then the table data concatenation matrix C is input into the complex table data error detection model for training to obtain a trained table data error detection model.

8. A complex table data error detection device, used to execute the complex table data error detection method according to any one of claims 1 to 7, characterized in that: include: An acquisition module (1) is used to acquire tabular data samples and construct a tabular data matrix and a tabular data label matrix; A modeling module (2) for constructing a complex table data error detection model, the complex table data error detection model comprising a grouping module based on a distance correlation coefficient, a joint representation module based on a WordVec model and a Sentence-transformer model, a special case pair sampling module based on a cluster sum of squares function and a bce-rerank model, and a classifier module based on inverse propensity score correction; A splicing module (3) is used to splice the table data matrix and the table data label matrix to obtain a table data splicing matrix; A training module (4) is used to input the table data concatenation matrix into the complex table data error detection model for training, thereby obtaining a trained table data error detection model; The error detection module (5) is used to input the table data matrix into the trained complex table data error detection model for error detection to obtain the final error detection result.

9. An electronic device, characterized in that: include: one or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to: execute the complex table data error detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer storage medium is used to store computer instructions, and when the computer storage medium is run on a computer, the computer is enabled to execute the complex table data error detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • A Chinese table column label recovery method and system based on text classification

    CN109710725A

  • Intelligent session method and server based on table data retrieval

    CN115495563A

  • Intelligent session method and server based on table data retrieval.

    MX2023003764A

  • Feature transformation and missing values

    US10534799B1

  • NLP and AIS of I / O, prompts, and collaborations of data, content, and correlations for evaluating, predicting, and ascertaining metrics for IP, creations, publishing, and communications ontologies

    US12094018B1

Cited By

  • Fetal health classification method, system and device and storage medium

    CN120260931A