Entity error correction method and apparatus for structured data, device and medium
By combining the prior knowledge of graph embedding models, language models and knowledge graphs, structured data is corrected, and the problem of low error correction accuracy in the existing technology is solved, and more efficient structured data error correction is achieved.
Patent Information
- Application Number
- PCT/CN2023/143445
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2023-12-29
- Publication Date
- 2025-06-26
AI Technical Summary
The prior art is difficult to effectively reason about structured data in combination with prior knowledge, resulting in low error correction accuracy when processing structured data.
By obtaining entity tuples in the structured data to be corrected, using the graph embedding model and language model combined with the prior knowledge of the knowledge graph, the correlation between attribute values is calculated, candidate values are determined and error correction is performed.
It improves the accuracy of structured data error correction, effectively combines prior knowledge to perform data inference, and enhances the reliability of data error correction.
Smart Images

Figure CN2023143445_26062025_PF_FP_ABST
Abstract
Description
Entity error correction method, device, equipment and medium for structured data
[0001] This application is based on the Chinese invention application with application number 202311773308.5 filed on December 20, 2023, and entitled “Entity error correction method, device, equipment and medium for structured data”, and claims its priority. Technical Field
[0002] The present application is applicable to the field of data cleaning technology, and in particular relates to a method, apparatus, device and medium for entity error correction of structured data. Background Art
[0003] In terms of data quality, error repair for structured relational data is particularly important. It focuses on identifying and repairing errors in cells in relational table data, including null value filling, spelling errors, inconsistency errors, and timing errors. Error repair is divided into error detection and data repair. For relational table data, the first step is to identify cells in the data that are suspected to have problems. In practice, cells with other problems, besides null values, are difficult to detect. For example, inconsistency errors rely on the dependencies between data, while spelling errors rely on data roles and common sense. Then, once the erroneous cell is detected, its error needs to be repaired, i.e., the correct value needs to be filled in. A common method is to link it to a strongly associated tuple or other attribute. When there are many candidate proofreading values, a technique is needed to fill in the most appropriate value for repair. For example, during the bank regulatory reporting process, the bank's transaction flow needs to be checked and proofread, and the bank needs to provide detailed detailed table data and indicator data. Due to the large amount of data and the existence of accuracy, manual filling, system failures and other reasons, it is difficult to make the data completely correct. Therefore, the data needs to be cleaned through data quality detection and repair.
[0004] Currently, the main methods for data detection and repair are: 1) rule-based methods, which use rule mining to mine the rules and patterns of the data itself and use the generated rules to discover inconsistencies in the data. This method often does not require manual injection of labeled data, but relies solely on the data itself for error detection and repair. However, the complexity of mining rules is usually high; 2) statistical learning-based methods, which often extract various features from the data, such as consistency features based on functional dependencies, pattern features, and outlier features based on anomaly detection. After the features are collected, the method often uses traditional machine learning methods such as random forests and support vector machines (SVM) for training and prediction. After the user performs a small amount of annotation, the model is trained and can be used to predict and correct errors in unseen data; 3) neural network-based methods, which use an end-to-end approach to encapsulate the problem of error detection into a binary classification task, train with a small amount of manual annotation, and predict unseen cells. For error repair, a generative model is used to generate the correct value. However, if the above-mentioned data detection and repair methods are applied to structured data, since the current embedding model is mainly trained through unstructured text, it will be unable to accurately learn the correlation between attributes due to the lack of prior knowledge when processing structured data, which in turn leads to low accuracy when processing structured data.
[0005] Therefore, how to effectively combine prior knowledge to reason about structured data to improve the accuracy of data error correction has become an urgent problem to be solved.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a method, apparatus, device, and medium for entity error correction of structured data to solve the problem of how to effectively combine prior knowledge to reason about structured data to improve the accuracy of data error correction.
[0008] A method for correcting entity errors in structured data, the method comprising:
[0009] Obtain any entity tuple in the structured data to be corrected, and construct the data to be corrected according to a first attribute value of at least one first attribute and a second attribute value of a second attribute in the entity tuple, wherein the first attribute is different from the second attribute;
[0010] Inputting the to-be-corrected data and a preset knowledge graph into a graph embedding model, and outputting a first confidence level representing the correlation between the first attribute value and the second attribute value;
[0011] Inputting the to-be-corrected data into a language model, and outputting a second confidence level representing the correlation between the first attribute value and the second attribute value;
[0012] If it is detected that the first confidence level and the second confidence level indicate that the second attribute value is incorrect, matching a set of candidate values for the second attribute from the preset knowledge graph based on the first attribute value;
[0013] Based on the graph embedding model, the correlation between the first attribute value and each candidate value in the candidate value set is calculated, and the candidate value with the highest correlation is determined. According to the candidate value with the highest correlation, the second attribute value in the entity tuple is corrected to obtain a corrected entity tuple.
[0014] A physical error correction device for structured data, the physical error correction device comprising:
[0015] a module for determining data to be corrected, configured to obtain any entity tuple in the structured data to be corrected, and construct the data to be corrected based on a first attribute value of at least one first attribute and a second attribute value of a second attribute in the entity tuple, wherein the first attribute is different from the second attribute;
[0016] A graph embedding model analysis module, configured to input the to-be-corrected data and a preset knowledge graph into a graph embedding model, and output a first confidence level representing the correlation between the first attribute value and the second attribute value;
[0017] a language model analysis module, configured to input the to-be-corrected data into a language model and output a second confidence level representing a correlation between the first attribute value and the second attribute value;
[0018] a candidate value determination module, configured to match a set of candidate values for the second attribute from the preset knowledge graph based on the first attribute value if it is detected that the first confidence level and the second confidence level indicate that the second attribute value is incorrect;
[0019] An attribute value error correction module is used to calculate the correlation between the first attribute value and each candidate value in the candidate value set based on the graph embedding model, determine the candidate value with the highest correlation, and correct the second attribute value in the entity tuple based on the candidate value with the highest correlation to obtain a corrected entity tuple.
[0020] A computer device includes a memory, a processor, and a readable storage medium stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the readable storage medium:
[0021] Obtain any entity tuple in the structured data to be corrected, and construct the data to be corrected according to a first attribute value of at least one first attribute and a second attribute value of a second attribute in the entity tuple, wherein the first attribute is different from the second attribute;
[0022] Inputting the to-be-corrected data and a preset knowledge graph into a graph embedding model, and outputting a first confidence level representing the correlation between the first attribute value and the second attribute value;
[0023] Inputting the to-be-corrected data into a language model, and outputting a second confidence level representing the correlation between the first attribute value and the second attribute value;
[0024] If it is detected that the first confidence level and the second confidence level indicate that the second attribute value is incorrect, matching a set of candidate values for the second attribute from the preset knowledge graph based on the first attribute value;
[0025] Based on the graph embedding model, the correlation between the first attribute value and each candidate value in the candidate value set is calculated, and the candidate value with the highest correlation is determined. According to the candidate value with the highest correlation, the second attribute value in the entity tuple is corrected to obtain a corrected entity tuple.
[0026] One or more computer-readable storage media storing computer-readable instructions, wherein the computer-readable instructions, when executed by one or more processors, cause the one or more processors to perform the following steps:
[0027] Obtain any entity tuple in the structured data to be corrected, and construct the data to be corrected according to a first attribute value of at least one first attribute and a second attribute value of a second attribute in the entity tuple, wherein the first attribute is different from the second attribute;
[0028] Inputting the to-be-corrected data and a preset knowledge graph into a graph embedding model, and outputting a first confidence level representing the correlation between the first attribute value and the second attribute value;
[0029] Inputting the to-be-corrected data into a language model, and outputting a second confidence level representing the correlation between the first attribute value and the second attribute value;
[0030] If it is detected that the first confidence level and the second confidence level indicate that the second attribute value is incorrect, matching a set of candidate values for the second attribute from the preset knowledge graph based on the first attribute value;
[0031] Based on the graph embedding model, the correlation between the first attribute value and each candidate value in the candidate value set is calculated, and the candidate value with the highest correlation is determined. According to the candidate value with the highest correlation, the second attribute value in the entity tuple is corrected to obtain a corrected entity tuple.
[0032] The present application obtains any entity tuple in the structured data to be corrected, constructs the data to be corrected based on the first attribute value of at least one first attribute and the second attribute value of a second attribute in the entity tuple, inputs the data to be corrected and the preset knowledge graph into the graph embedding model, outputs a first confidence level representing the correlation between the first attribute value and the second attribute value, inputs the data to be corrected into the language model, outputs a second confidence level representing the correlation between the first attribute value and the second attribute value, if it is detected that the first confidence level and the second confidence level represent that the second attribute value is wrong, then according to the first attribute value, a set of candidate values for the second attribute is matched from the preset knowledge graph, based on the graph embedding model, calculates the correlation between the first attribute value and each candidate value in the candidate value set, determines the candidate value with the highest correlation, and corrects the second attribute value in the entity tuple according to the candidate value with the highest correlation to obtain a corrected entity tuple. Among them, combining the prior knowledge of the knowledge graph, the graph embedding model and the language model for error detection, and using the model in the error detection process to correct the erroneous attributes, can effectively combine the prior knowledge to reason about structured data, thereby improving the accuracy of data error correction. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0034] FIG1 is a schematic diagram of an application environment of a physical error correction method for structured data provided in Example 1 of the present application;
[0035] FIG2 is a flow chart of a method for entity error correction of structured data provided in Example 2 of the present application;
[0036] FIG3 is a flow chart of a method for entity error correction of structured data provided in Example 3 of the present application;
[0037] FIG4 is a flow chart of a method for entity error correction of structured data provided in Example 4 of the present application;
[0038] FIG5 is a flow chart of a method for entity error correction of structured data provided in Example 5 of the present application;
[0039] FIG6 is a schematic diagram of a model framework of a physical error correction method provided in Example 5 of the present application;
[0040] FIG7 is a flow chart of a method for entity error correction of structured data provided in Example 6 of the present application;
[0041] FIG8 is a flow chart of a method for entity error correction of structured data provided in Example 7 of the present application;
[0042] FIG9 is a schematic structural diagram of a physical error correction device for structured data provided in Example 8 of the present application;
[0043] FIG10 is a schematic structural diagram of a computer device provided in Example 9 of the present application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0045] In order to illustrate the technical solution of the present application, specific embodiments are provided below.
[0046] A method for correcting the entity error of structured data provided in the first embodiment of the present application can be applied in an application environment such as that shown in FIG1 , wherein a server communicates with a client, and the server is used to carry the method for correcting the entity error of the data to provide a data cleaning service. The client can request the data cleaning service from the server by providing corresponding structured data, thereby obtaining cleaned data. Of course, the server can also clean the structured data stored in itself. The client includes but is not limited to a PDA, a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud computer device, a personal digital assistant (PDA), and other devices. The computer device corresponding to the server can be implemented using an independent server or a server cluster consisting of multiple servers.
[0047] Referring to Figure 2, which is a flow chart of a method for entity error correction of structured data provided in Example 2 of the present application, the entity error correction method for structured data is applied to the server in Figure 1. The user corresponding to the client sends the structured data to be corrected to the server, triggering the server to perform data error detection and repair tasks. As shown in Figure 2, the entity error correction method for structured data may include the following steps:
[0048] Step S201 : obtaining any entity tuple in the structured data to be corrected, and constructing the data to be corrected according to a first attribute value of at least one first attribute and a second attribute value of a second attribute in the entity tuple.
[0049] In this embodiment, the structured data is a row or a column in a relational table in a database. If a row is represented as a tuple, the row of data corresponds to an entity. Each column in the row of data represents the attributes of the entity. The entity tuple is a row of data in the relational table. The structured data includes at least one entity tuple. Specifically, the structured data can be a relational table.
[0050] In the process of collecting and organizing data, the attributes of two different entities may be mistakenly merged into one entity. For example, there are two entities named A. The gender attribute of the first entity A is "male (that is, the attribute value of the gender attribute)", and the age attribute of the second entity A is "25 years old (the attribute value of the age attribute)". After the two entities are mistakenly merged, the entity A is obtained, with the gender attribute of "male" and the age attribute of "25 years old". Obviously, this is an incorrect fusion result and a manifestation of entity anomaly.
[0051] Furthermore, during data collection and organization, the attribute values of an entity may conflict. For example, entity B may have a gender attribute of "male" and a medical treatment attribute of "gynecology." Since men are not advised to seek medical treatment at gynecology clinics, the attributes of entity B may be incorrect, which is a manifestation of an entity anomaly. Of course, other entity anomalies include instances where an attribute is empty or exceeds a limit.
[0052] For structured data, the detection method can be used to filter out abnormal attributes in entity tuples, wherein the detection object can be the attribute value.
[0053] For any entity tuple in the structured data to be corrected, erroneous data may appear. The method of the present application can perform error detection and repair on each entity tuple separately. The objects of detection and correction are specifically one or more attributes in the entity tuple. During the detection and correction process, one or more attributes in the entity tuple should be considered correct, and an attribute other than the correct attribute is taken as the attribute to be corrected. The first attribute recorded in the above steps is the attribute considered correct in this error correction, and the second attribute is the attribute to be corrected. It can be seen that the first attribute is different from the second attribute.
[0054] There may be one first attribute in the above-mentioned data to be corrected. Of course, there may also be multiple first attributes, that is, multiple attributes different from the second attribute. For example, attributes such as name, birthday, school, etc. are multiple first attributes, and the political relationship attribute is the second attribute. In this way, multiple first attributes can be used to detect the second attribute, and correction can be made when the second attribute is an erroneous attribute, thereby achieving more accurate correction of the second attribute.
[0055] Step S202: Input the data to be corrected and the preset knowledge graph into the graph embedding model, and output a first confidence level representing the correlation between the first attribute value and the second attribute value.
[0056] In this embodiment, the preset knowledge graph may refer to graph data containing some known data and the association relationships between data. Based on the association relationships between the known data and data, some prior knowledge about the data to be corrected can be learned. For example, if it is known that "DOK.fest is an annual event in Munich", then t[A] = "Munich" and t[b] = "DOK.fest" are likely to be related, where t is an entity tuple, A is the first attribute (which can contain multiple attributes), and b is the second attribute.
[0057] By pre-training the knowledge graph for use in graph embedding models, graph embedding models can implicitly learn rich contextual information from the pre-trained graph data, for example, learning that DOK.fes was held in Munich. Given a knowledge graph G(V,E,L), a graph representation method is used to learn node / relationship embeddings such that for node embeddings u, v, v′ and relationship embedding r with edge label r, where e = (u,v) and L(e) = r is in E, and e′ = (u,v′) and L(e′) = r is not in E, the score function g(u,r,v) is maximized while g(u,r,v′) is minimized. In other words, embeddings are learned such that if two vertices have close embeddings, they are related. This pre-training can be performed once as a preprocessing step and is not performed during inference, thus not increasing the cost of inference.
[0058] For example, a preset knowledge graph can find corresponding movie matches in the Internet Movie Database (IMDb), and then use the information in IMDb to check and correct structured data such as movie information.
[0059] In this embodiment, the graph embedding model is a model trained according to a pre-set training data set. It can access the graph data after pre-training of the knowledge graph to label and embed the input data, thereby forming a graph embedding result, and combined with the encoder in the model to realize the prediction of correlation.
[0060] In the graph embedding model, by analyzing the pre-trained graph data of the preset knowledge graph, some prior knowledge between the first attribute and the second attribute in the data to be corrected can be obtained, so that the correlation between the first attribute and the second attribute can be predicted. The confidence of the prediction result is the first confidence that characterizes the correlation between the first attribute value of the first attribute and the second attribute value of the second attribute.
[0061] Step S203: input the data to be corrected into the language model, and output a second confidence level representing the correlation between the first attribute value and the second attribute value.
[0062] In this embodiment, the data to be corrected can be expressed as a language, so that the language model can be used to analyze the data to be corrected to predict the correlation between the first attribute and the second attribute. The confidence of the prediction result is the second confidence that represents the correlation between the first attribute value of the first attribute and the second attribute value of the second attribute.
[0063] Among them, the language model can adopt an N-gram language model, a Bert model, etc., or even a large language model (LLM). When the data to be corrected is input into the language model for analysis, the language expression of the data to be corrected can be set according to needs, and the language model can be pre-trained using the corresponding training data set to obtain a pre-trained language model. The language model used in step S203 is the language model pre-trained in this usage scenario.
[0064] During training, the graph embedding model in the above step S202 and the language model in this step S203 can be jointly trained, and the confidence output by the two models can be fused to obtain a final confidence to guide the joint training. For example, let Tc = {(xi,yi)} be the training data set, where xi = (ti[A],ti[b]), xi = (ti[A],ti[b]) is the i-th training data, and yi∈{0,1} is the label of xi; if ti[A] and ti[b] are related, yi = 1, otherwise yi = 0. During joint training, the cross entropy loss function is used. In practice, for t∈D and A∈R, t[RA] is usually related to t[A], where RA is all attributes except A. The above training data set Tc can be generated by self-supervised learning, and the specific generation method is as follows:
[0065] 1. Randomly extract a tuple t from D; 2. Randomly select an attribute b and use the remaining attributes as A; 3. With a certain probability p, use (t[A], t[b], 1) as a positive example, with probability 1-p, randomly select a value c from attribute b, and use (t[A], c, 0) as a negative example, thus obtaining N samples in Tc.
[0066] In step S204, if it is detected that the first confidence level and the second confidence level indicate that the second attribute value is wrong, a candidate value set of the second attribute is matched from a preset knowledge graph based on the first attribute value.
[0067] In this embodiment, by analyzing the above-mentioned first confidence level and second confidence level, it is possible to detect whether the second attribute value is wrong or correct. If the second attribute value is wrong, that is, the corresponding first confidence level and second confidence level are both small, then it means that the second attribute is detected and needs to be corrected. If the second attribute value is correct, that is, the corresponding first confidence level and second confidence level are both large, then it means that the second attribute does not need to be corrected.
[0068] When the second attribute value needs to be corrected, the knowledge graph above is used to provide a suggested value, which is the attribute value used for error correction of the second attribute value. Specifically, based on the first attribute value, the candidate value set CandB is retrieved from the knowledge graph through HER to obtain the candidate value of t[b].
[0069] For example, given a partial tuple t[A], an attribute b and a knowledge graph G = (V, E, L), we first extract the set of vertices Vt that match t[A] from G through the heterogeneous entity resolution method HER; then, for each vertex v in Vt, we check each of its k-hop neighbors v' within a predefined boundary k in G; let ρ be the label path from v to v', sim() be the similarity measure, and τ be the predefined similarity threshold in the BERT-based function. If sim(ρ,B)≥τ, add L(v') to the candidate value set CandB.
[0070] Step S205: Based on the graph embedding model, calculate the correlation between the first attribute value and each candidate value in the candidate value set, determine the candidate value with the highest correlation, and correct the second attribute value in the entity tuple based on the candidate value with the highest correlation to obtain a corrected entity tuple.
[0071] In this embodiment, if CandB is not empty, a ranking model is used to obtain a recommended value for t[b]; otherwise, other remedial strategies may be needed to predict the value of t[b].
[0072] In the graph embedding model in the above step S202, the context information of the knowledge graph has been learned, and the graph embedding model can predict the correlation between the first attribute and the second attribute. Therefore, each candidate value in the set of all candidate values is assigned to the second attribute one by one, and the correlation between the first attribute value and each candidate value can be calculated, so that the candidate value corresponding to the highest correlation can be determined, and then the candidate value with the highest correlation can be used as the attribute value of the second attribute (that is, the second attribute value in the entity tuple is corrected according to the candidate value with the highest correlation), and finally the corrected entity tuple is obtained.
[0073] The embodiment of the present application obtains any entity tuple in the structured data to be corrected, constructs the data to be corrected based on the first attribute value of at least one first attribute and the second attribute value of the second attribute in the entity tuple, inputs the data to be corrected and a preset knowledge graph into a graph embedding model, outputs a first confidence level representing the correlation between the first attribute value and the second attribute value, inputs the data to be corrected into a language model, outputs a second confidence level representing the correlation between the first attribute value and the second attribute value, and if it is detected that the first confidence level and the second confidence level indicate that the second attribute value is incorrect, then, based on the first attribute value, a set of candidate values for the second attribute is matched from the preset knowledge graph, and based on the graph embedding model, the correlation between the first attribute value and each candidate value in the candidate value set is calculated, and the candidate value with the highest correlation is determined. Based on the candidate value with the highest correlation, the second attribute value in the entity tuple is corrected to obtain a corrected entity tuple. In this case, the prior knowledge of the knowledge graph, the graph embedding model, and the language model are combined to perform error detection, and the model in the error detection process is used to correct the incorrect attribute, which can effectively combine the prior knowledge to reason about the structured data, thereby improving the accuracy of data error correction.
[0074] See Figure 3, which is a flow chart of a method for entity error correction of structured data provided in Example 3 of the present application. As shown in Figure 3, in step S205, based on the graph embedding model, the correlation between the first attribute value and each candidate value in the candidate value set is calculated to determine the candidate value with the highest correlation. The following steps may also be included:
[0075] Step S301 : constructing error correction data corresponding to a candidate value according to a first attribute value and any candidate value in a candidate value set.
[0076] In this embodiment, the first attribute value of the above-mentioned first attribute and each candidate value in the candidate value set are respectively formed into an error correction data to obtain the error correction data corresponding to each candidate value, that is, the number of candidate values in the candidate value data set corresponds to the number of error correction data.
[0077] Step S302: Input the error correction data and the preset knowledge graph into the graph embedding model, and output a third confidence level representing the correlation between the first attribute value and the candidate value.
[0078] Among them, the data format, length, etc. of the error correction data are the same as those of the above-mentioned data to be corrected. The difference between the error correction data and the data to be corrected is only the difference between the second attribute value and the candidate value. Therefore, the preset knowledge graph and graph embedding model can be used for the error correction data to calculate the third confidence between the first attribute value and the candidate value.
[0079] The calculation process of the third confidence level is similar to the calculation process of the first confidence level. Please refer to the description of the calculation of the first confidence level, which will not be repeated here.
[0080] In addition, at this time in step S302, the graph encoder of the graph embedding model and the pre-trained graph data of the preset knowledge graph can be reused to realize the calculation of the third confidence level of the error correction data. That is, in actual use, a graph encoder and graph data of a graph embedding model are copied for separate use of the error correction data, so that the calculation of the error correction data is isolated from the calculation of the data to be corrected.
[0081] Step S303 , traverse all candidate values in the candidate value set, obtain the third confidence level of each candidate value, and determine the candidate value corresponding to the highest confidence level among all the third confidence levels as the candidate value with the highest correlation.
[0082] In this embodiment, each candidate value is assigned a third confidence level for the error correction data corresponding to it. A higher confidence level indicates a higher correlation between the candidate value and the first attribute value, and vice versa. Therefore, the highest confidence level can be selected from all third confidence levels, and the candidate value corresponding to the highest confidence level is the candidate value with the highest correlation.
[0083] The embodiment of the present application uses a preset knowledge graph and graph embedding model to select a suggested value from a set of candidate values obtained from the knowledge graph as the attribute value for correcting the second attribute, which can effectively and quickly combine prior knowledge to achieve the final correction.
[0084] See Figure 4, which is a flow chart of a method for entity error correction of structured data provided in Example 4 of the present application. As shown in Figure 4, in step S202, the data to be corrected and the preset knowledge graph are input into the graph embedding model, and a first confidence level representing the correlation between the first attribute value and the second attribute value is output. The following steps may also be included:
[0085] Step S401: Input the data to be corrected and the preset knowledge graph into the graph embedding model, and for any attribute value in the data to be corrected, find the label corresponding to the attribute value from the preset knowledge graph.
[0086] In this embodiment, the data to be corrected is serialized and other operations are performed using preset input rules to facilitate processing by the input value graph embedding model. For example, the sequence I = (t[A], t[b]). At this time, each attribute value is labeled, that is, I is labeled. The labels are obtained based on the lookup table Dict obtained by pre-training the preset knowledge graph to implement subsequent label embedding.
[0087] Step S402 : embedding a mark into a position corresponding to the attribute value in the data to be corrected to obtain the embedded data to be corrected, and converting the embedded data to be corrected into a first embedding matrix.
[0088] In this embodiment, the marker is embedded into the corresponding attribute value in the data to be corrected to represent the embedded data, and the embedded data to be corrected is represented using an embedding matrix. Specifically, if the marker Tt is found in Dict, the marker Tt is embedded as T = Dict[Tt]. If not, a normal (Gaussian) distribution is used for random initialization and embedding. Then, I is converted into a matrix MG = [T1; ...; T|I|]∈R|I|×d1, where d1 is the dimension of the graph embedding and |I| represents the integer number of attributes in I.
[0089] Step S403: perform graph coding on the first embedding matrix to obtain a first coding result.
[0090] In this embodiment, after obtaining the first embedding matrix, the graph encoder in the graph embedding model is used to perform graph encoding on the matrix to obtain the encoding result. Specifically, MG is encoded according to the graph encoder, represented by H(MG)∈Rd1×1: H(MG)=EncoderG(MG)=Poolmax(Attention(MG))
[0091] Among them, EncoderG is a graph encoder with attention mechanism Attention and maximum pooling strategy Poolmax.
[0092] Step S404: Calculate the probability of the correlation between the first attribute value and the second attribute value in the first encoding result, and obtain a first calculation result as a first confidence level.
[0093] In this embodiment, the probability of the correlation between the first attribute value and the second attribute value in the first encoding result is calculated based on the fully connected layer and the corresponding activation function, thereby characterizing the probability when the first attribute value and the second attribute value are correlated, that is, the first confidence level of the correlation between the two attribute values.
[0094] Specifically, the two-dimensional probability is calculated through the above fully connected layer and softmax activation to generate the confidence of the correlation between t[A] and t[b]: pG=Softmax(FCG(h(MG)))
[0095] Among them, pG[0] is the probability based on the graph embedding model, which indicates the probability that t[A] and t[b] are correlated, and pG[1] is the probability based on the graph embedding model, which indicates the probability that t[A] and t[b] are uncorrelated.
[0096] The embodiment of the present application can achieve the fusion of the knowledge graph and the graph embedding model through the above-mentioned graph embedding process, thereby helping to obtain the prediction result of the correlation between two attribute values through prior knowledge, so as to accurately obtain the confidence level of the correlation between the two attribute values, and provide a more accurate prerequisite for subsequent error detection.
[0097] 5 is a flow chart illustrating a method for entity error correction of structured data provided in a fifth embodiment of the present application. As shown in FIG5 , step S203 of inputting the data to be corrected into the language model and outputting a second confidence level representing the correlation between the first attribute value and the second attribute value may include the following steps:
[0098] Step S501 : Based on the input data expression rules of the language model, the data to be corrected is serialized to obtain sequence data.
[0099] In this embodiment, for the language model, the expression rules of the input data of the language model are obtained, so as to serialize the large error correction data and obtain the corresponding sequence data. Specifically, I is serialized as follows:
[0100] Where A={A1,…,Ak}, and are special markers that indicate the beginning of an attribute and a value, respectively. After I is serialized, it is input into the language model.
[0101] Step S502: input the sequence data into the language model, perform embedding encoding on each attribute value in the sequence data, and obtain a second encoding result.
[0102] In this embodiment, the language model can be used to analyze and obtain the tag after each attribute value in the sequence data, thereby implementing embedded encoding of the attribute value and obtaining the encoding result. Specifically, for the tag T in the serialized I, T = LM(T) is used as the embedding of the tag T. Then, I is converted into the matrix MLM = [T1; ...; T|serial(I)|]∈R|serial(I)|×d2, where d2 is the dimension of the language model embedding and |serial(I)| is the number of data in the serialized I.
[0103] Step S503: Calculate the probability of the correlation between the first attribute value and the second attribute value in the second encoding result, and obtain a second calculation result as a second confidence level.
[0104] In this embodiment, the probability of the correlation between the first attribute value and the second attribute value in the second encoding result is calculated based on the fully connected layer and the corresponding activation function, thereby characterizing the probability when the first attribute value and the second attribute value are correlated, that is, the second confidence level of the correlation between the two attribute values.
[0105] Specifically, the two-dimensional probability is calculated through the above fully connected layer and softmax activation to generate the confidence of the correlation between t[A] and t[b]: pLM=Softmax(FCLM(h(MLM)
[0106] Among them, pLM[0] is the probability based on the graph embedding model, which indicates the probability that t[A] and t[b] are correlated, and pLM[1] is the probability based on the graph embedding model, which indicates the probability that t[A] and t[b] are uncorrelated.
[0107] The embodiment of the present application realizes a multi-angle and multi-modal data analysis method through a language model, thereby providing a possibility for improving the accuracy of the prediction of the correlation between two attribute values, and also helps to improve the comprehensiveness of the analysis, so as to provide a more accurate prerequisite for subsequent error detection.
[0108] As shown in Figure 6, a schematic diagram of the model framework of an entity error correction method provided in Example 5 of the present application is provided, wherein a model Mc is proposed. The model Mc includes pre-training of the knowledge graph, a graph embedding model, and a language model. Through the pre-training process, the prior knowledge corresponding to the knowledge graph is integrated into (t[A], t[b]), and then encoded through the graph embedding model and the language model to realize the process of predicting correlation based on embedding technology. In addition, a self-supervised learning mechanism can be used to train the model Mc without the need for a large amount of manual labeling. As shown in Figure 6, the model Mc takes t[A] and t[b] (the table in Figure 6, where Name, Born, and College are t[A], and Affiliation is t[b]) as input, and outputs the confidence level representing their correlation after full connection. Finally, the two confidence levels are used to judge the detection results.
[0109] In Figure 6, HER is the process in which the user filters the knowledge graph to obtain a set of candidate values. The screening process of the candidate value set requires the use of the graph encoder of the graph embedding model to obtain the final candidate value to correct the incorrect attribute value. For obtaining the top-ranked value in CandB as the recommended value of t[b]. To this end, the lookup table pre-trained by the knowledge graph and the graph encoder of the graph embedding model are repeatedly used. Specifically, Idt = (t[A], b) is converted to a matrix M. Then, Idt is encoded as Idt = EncoderG(M). For each value c in CandB, its graph embedding c = Dict[c] is calculated. In order to map the embedding to the same latent space, two encoders EncoderI and Encoderc are used as follows: E[A] = EncoderI(Idt) = σ(FCI(Idt))
[0110] Where FC is the fully connected layer, σ is the sigmoid function, and E[A] (or E[c]) represents the final embedding of t[A] (or c). Intuitively, if E[A] is correlated with E[c], then t[B] is likely to take the value c. To measure the correlation, we have In addition, to make up for the lack of training data, a loss function with triples can be used to train the two encoders.
[0111] Therefore, the model Mc in Figure 6 can effectively integrate the process results of error detection and error repair of attributes, thereby achieving accurate error correction.
[0112] See Figure 7, which is a flow chart of a method for entity error correction of structured data provided in Example 6 of the present application. As shown in Figure 7, in step S204, if it is detected that the first confidence level and the second confidence level indicate that the second attribute value is incorrect, then before matching a set of candidate values for the second attribute from a preset knowledge graph based on the first attribute value, the following steps may be included:
[0113] Step S701: Perform weighted summation on the first confidence level and the second confidence level to obtain a weighted summation result as a final confidence level.
[0114] Step S702 : detecting whether the final confidence is greater than a confidence threshold; if it is detected that the final confidence is not greater than the confidence threshold, determining that the characterization result corresponding to the first confidence and the second confidence is that the second attribute value is wrong.
[0115] Step S703: If it is detected that the final confidence level is greater than the confidence level threshold, it is determined that the characterization result corresponding to the first confidence level and the second confidence level is that the second attribute value is correct.
[0116] In this embodiment, a confidence threshold is set so that the error detection result can be obtained according to the confidence level. The detection result is either an error or a correct result. An error can be represented by 0, and a correct result can be represented by 1. A higher confidence level indicates a higher probability of correctness, while a lower confidence level indicates a higher probability of error.
[0117] The correlation between attributes is determined by the model Mc shown in FIG6 , so that the final confidence value of the correlation between t[A] and t[b] is: Mc(t[A], t[b])=αpG[0]+(1-α)pLM[0],
[0118] Among them, α is a hyperparameter that balances the two mechanisms. In this way, not only the contextual knowledge from the knowledge graph is utilized, but also the results are enhanced by rich semantics, making the detection results more accurate.
[0119] See Figure 8, which is a flow chart of a method for entity error correction of structured data provided in Example 7 of the present application. As shown in Figure 8, after matching a set of candidate values for the second attribute from a preset knowledge graph based on the first attribute value in step S204 above, the following steps may also be included:
[0120] Step S801: Check whether the candidate value set is empty. If it is detected that the candidate value set is not empty, execute the graph embedding model to calculate the correlation between the first attribute value and each candidate value in the candidate value set. According to the candidate value with the highest correlation, correct the second attribute value in the first structured data to obtain the corrected first structured data.
[0121] Step S802: If it is detected that the candidate value set is empty, the first attribute value is input into the prediction model, and the predicted value corresponding to the second attribute is output. According to the predicted value, the second attribute value in the first structured data is corrected to obtain the corrected first structured data.
[0122] In this embodiment, if a candidate value set can be obtained from the knowledge graph and the candidate value set is not an empty set, then step S205 can be used. If the candidate value set is an empty set, then no error correction value can be obtained from the candidate value set. This embodiment provides a remedial strategy to predict an error correction value. Specifically, a prediction model is required. This prediction model and the first attribute value are used to predict the attribute value of the second attribute, and ultimately to achieve error correction.
[0123] Optionally, the prediction model is a twin network model, and the first attribute value is input into the prediction model, and the prediction value corresponding to the second attribute is output, including:
[0124] Inputting the first attribute value into the first network of the twin network model to obtain a first embedding result output by the first network;
[0125] Obtaining a predicted value range for the second attribute, selecting a target value from the predicted value range and inputting it into the second network of the twin network model to obtain a second embedding result for the second output;
[0126] Calculate the correlation between the first embedding result and the second embedding result, return to execute the second network of the twin network model by selecting a target value from the predicted value range, and obtain the second embedding result of the second output until the highest correlation is detected. Determine that the target value corresponding to the highest correlation is the predicted value of the second attribute.
[0127] Among them, when CandB is empty, a prediction model is trained to predict t[b]. A twin network model such as sentenceBert can be used as the basic model, t[A] is regarded as a sequence, and its embedding is returned, and sentenceBert(t[A]) is represented by E[A]. Similarly, for a value c in dom(b), where dom(b) is the predicted value range of b (i.e., the activity domain), E[c] = sentenceBert(c) is calculated. Then, let
[0128] The above prediction logic combines two parameter-sharing sentenceBert models. During inference, the Faiss vector search library is used to retrieve the top-ranked value from dom(b). This twin network model allows for better prediction results and improves error correction accuracy.
[0129] Corresponding to the entity error correction method for structured data in the above embodiment, FIG9 shows a block diagram of the structured data entity error correction device provided in the eighth embodiment of the present application. The above entity error correction device is applied to the server in FIG1 . The user corresponding to the client sends the structured data to be corrected to the server, triggering the server to perform the data error detection and repair task. For ease of explanation, only the part relevant to the embodiment of the present application is shown.
[0130] Referring to FIG9 , the physical error correction device includes:
[0131] The to-be-corrected data determining module 91 is configured to obtain any entity tuple in the to-be-corrected structured data, and construct the to-be-corrected data based on a first attribute value of at least one first attribute and a second attribute value of a second attribute in the entity tuple, wherein the first attribute is different from the second attribute;
[0132] A graph embedding model analysis module 92 is configured to input the data to be corrected and a preset knowledge graph into a graph embedding model, and output a first confidence level representing the correlation between the first attribute value and the second attribute value;
[0133] A language model analysis module 93 is configured to input the data to be corrected into a language model and output a second confidence level representing the correlation between the first attribute value and the second attribute value;
[0134] A candidate value determination module 94 is configured to match a set of candidate values for the second attribute from a preset knowledge graph based on the first attribute value if the first confidence level and the second confidence level indicate that the second attribute value is incorrect.
[0135] The attribute value error correction module 95 is used to calculate the correlation between the first attribute value and each candidate value in the candidate value set based on the graph embedding model, determine the candidate value with the highest correlation, and correct the second attribute value in the entity tuple based on the candidate value with the highest correlation to obtain a corrected entity tuple.
[0136] Optionally, the attribute value error correction module 95 includes:
[0137] an error correction data determination unit, configured to construct error correction data corresponding to a candidate value based on the first attribute value and any candidate value in the candidate value set;
[0138] a confidence determination unit, configured to input the error correction data and a preset knowledge graph into a graph embedding model, and output a third confidence level representing the correlation between the first attribute value and the candidate value;
[0139] The candidate value determination unit is configured to traverse all candidate values in the candidate value set, obtain a third confidence level for each candidate value, and determine the candidate value corresponding to the highest confidence level among all the third confidence levels as the candidate value with the highest correlation.
[0140] Optionally, the graph embedding model analysis module 92 includes:
[0141] A tag search unit is used to input the data to be corrected and a preset knowledge graph into the graph embedding model, and for any attribute value in the data to be corrected, find the tag corresponding to the attribute value from the preset knowledge graph;
[0142] A first matrix determination unit is configured to embed a marker into a position corresponding to an attribute value in the data to be corrected, obtain the embedded data to be corrected, and convert the embedded data to be corrected into a first embedding matrix;
[0143] A first encoding unit, configured to perform graph encoding on the first embedding matrix to obtain a first encoding result;
[0144] The first confidence calculation unit is configured to calculate the probability of the correlation between the first attribute value and the second attribute value in the first encoding result, and obtain a first calculation result as a first confidence.
[0145] Optionally, the language model analysis module 93 includes:
[0146] A sequence data determination unit is used to serialize the data to be corrected based on the input data expression rules of the language model to obtain sequence data;
[0147] A second encoding unit is used to input the sequence data into the language model, perform embedding encoding on each attribute value in the sequence data, and obtain a second encoding result;
[0148] The second confidence calculation unit is configured to calculate the probability of the correlation between the first attribute value and the second attribute value in the second encoding result, and obtain a second calculation result as a second confidence.
[0149] Optionally, the physical error correction device further includes:
[0150] A final confidence calculation module is configured to, if it is detected that the first confidence and the second confidence indicate that the second attribute value is incorrect, perform a weighted sum of the first confidence and the second confidence before matching a set of candidate values for the second attribute from a preset knowledge graph based on the first attribute value, and obtain the weighted sum as a final confidence;
[0151] A first determination module is configured to detect whether a final confidence level is greater than a confidence threshold, and if it is detected that the final confidence level is not greater than the confidence threshold, determine that a characterization result corresponding to the first confidence level and the second confidence level is that the second attribute value is incorrect;
[0152] The second determination module is configured to determine that the characterization result corresponding to the first confidence level and the second confidence level is correct if it is detected that the final confidence level is greater than the confidence level threshold.
[0153] Optionally, the physical error correction device further includes:
[0154] A first error correction module is configured to, after matching a set of candidate values for the second attribute from a preset knowledge graph based on the first attribute value, detect whether the candidate value set is empty; if it is detected that the candidate value set is not empty, perform a graph embedding model-based calculation to calculate a correlation between the first attribute value and each candidate value in the candidate value set; and perform error correction on the second attribute value in the first structured data based on the candidate value with the highest correlation, thereby obtaining the corrected first structured data;
[0155] The second error correction module is used to input the first attribute value into the prediction model if it is detected that the candidate value set is empty, output the predicted value corresponding to the second attribute, and correct the second attribute value in the first structured data according to the predicted value to obtain the corrected first structured data.
[0156] Optionally, the prediction model is a twin network model, and the second error correction module includes:
[0157] A first network execution unit is configured to input the first attribute value into the first network of the twin network model to obtain a first embedding result output by the first network;
[0158] A second network execution unit is configured to obtain a predicted value range of a second attribute, select a target value from the predicted value range, input the target value into the second network of the twin network model, and obtain a second embedding result of a second output;
[0159] The correlation calculation unit is used to calculate the correlation between the first embedding result and the second embedding result, return to execute the second network of the twin network model by selecting a target value from the predicted value range, and obtain the second embedding result of the second output until the highest correlation is detected, and determine that the target value corresponding to the highest correlation is the predicted value of the second attribute.
[0160] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0161] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as shown in FIG10 . The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a readable storage medium, and a database. The internal memory provides an environment for the operation of the operating system and the readable storage medium in the non-volatile storage medium. The database of the computer device is used to store user original data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the readable storage medium is executed by the processor, an automated entity splitting method is implemented.
[0162] In one embodiment, a computer device is provided, comprising a memory, a processor, and a readable storage medium stored in the memory and operable on the processor. When the processor executes the readable storage medium, the steps of the automated entity splitting method in the above embodiment are implemented, such as steps S201-S205 shown in FIG2 , or the steps shown in FIG3 to FIG8 . To avoid repetition, they are not described here. Alternatively, when the processor executes the readable storage medium, the functions of the modules / units in the embodiment of the user data processing device are implemented, such as the functions of the data to be corrected determination module 91, the graph embedding model analysis module 92, the language model analysis module 93, the candidate value determination module 94, and the attribute value correction module 95 shown in FIG9 . To avoid repetition, they are not described here.
[0163] In one embodiment, one or more readable storage media storing computer-readable instructions are provided. The computer-readable storage media stores computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors implement the steps of the automated entity splitting method in the above-mentioned embodiment, such as steps S201-S205 shown in FIG2 , or the steps shown in FIG3 to FIG8 . To avoid repetition, they are not described here. Alternatively, when the processor executes the readable storage medium, the functions of each module / unit in this embodiment of the user data processing device are implemented, such as the functions of the data to be corrected determination module 91, the graph embedding model analysis module 92, the language model analysis module 93, the candidate value determination module 94, and the attribute value correction module 95 shown in FIG9 . To avoid repetition, they are not described here. The readable storage medium in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0164] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a readable storage medium, and the readable storage medium can be stored in a non-volatile computer-readable storage medium. When the readable storage medium is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0165] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0166] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. An entity error correction method for structured data, characterized in that The entity error correction method includes: Obtain any entity tuple in the structured data to be error-corrected, and construct error-correction data according to the first attribute value of at least one first attribute and the second attribute value of a second attribute in the entity tuple, where the first attribute is different from the second attribute; Input the error-correction data and a preset knowledge graph into a graph embedding model, and output a first confidence level representing the correlation between the first attribute value and the second attribute value; Input the error-correction data into a language model, and output a second confidence level representing the correlation between the first attribute value and the second attribute value; If it is detected that the first confidence level and the second confidence level indicate that the second attribute value is incorrect, then according to the first attribute value, match a candidate value set of the second attribute from the preset knowledge graph; Based on the graph embedding model, calculate the correlation between the first attribute value and each candidate value in the candidate value set, determine the candidate value with the highest correlation, and correct the second attribute value in the entity tuple according to the candidate value with the highest correlation to obtain an error-corrected entity tuple.
2. The entity error correction method according to claim 1, characterized in that, The calculating the correlation between the first attribute value and each candidate value in the candidate value set based on the graph embedding model and determining the candidate value with the highest correlation includes: According to the first attribute value and any candidate value in the candidate value set, construct error-correction data corresponding to the candidate value; Input the error-correction data and the preset knowledge graph into the graph embedding model, and output a third confidence level representing the correlation between the first attribute value and the candidate value; Traverse all candidate values in the candidate value set to obtain the third confidence level of each candidate value, and determine the candidate value corresponding to the highest confidence level among all third confidence levels as the candidate value with the highest correlation.
3. The entity error correction method according to claim 1, characterized in that The inputting the error-correction data and a preset knowledge graph into a graph embedding model and outputting a first confidence level representing the correlation between the first attribute value and the second attribute value includes: Input the error-correction data and a preset knowledge graph into a graph embedding model, and for any attribute value in the error-correction data, find the mark corresponding to the attribute value from the preset knowledge graph; Embed the mark into the position corresponding to the attribute value in the error-correction data to obtain the embedded error-correction data, and convert the embedded error-correction data into a first embedding matrix; Perform graph encoding on the first embedding matrix to obtain a first encoding result; Calculate the probability of the correlation between the first attribute value and the second attribute value in the first encoding result to obtain a first calculation result as the first confidence level.
4. The entity error correction method according to claim 1, characterized in that, The inputting the error-correction data into a language model and outputting a second confidence level representing the correlation between the first attribute value and the second attribute value includes: Based on the input data expression rule of the language model, perform serialization expression on the error-correction data to obtain sequence data; Input the sequence data into the language model, and perform embedding encoding on each attribute value in the sequence data to obtain a second encoding result; Calculate the probability of the correlation between the first attribute value and the second attribute value in the second coding result, and obtain the second calculation result as the second confidence level.
5. The entity error correction method according to claim 1, characterized in that, Before the step of, if it is detected that the first confidence level and the second confidence level indicate that the second attribute value is incorrect, then according to the first attribute value, matching the candidate value set of the second attribute from the preset knowledge graph, it further includes: Perform a weighted sum of the first confidence level and the second confidence level, and obtain the weighted sum result as the final confidence level; Detect whether the final confidence level is greater than the confidence level threshold. If it is detected that the final confidence level is not greater than the confidence level threshold, then determine that the representation result corresponding to the first confidence level and the second confidence level is that the second attribute value is incorrect; If it is detected that the final confidence level is greater than the confidence level threshold, then determine that the representation result corresponding to the first confidence level and the second confidence level is that the second attribute value is correct.
6. The entity error correction method according to any one of claims 1 to 5, characterized in that After the step of, according to the first attribute value, matching the candidate value set of the second attribute from the preset knowledge graph, it further includes: Detect whether the candidate value set is empty. If it is detected that the candidate value set is not empty, then execute the step of calculating the correlation between the first attribute value and each candidate value in the candidate value set based on the graph embedding model, and correcting the second attribute value in the first structured data according to the candidate value with the highest correlation to obtain the corrected first structured data; If it is detected that the candidate value set is empty, then input the first attribute value into the prediction model, output the predicted value corresponding to the second attribute, and correct the second attribute value in the first structured data according to the predicted value to obtain the corrected first structured data.
7. The entity error correction method according to claim 6, wherein The prediction model is a siamese network model. The step of inputting the first attribute value into the prediction model and outputting the predicted value corresponding to the second attribute includes: Input the first attribute value into the first network of the siamese network model, and obtain the first embedding result output by the first network; Obtain the range of the predicted value of the second attribute, select a target value from the range of the predicted value and input it into the second network of the siamese network model, and obtain the second embedding result output by the second network; Calculate the correlation between the first embedding result and the second embedding result, and return to execute the step of selecting a target value from the range of the predicted value and inputting it into the second network of the siamese network model to obtain the second embedding result output by the second network until the highest correlation is detected, and determine the target value corresponding to the highest correlation as the predicted value of the second attribute.
8. An entity error correction device for structured data, characterized in that, The entity error correction device includes: A to-be-corrected data determination module, configured to obtain any entity tuple in the to-be-corrected structured data, and construct to-be-corrected data according to the first attribute value of at least one first attribute and the second attribute value of a second attribute in the entity tuple, where the first attribute is different from the second attribute; A graph embedding model analysis module, configured to input the data to be corrected and a preset knowledge graph into a graph embedding model, and output a first confidence level representing the correlation between the first attribute value and the second attribute value; A language model analysis module, configured to input the data to be corrected into a language model, and output a second confidence level representing the correlation between the first attribute value and the second attribute value; A candidate value determination module, configured to, if it is detected that the first confidence level and the second confidence level represent that the second attribute value is incorrect, match a set of candidate values of the second attribute from the preset knowledge graph according to the first attribute value; An attribute value correction module, configured to calculate the correlation between the first attribute value and each candidate value in the set of candidate values based on the graph embedding model, determine the candidate value with the highest correlation, and correct the second attribute value in the entity tuple according to the candidate value with the highest correlation to obtain a corrected entity tuple.
9. A computer device, comprising a memory, a processor, and a readable storage medium stored in the memory and operable on the processor, wherein, When the processor executes the readable storage medium, the following steps are implemented: Obtain any entity tuple in the structured data to be corrected, and construct data to be corrected according to the first attribute value of at least one first attribute and the second attribute value of a second attribute in the entity tuple, where the first attribute is different from the second attribute; Input the data to be corrected and a preset knowledge graph into a graph embedding model, and output a first confidence level representing the correlation between the first attribute value and the second attribute value; Input the data to be corrected into a language model, and output a second confidence level representing the correlation between the first attribute value and the second attribute value; If it is detected that the first confidence level and the second confidence level represent that the second attribute value is incorrect, match a set of candidate values of the second attribute from the preset knowledge graph according to the first attribute value; Based on the graph embedding model, calculate the correlation between the first attribute value and each candidate value in the set of candidate values, determine the candidate value with the highest correlation, and correct the second attribute value in the entity tuple according to the candidate value with the highest correlation to obtain a corrected entity tuple.
10. The computer device according to claim 9, wherein, The calculating the correlation between the first attribute value and each candidate value in the set of candidate values based on the graph embedding model and determining the candidate value with the highest correlation includes: Constructing corrected data corresponding to the candidate value according to the first attribute value and any candidate value in the set of candidate values; Inputting the corrected data and the preset knowledge graph into the graph embedding model, and outputting a third confidence level representing the correlation between the first attribute value and the candidate value; Traversing all candidate values in the set of candidate values to obtain the third confidence level of each candidate value, and determining the candidate value corresponding to the highest confidence level among all the third confidence levels as the candidate value with the highest correlation.
11. The computer device according to claim 9, wherein, The inputting the data to be corrected and a preset knowledge graph into a graph embedding model and outputting a first confidence level representing the correlation between the first attribute value and the second attribute value includes: Input the data to be error-corrected and the preset knowledge graph into the graph embedding model. For any attribute value in the data to be error-corrected, find the corresponding tag from the preset knowledge graph; Embed the tag into the position corresponding to the attribute value in the data to be error-corrected to obtain the data to be error-corrected after embedding, and convert the data to be error-corrected after embedding into the first embedding matrix; Perform graph encoding on the first embedding matrix to obtain the first encoding result; Calculate the probability of the correlation between the first attribute value and the second attribute value in the first encoding result to obtain the first calculation result as the first confidence level.
12. The computer device according to claim 9, wherein, The step of inputting the data to be error-corrected into the language model and outputting the second confidence level representing the correlation between the first attribute value and the second attribute value includes: Based on the input data expression rule of the language model, perform serialization expression on the data to be error-corrected to obtain sequence data; Input the sequence data into the language model, and perform embedding encoding on each attribute value in the sequence data to obtain the second encoding result; Calculate the probability of the correlation between the first attribute value and the second attribute value in the second encoding result to obtain the second calculation result as the second confidence level.
13. The computer device according to claim 9, wherein, Before the step of, if it is detected that the first confidence level and the second confidence level represent that the second attribute value is incorrect, then match the candidate value set of the second attribute from the preset knowledge graph according to the first attribute value, further includes: Perform weighted summation on the first confidence level and the second confidence level to obtain the weighted summation result as the final confidence level; Detect whether the final confidence level is greater than the confidence level threshold. If it is detected that the final confidence level is not greater than the confidence level threshold, then determine that the representation result corresponding to the first confidence level and the second confidence level is that the second attribute value is incorrect; If it is detected that the final confidence level is greater than the confidence level threshold, then determine that the representation result corresponding to the first confidence level and the second confidence level is that the second attribute value is correct.
14. The computer device according to any one of claims 9 to 13, wherein, After the step of matching the candidate value set of the second attribute from the preset knowledge graph according to the first attribute value, further includes: Detect whether the candidate value set is empty. If it is detected that the candidate value set is not empty, then perform the following operations based on the graph embedding model: calculate the correlation between the first attribute value and each candidate value in the candidate value set, and correct the second attribute value in the first structured data according to the candidate value with the highest correlation to obtain the first structured data after error correction; If it is detected that the candidate value set is empty, then input the first attribute value into the prediction model, output the predicted value corresponding to the second attribute, and correct the second attribute value in the first structured data according to the predicted value to obtain the first structured data after error correction.
15. The computer device according to claim 14, wherein, The prediction model is a siamese network model. The step of inputting the first attribute value into the prediction model and outputting the predicted value corresponding to the second attribute includes: Input the first attribute value into the first network of the siamese network model to obtain the first embedding result output by the first network; Obtain the predicted value range of the second attribute, select a target value from the predicted value range and input it into the second network of the twin network model to obtain the second embedding result of the second output; Calculate the correlation between the first embedding result and the second embedding result, return to execute the step of selecting a target value from the predicted value range and inputting it into the second network of the twin network model to obtain the second embedding result of the second output, until the highest correlation is detected, and determine the target value corresponding to the highest correlation as the predicted value of the second attribute.
16. One or more readable storage media storing computer-readable instructions, the computer-readable storage media storing computer-readable instructions, wherein, When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the following steps: Obtain any entity tuple in the structured data to be corrected, and construct the data to be corrected according to the first attribute values of at least one first attribute and the second attribute value of a second attribute in the entity tuple, where the first attribute is different from the second attribute; Input the data to be corrected and a preset knowledge graph into a graph embedding model, and output a first confidence level representing the correlation between the first attribute value and the second attribute value; Input the data to be corrected into a language model, and output a second confidence level representing the correlation between the first attribute value and the second attribute value; If it is detected that the first confidence level and the second confidence level represent that the second attribute value is incorrect, then according to the first attribute value, match a set of candidate values of the second attribute from the preset knowledge graph; Based on the graph embedding model, calculate the correlation between the first attribute value and each candidate value in the set of candidate values, determine the candidate value with the highest correlation, and correct the second attribute value in the entity tuple according to the candidate value with the highest correlation to obtain the corrected entity tuple.
17. The readable storage medium according to claim 16, wherein, The calculating the correlation between the first attribute value and each candidate value in the set of candidate values based on the graph embedding model and determining the candidate value with the highest correlation includes: Construct the data to be corrected corresponding to the candidate value according to the first attribute value and any candidate value in the set of candidate values; Input the data to be corrected and the preset knowledge graph into the graph embedding model, and output a third confidence level representing the correlation between the first attribute value and the candidate value; Traverse all candidate values in the set of candidate values to obtain the third confidence level of each candidate value, and determine the candidate value corresponding to the highest confidence level among all the third confidence levels as the candidate value with the highest correlation.
18. The readable storage medium according to claim 16, wherein, The inputting the data to be corrected and the preset knowledge graph into the graph embedding model and outputting a first confidence level representing the correlation between the first attribute value and the second attribute value includes: Input the data to be corrected and the preset knowledge graph into the graph embedding model, and for any attribute value in the data to be corrected, find the label corresponding to the attribute value from the preset knowledge graph; Embed the label into the position corresponding to the attribute value in the data to be corrected to obtain the embedded data to be corrected, and convert the embedded data to be corrected into a first embedding matrix; Perform graph encoding on the first embedding matrix to obtain a first encoding result; Calculate the probability of the correlation between the first attribute value and the second attribute value in the first encoding result, and obtain a first calculation result as the first confidence level.
19. The readable storage medium according to claim 16, wherein, The step of inputting the data to be error-corrected into the language model and outputting a second confidence level representing the correlation between the first attribute value and the second attribute value includes: Based on the input data expression rule of the language model, serialize the data to be error-corrected to obtain sequence data; Input the sequence data into the language model, and perform embedding encoding on each attribute value in the sequence data to obtain a second encoding result; Calculate the probability of the correlation between the first attribute value and the second attribute value in the second encoding result, and obtain a second calculation result as the second confidence level.
20. The readable storage medium according to claim 16, wherein Before the step of, if it is detected that the first confidence level and the second confidence level represent that the second attribute value is incorrect, then according to the first attribute value, match the candidate value set of the second attribute from the preset knowledge graph, further includes: Perform weighted summation on the first confidence level and the second confidence level to obtain a weighted summation result as the final confidence level; Detect whether the final confidence level is greater than the confidence level threshold. If it is detected that the final confidence level is not greater than the confidence level threshold, then determine that the representation result corresponding to the first confidence level and the second confidence level is that the second attribute value is incorrect; If it is detected that the final confidence level is greater than the confidence level threshold, then determine that the representation result corresponding to the first confidence level and the second confidence level is that the second attribute value is correct.
Citation Information
Patent Citations
Case information semantic retrieval method and device based on knowledge graph
CN111475623A
Text error correction method, device, equipment and computer readable medium
CN114036930A
Government affair field multi-stage fusion text error correction method based on knowledge graph
CN116502628A
Knowledge graph correction method and device
CN116628225A
Knowledge graph construction method based on reliability of aircraft parts
CN116644192A
Cited By
Patent retrieval method based on feature-efficacy matrix and large language model
CN121722903A
Data governance anomaly detection and identification method for knowledge graph analysis
CN121980469A