A conflict elimination query method and device for multi-source heterogeneous data

By converting multi-source heterogeneous data into knowledge graphs and matching them in line graphs, combined with iterative estimation of confidence, the problem of query results conflict in centralized data management systems is solved, and efficient and flexible data fusion and query result output are achieved.

CN118484542BActive Publication Date: 2025-05-16ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410566416.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-09
Publication Date
2025-05-16
Estimated Expiration
2044-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of query result conflict caused by multi-source heterogeneous data in centralized data management systems, and has high spatiotemporal complexity and low scalability, so it cannot adapt to scenarios where data updates and dynamic changes.

Method used

By converting the data set of heterogeneous data sources to be found into an equivalent knowledge graph, and converting the query input by the user into an equivalent knowledge graph, matching using line graph representation, finding nodes that match semantics and network structures, forming candidate query results, and estimating the trustworthiness of the data source and data through iteratively, finally selecting the entity with the highest confidence as the output of the query result.

Benefits of technology

It realizes the acquisition of trusted and consistent query results in multi-source heterogeneous data sources, improves the flexibility and efficiency of data fusion, and adapts to the dynamic changes of multi-source heterogeneous data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118484542B_ABST
    Figure CN118484542B_ABST
Patent Text Reader

Abstract

The present invention discloses a conflict elimination query method and device for multi-source heterogeneous data. The method converts each data set and user query in a heterogeneous data source into a data map and a query map; converts the data map and the query map into a line graph; according to the semantic information and network structure of the nodes in the query line graph, finds all nodes that match the semantics and network structure of the query line graph nodes from all data line graphs; obtains nodes that simultaneously meet the semantic matching and network structure matching, and obtains entities that match the entities to be queried in the query map to form candidate query results; assigns initial credibility to each data set in the data source and each entity in the candidate query result; according to the candidate query result and the historical query results of each data source, iteratively estimates the data source credibility and data credibility until convergence; according to the converged data credibility, selects the entity with the highest credibility in the candidate query result as the query result output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of centralized data management, and in particular relates to a conflict elimination query method and device for multi-source heterogeneous data. Background Art

[0002] With the explosive growth of data volume, major enterprises choose to store multi-source heterogeneous data in centralized data management systems (such as data lakes), which usually support queries on multi-source heterogeneous data. However, due to the uneven quality of data from different sources, query results from different data sources often conflict, which in turn leads to inaccurate and unreliable query results, seriously reducing the availability of centralized data management systems. Therefore, resolving conflicts in cross-source query results is one of the most important issues facing centralized data management systems.

[0003] In order to solve this problem, many methods based on batch data fusion have been proposed. These methods estimate the credibility of each data and each data source one by one by traversing all the data in the data management system, and finally select some data with credibility greater than the threshold according to the artificially set credibility threshold as the data fusion result, and finally form a clean and consistent data to serve the downstream data analysis tasks.

[0004] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:

[0005] Most of these methods have high spatiotemporal complexity and low scalability, making them difficult to apply to large centralized data management systems in reality. Secondly, these methods cannot cope with frequent data updates and are difficult to adapt to the dynamic changes of data lakes in actual application scenarios. In addition, these methods rely on data matching preprocesses such as metadata matching and entity matching. The accuracy and efficiency of the data matching preprocess will directly affect the accuracy and efficiency of batch data fusion. The above problems make it difficult for existing data fusion technologies to be widely applied to the centralized management scenarios of multi-source heterogeneous data in reality. Summary of the invention

[0006] The purpose of the embodiments of the present application is to provide a conflict elimination query method and device for multi-source heterogeneous data, so as to solve the technical problem existing in the related art that it is difficult to apply to the centralized management of multi-source heterogeneous data.

[0007] According to a first aspect of an embodiment of the present application, a conflict elimination query method for multi-source heterogeneous data is provided, comprising:

[0008] Convert each data set in the to-be-queried heterogeneous data source into an equivalent knowledge graph, denoted as a data graph, and convert the user input query into an equivalent knowledge graph in the same way, denoted as a query graph, wherein each knowledge graph includes a number of knowledge tuples, each of which includes two entities, and there are a number of to-be-queried entities in the query graph, representing the user's target query result;

[0009] Convert the query graph input by the user and the data graph converted from each data set in the heterogeneous data source into a line graph representation to obtain a corresponding query line graph and data line graph;

[0010] According to the semantic information of each node in the query line graph and the network structure characteristics of the query line graph, all nodes that match the semantics and network structure of the query line graph nodes are found from all data line graphs;

[0011] Based on the nodes that satisfy both semantic matching and network structure matching, the entities in the nodes that match the entities to be queried in the query graph are obtained to form candidate query results;

[0012] Assign initial data source credibility and initial data credibility to each data set in the heterogeneous data source and each data in the candidate query result;

[0013] According to the candidate query results and the historical query results of each data set, the data source credibility and data credibility are iteratively estimated until the two converge;

[0014] According to the converged data credibility, the entity with the highest credibility among the candidate query results is selected as the query result output.

[0015] Furthermore, the query graph input by the user and the data graph converted from each data set in the heterogeneous data source are converted into a line graph representation to obtain corresponding query line graphs and data line graphs, including:

[0016] Create a one-to-one correspondence line graph for the query graph and the data graph, which correspond to the query line graph and the data line graph respectively. The created line graph is initially empty;

[0017] Store the multi-tuples in the query graph and the data graph as nodes in the corresponding line graph;

[0018] Consider two nodes in the line graph, which correspond to two tuples in the knowledge graph. If the same entity exists in the two tuples, an edge is formed between the two nodes in the line graph. Otherwise, no edge can be formed between the two nodes, thereby completing the edge between the nodes in the query line graph and the edge between the nodes in the data line graph.

[0019] Furthermore, according to the semantic information of each node in the query line graph and the network structure characteristics of the query line graph, all nodes that match the semantics and network structure of the query line graph nodes are found from all data line graphs, including:

[0020] According to the semantic information of each node on the query line graph and the data line graph, the node semantic information is represented by using a pre-trained language model to form a semantic representation vector of each node on the query line graph and the data line graph respectively;

[0021] According to the semantic representation vectors of each node in the query line graph and the data line graph, the similarity between each node in the query line graph and all nodes in the data line graph is obtained, and all nodes in the data line graph whose similarity exceeds a threshold are found;

[0022] According to the network structure characteristics of the query line graph, the subgraph matching technology is used to find the subgraph structure that matches the network structure of the query line graph from the data line graph, and all nodes contained in the found matching subgraph are recorded.

[0023] Furthermore, based on the candidate query results and the historical query results of each data set in the data source, the data source credibility and data credibility are iteratively estimated until convergence, including:

[0024] Estimate and update the data source credibility based on the candidate query results and their corresponding data credibility and the historical query results of the data source;

[0025] Estimate and update data credibility based on the candidate query results and their corresponding data credibility and the updated data source credibility;

[0026] Iterate the data source credibility and data credibility calculation update process until the two converge.

[0027] Furthermore, for a data set D in a data source, the data source credibility Pr(D) is calculated as follows:

[0028]

[0029] Where Data(Q,D) represents the partial query results provided by the dataset D in the candidate query results of the query Q, v represents a specific entity in the partial query result Data(Q,D), Pr(v) represents the data credibility of one of the query results, namely the entity v, and Pr(D|v) represents the probability that the dataset D is credible under the premise that v is credible. The calculation method of Pr(D|v) is as follows:

[0030]

[0031] Among them, Pr h(D) represents the historical value of the data source credibility of dataset D, H represents the number of query results provided by dataset D in historical queries, and D v [q] represents the entity subset in Data(Q,Q) whose data credibility is greater than or equal to entity v, that is,

[0032] Furthermore, for an entity v in the candidate query results, its data credibility Pr(v) is calculated as follows:

[0033]

[0034] in, represents the set of all data sets in the heterogeneous data sources. The calculation of Pr(D|v) is as described in step S31. Pr(v|D) represents the probability that entity v is credible under the premise that data set D is credible. The calculation method of Pr(v|D) is as follows:

[0035]

[0036] According to a second aspect of an embodiment of the present application, a conflict elimination query device for multi-source heterogeneous data is provided, comprising:

[0037] An acquisition module is used to acquire a knowledge graph converted from each data set in a heterogeneous data source and a knowledge graph converted from a user input query, which are respectively recorded as a data graph and a query graph, wherein each knowledge graph includes a number of knowledge tuples, each of which includes two entities, and there are a number of entities to be queried in the query graph, representing the user's target query result;

[0038] A conversion module, used to convert the query graph input by the user and the data graph converted from each data set in the heterogeneous data source into a line graph representation to obtain a corresponding query line graph and data line graph;

[0039] A matching module is used to find all nodes that are semantically matched and network-structured with the query line graph nodes from all data line graphs according to the semantic information of each node in the query line graph and the network structure characteristics of the query line graph;

[0040] A matching and merging module is used to obtain entities in the nodes that match the entities to be queried in the query graph based on nodes that satisfy both semantic matching and network structure matching, thereby forming candidate query results;

[0041] A credibility initialization module is used to assign initial data source credibility and initial data credibility to each data set in the heterogeneous data source and each data in the candidate query result;

[0042] The credibility iterative calculation module is used to iteratively estimate the credibility of the data source and the credibility of the data according to the candidate query results and the historical query results of each data set in the data source until the two converge;

[0043] The output module is used to select the entity with the highest credibility among the candidate query results as the query result output according to the converged data credibility.

[0044] According to a third aspect of an embodiment of the present application, a computer program product is provided, comprising a computer program / instruction, which implements the method described in the first aspect when executed by a processor.

[0045] According to a fourth aspect of an embodiment of the present application, there is provided an electronic device, including:

[0046] one or more processors;

[0047] A memory for storing one or more programs;

[0048] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.

[0049] According to a fifth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in the first aspect are implemented.

[0050] The technical solution provided by the embodiments of the present application may have the following beneficial effects:

[0051] It can be seen from the above embodiments that the present application converts the query problem of multiple heterogeneous data sources into a knowledge graph matching problem, and introduces line graph conversion to accelerate the knowledge graph matching, thereby obtaining candidate query results that match the user query in each heterogeneous data source. Considering that each data source may provide conflicting data and unreliable data, a small batch data fusion algorithm that only considers the candidate query results is designed to eliminate unreliable conflicting data in the candidate query results, and finally obtain reliable and consistent data from multiple heterogeneous data sources as query results. On the one hand, the present application realizes trusted queries for multiple heterogeneous data sources, and on the other hand, it realizes more flexible and efficient data fusion.

[0052] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0054] Figure 1 The present invention is a flowchart showing a conflict elimination query method for multi-source heterogeneous data according to an exemplary embodiment.

[0055] Figure 2 It is a schematic diagram of knowledge graph line graph conversion according to an exemplary embodiment.

[0056] Figure 3 is a flowchart of step S13 according to an exemplary embodiment.

[0057] Figure 4 The diagram is an example of node semantic matching and network structure matching according to an exemplary embodiment.

[0058] Figure 5 is a flowchart of step S16 according to an exemplary embodiment.

[0059] Figure 6 The present invention is a block diagram showing a conflict elimination query device for multi-source heterogeneous data according to an exemplary embodiment.

[0060] Figure 7 The diagram is a schematic diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0061] Here, exemplary embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application.

[0062] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0063] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0064] Glossary:

[0065] Heterogeneous data sources: Heterogeneous data sources usually contain multiple copies of data, each stored in different formats, such as tables, knowledge graphs, text, etc. Different data in heterogeneous data sources often have characteristics such as uneven quality and semantic misalignment, which are difficult to support with traditional query methods for a single data source.

[0066] Knowledge graph: A knowledge graph is a semantic network that reveals the relationship between entities. It contains several knowledge tuples, each of which includes a head entity, a tail entity, and an entity reference relationship between the head entity and the tail entity.

[0067] Figure 1 is a flowchart of a conflict elimination query method for multi-source heterogeneous data according to an exemplary embodiment. Figure 1 As shown, the method is applied in a terminal and may include the following steps:

[0068] Step S11: convert each data set in the to-be-queried heterogeneous data source into an equivalent knowledge graph, recorded as a data graph, and convert the user input query into an equivalent knowledge graph in the same way, recorded as a query graph, wherein each knowledge graph includes a number of knowledge tuples, each of the knowledge tuples includes two entities, and there are a number of to-be-queried entities in the query graph, representing the user's target query result;

[0069] Step S12: converting the query graph input by the user and the data graph converted from each data set in the heterogeneous data source into a line graph representation to obtain a corresponding query line graph and data line graph;

[0070] Step S13: According to the semantic information of each node in the query line graph and the network structure characteristics of the query line graph, all nodes that match the semantics and network structure of the query line graph nodes are found from all data line graphs;

[0071] Step S14: Based on the nodes that satisfy both semantic matching and network structure matching, entities in the nodes that match the entities to be queried in the query graph are obtained to form candidate query results;

[0072] Step S15: assigning initial data source credibility and initial data credibility to each data set in the heterogeneous data source and each data in the candidate query result;

[0073] Step S16: iteratively estimating data source credibility and data credibility according to the candidate query results and the historical query results of each data set until the two converge;

[0074] Step S17: According to the converged data credibility, select the entity with the highest credibility among the candidate query results as the query result output.

[0075] It can be seen from the above embodiments that the present application converts the query problem of multiple heterogeneous data sources into a knowledge graph matching problem, and introduces line graph conversion to accelerate the knowledge graph matching, thereby obtaining candidate query results that match the user query in each heterogeneous data source. Considering that each data source may provide conflicting data and unreliable data, a small batch data fusion algorithm that only considers the candidate query results is designed to eliminate unreliable conflicting data in the candidate query results, and finally obtain reliable and consistent data from multiple heterogeneous data sources as query results. On the one hand, the present application realizes trusted queries for multiple heterogeneous data sources, and on the other hand, it realizes more flexible and efficient data fusion.

[0076] In the specific implementation of step S11, each data set in the to-be-queried heterogeneous data source is converted into an equivalent knowledge graph, which is recorded as a data graph, and the user input query is converted into an equivalent knowledge graph in the same way, which is recorded as a query graph, wherein each knowledge graph includes a number of knowledge tuples, each of which includes two entities, and there are a number of to-be-queried entities in the query graph, which represent the user's target query result;

[0077] Specifically, each data set in the heterogeneous data source to be queried and the user input query are converted into corresponding equivalent knowledge graphs, the former is called a data graph, and the latter is called a query graph, wherein each of the knowledge graphs includes a number of knowledge tuples, each of the knowledge tuples includes a head entity, a tail entity, and an entity reference relationship between the head entity and the tail entity. In other words, the knowledge graph is composed of pieces of knowledge, each of which is represented by an SPO (Subject-Predicate-Object) triple, wherein the head entity and the tail entity are concrete things that exist objectively, usually referring to actual, functional organizations, institutions, objects, people, and other nouns. This method does not distinguish between the head entity and the tail entity, so they are collectively referred to as entities. In addition, the knowledge graph forms a graph structure based on the entity reference relationship between the entities. In addition, in particular, there are several entities to be queried in the query graph, representing the user's target query results;

[0078] In the specific implementation of step S12, the query graph and the data graph converted from each data set in the heterogeneous data source are converted into a line graph representation to obtain a corresponding query line graph and data line graph;

[0079] Specifically, a line graph corresponding to the query graph and the data graph is created. To distinguish them, the line graph corresponding to the query graph is named the query line graph, and the line graph corresponding to the data graph is named the data line graph. The created line graph is initially empty. The tuples in the query graph and the data graph are stored as nodes in the corresponding line graph. Consider two nodes in the line graph, that is, two tuples in the original graph. If the same entities exist in the two tuples, an edge is formed between the two nodes in the line graph. Otherwise, no edge can be formed between the two nodes. According to the above rules, the edges between the nodes in the query line graph and the edges between the nodes in the data line graph are connected respectively. For Figure 2 In one embodiment shown, any set of tuples in any knowledge graph (regardless of data graph and query graph) will be converted into a node in the corresponding line graph, such as the tuple will be transformed into the node u1 in the corresponding line graph. When two tuples share the same entity, an edge is formed between the nodes in their corresponding line graphs, such as the tuple and If the entity v1 is shared between them, then an edge is formed between the corresponding nodes u1 and u2 in the line graph. And so on. Figure 2 The displayed knowledge graph can be converted into a corresponding line graph representation.

[0080] In the specific implementation of step S13, according to the semantic information of each node in the query line graph and the network structure characteristics of the query line graph, all nodes that match the semantics and network structure of the query line graph nodes are found from all data line graphs, such as Figure 3 As shown, this step includes the following sub-steps:

[0081] Step S21: According to the semantic information of each node on the query line graph and the data line graph, the node semantic information is represented by using a pre-trained language model to form semantic representation vectors of each node on the query line graph and the data line graph respectively.

[0082] Specifically, Figure 4 In a specific embodiment shown in Table 1, the semantic information of a node in the query line graph and the data line graph is represented by the direct concatenation of its corresponding multi-tuple entity pairs and relationships.

[0083] Table 1 Semantic information of nodes in query line graph and data line graph

[0084]

[0085] As for Its semantic information is the corresponding multi-tuple In a specific embodiment, the pre-trained language model Sentence-BERT is used to represent the semantic information, and the semantic representation vector formed is expressed as For other nodes Can be formed into semantic representation vectors respectively

[0086] Step S22: Based on the semantic representation vectors of each node in the query line graph and the data line graph, the similarity between each node in the query line graph and all nodes in the data line graph is obtained, and all nodes in the data line graph whose similarity exceeds a similarity threshold are found.

[0087] Specifically, in a specific embodiment, the vector cosine similarity is used to calculate the similarity of the node semantic representation vector, such as Figure 4 As shown, for the query line graph node and nodes in the data line graph The calculated cosine similarity is 0.92. Set the similarity threshold to 0.9, then the nodes in the data line graph and If the threshold condition is met, it is regarded as a semantic matching node.

[0088] Step S23: According to the network structure characteristics of the query line graph, a subgraph structure matching the network structure of the query line graph is found from the data line graph using subgraph matching technology, and all nodes contained in the found matching subgraph are recorded.

[0089] Specifically, any subgraph matching technique is used to find a subgraph structure that matches the query line graph structure from the data line graph. In a specific embodiment, the VF2 subgraph matching algorithm is used to complete the matching subgraph search. Figure 4 In the specific embodiment shown, the subgraph structure formed by any two nodes in the data line graph matches the query line graph, so all nodes in the data line graph are recorded as structure matching nodes.

[0090] In the specific implementation of step S14, based on the nodes that simultaneously satisfy the semantic matching and the network structure matching, the entities in the nodes that match the entities to be queried in the query graph are obtained to form candidate query results;

[0091] Specifically, in Figure 4 In a specific embodiment shown in the figure, after step S13, nodes satisfying semantic matching are found from the data line graph. and And nodes that meet the network structure matching (The four nodes can form a "two-point-one-side" structure that is the same as the query line graph structure, so the four nodes all meet the network structure matching). The node that satisfies both semantic matching and network structure matching is and in, Contains the entity to be queried (query line graph node Included v ? ) So choose A candidate query result provided as a data source.

[0092] In the specific implementation of step S15, initial data source credibility and initial data credibility are respectively assigned to each data set in the heterogeneous data source and each data in the candidate query result;

[0093] Specifically, in one embodiment, the inverse of the missing value ratio contained in a data set in a heterogeneous data source is assigned to the data set as the initial data source credibility; and the data similarity obtained in step S13 is assigned to the corresponding candidate query result as the initial data credibility.

[0094] In the specific implementation of step S16, the data source credibility and data credibility are iteratively estimated according to the candidate query results and the historical query results of each data set until the two converge; Figure 5 As shown, this step may include the following sub-steps:

[0095] Step S31: Estimate and update the data source credibility based on the candidate query results, their corresponding data credibility and the data source historical query results.

[0096] Specifically, for a data set D in a data source, the data source credibility Pr(D) is calculated as follows:

[0097]

[0098] Where Data(Q,D) represents the partial query results provided by the dataset D in the candidate query results of the query Q, v represents a specific entity in the partial query results Data(Q,D), Pr(v) represents the data credibility of one of the query results, namely the entity v, and Pr(D|v) represents the probability that the dataset D is credible under the premise that v is credible. The calculation method of Pr(D|v) is as follows:

[0099]

[0100] Among them, Pr h (D) represents the historical value of the data source credibility of dataset D, H represents the number of query results provided by dataset D in historical queries, and Dv [Q] represents the entity subset in Data(Q,D) whose data credibility is greater than or equal to entity v, that is,

[0101] Step S32: Estimate and update the data credibility based on the candidate query results, the corresponding data credibility and the updated data source credibility.

[0102] Specifically, for an entity v in the candidate query result, its data credibility Pr(v) is calculated as follows:

[0103]

[0104] in, represents the set of all data sets in the heterogeneous data sources. The calculation of Pr(D|v) is as described in step S31. Pr(v|D) represents the probability that entity v is credible under the premise that data set D is credible. The calculation method of Pr(v|D) is as follows:

[0105]

[0106] Step S33: Iterate the data source credibility and data credibility calculation update process until the two converge.

[0107] Specifically, step S31 and step S32 are completed through multiple iterations until the data source credibility Pr(D) and the data credibility Pr(v) converge.

[0108] In the specific implementation of step S17, according to the converged data credibility, the entity with the highest credibility among the candidate query results is selected as the query result output;

[0109] Specifically, according to the converged data credibility calculated in step S16, the candidate query results are sorted, and the result with the highest data credibility among the candidate query results is output as the final query result.

[0110] Corresponding to the aforementioned embodiment of the multi-channel entity alignment method for large-scale knowledge graphs, the present application also provides an embodiment of a multi-channel entity alignment device for large-scale knowledge graphs.

[0111] Figure 6 1 is a block diagram of a conflict elimination query device for multi-source heterogeneous data according to an exemplary embodiment. Figure 6 , the device may include:

[0112] The acquisition module 21 converts each data set in the to-be-queried heterogeneous data source into an equivalent knowledge graph, which is recorded as a data graph, and converts the user input query into an equivalent knowledge graph in the same way, which is recorded as a query graph, wherein each knowledge graph includes a number of knowledge tuples, each of which includes two entities, and there are a number of to-be-queried entities in the query graph, which represent the user's target query result;

[0113] The conversion module 22 converts the query graph input by the user and the data graph converted from each data set in the heterogeneous data source into a line graph representation to obtain a corresponding query line graph and data line graph;

[0114] The matching module 23 finds all nodes that match the query line graph nodes semantically and in network structure from all data line graphs according to the semantic information of each node in the query line graph and the network structure characteristics of the query line graph;

[0115] The matching and merging module 24 obtains entities in the nodes that match the entities to be queried in the query graph based on the nodes that satisfy both semantic matching and network structure matching, thereby forming candidate query results;

[0116] The credibility initialization module 25 allocates initial data source credibility and initial data credibility to each data set in the heterogeneous data source and each data in the candidate query result;

[0117] The credibility iterative calculation module 26 iteratively estimates the data source credibility and data credibility according to the candidate query results and the historical query results of each data set until the two converge;

[0118] The output module 27 selects the entity with the highest credibility among the candidate query results as the query result to be output according to the converged data credibility.

[0119] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0120] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0121] Accordingly, the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the above-mentioned conflict elimination query method for multi-source heterogeneous data.

[0122] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned conflict elimination query method for multi-source heterogeneous data. Figure 7 As shown, a hardware structure diagram of a device having data processing capability, in which a conflict elimination query device for multi-source heterogeneous data provided by an embodiment of the present invention is located, except Figure 7 In addition to the processor, memory and network interface shown, any device with data processing capability in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capability, which will not be described in detail.

[0123] Accordingly, the present application also provides a computer-readable storage medium on which computer instructions are stored, and when the instructions are executed by the processor, the conflict elimination query method for multi-source heterogeneous data as described above is implemented. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card (Flash Card), etc. equipped on the device. Furthermore, the computer-readable storage medium can also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or is to be output.

[0124] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.

[0125] It should be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. A conflict elimination query method for multi-source heterogeneous data, characterized in that: include: Convert each data set in the to-be-queried heterogeneous data source into an equivalent knowledge graph, denoted as a data graph, and convert the user input query into an equivalent knowledge graph in the same way, denoted as a query graph, wherein each knowledge graph includes a number of knowledge tuples, each of which includes two entities, and there are a number of to-be-queried entities in the query graph, representing the user's target query result; Convert the query graph input by the user and the data graph converted from each data set in the heterogeneous data source into a line graph representation to obtain a corresponding query line graph and data line graph; According to the semantic information of each node in the query line graph and the network structure characteristics of the query line graph, all nodes that match the semantics and network structure of the query line graph nodes are found from all data line graphs; Based on the nodes that satisfy both semantic matching and network structure matching, the entities in the nodes that match the entities to be queried in the query graph are obtained to form candidate query results; Assign initial data source credibility and initial data credibility to each data set in the heterogeneous data source and each data in the candidate query result; According to the candidate query results and the historical query results of each data set, the data source credibility and data credibility are iteratively estimated until the two converge; According to the converged data credibility, the entity with the highest credibility among the candidate query results is selected as the query result output; The query graph input by the user and the data graph converted from each data set in the heterogeneous data source are converted into a line graph representation to obtain corresponding query line graphs and data line graphs, including: Create a one-to-one correspondence line graph for the query graph and the data graph, which correspond to the query line graph and the data line graph respectively. The created line graph is initially empty; Store the multi-tuples in the query graph and the data graph as nodes in the corresponding line graph; Consider two nodes in the line graph, which correspond to two tuples in the knowledge graph. If the same entity exists in the two tuples, an edge is formed between the two nodes in the line graph. Otherwise, no edge can be formed between the two nodes, thereby completing the edge between the nodes in the query line graph and the edge between the nodes in the data line graph.

2. The method according to claim 1, characterized in that According to the semantic information of each node in the query line graph and the network structure characteristics of the query line graph, all nodes that match the semantics and network structure of the query line graph nodes are found from all data line graphs, including: According to the semantic information of each node on the query line graph and the data line graph, the node semantic information is represented by using a pre-trained language model to form a semantic representation vector of each node on the query line graph and the data line graph respectively; According to the semantic representation vectors of each node in the query line graph and the data line graph, the similarity between each node in the query line graph and all nodes in the data line graph is obtained, and all nodes in the data line graph whose similarity exceeds a threshold are found; According to the network structure characteristics of the query line graph, the subgraph matching technology is used to find the subgraph structure that matches the network structure of the query line graph from the data line graph, and all nodes contained in the found matching subgraph are recorded.

3. The method according to claim 1, characterized in that According to the candidate query results and the historical query results of each data set in the data source, iteratively estimate the data source credibility and data credibility until convergence, including: Estimate and update the data source credibility based on the candidate query results and their corresponding data credibility and the historical query results of the data source; Estimate and update data credibility based on the candidate query results and their corresponding data credibility and the updated data source credibility; Iterate the data source credibility and data credibility calculation update process until the two converge.

4. The method according to claim 3, characterized in that For a data set D in a data source, the data source credibility Pr(D) is calculated as follows: Where Data(Q, D) represents the partial query results provided by the data set D in the candidate query results of the query Q, v represents a specific entity in the partial query result Data(Q, D), Pr(v) represents the data credibility of one of the query results, namely the entity v, and Pr(D|v) represents the probability that the data set D is credible under the premise that v is credible, where Pr(D|v) is calculated as follows: Among them, pr h (D) represents the historical value of the data source credibility of dataset D, H represents the number of query results provided by dataset D in historical queries, and D v [Q] represents the entity subset in Data(Q, D) whose data credibility is greater than or equal to entity v, that is, 5. The method according to claim 4, characterized in that For an entity v in the candidate query results, its data credibility Pr(v) is calculated as follows: in, represents the set of all data sets in the heterogeneous data sources. The calculation of Pr(D|v) is as described in step S31. Pr(v|D) represents the probability that entity v is credible under the premise that data set D is credible. The calculation method of Pr(v|D) is as follows:

6. A conflict elimination query device for multi-source heterogeneous data, characterized in that: include: An acquisition module is used to acquire a knowledge graph converted from each data set in a heterogeneous data source and a knowledge graph converted from a user input query, which are respectively recorded as a data graph and a query graph, wherein each knowledge graph includes a number of knowledge tuples, each of which includes two entities, and there are a number of entities to be queried in the query graph, representing the user's target query result; A conversion module, used to convert the query graph input by the user and the data graph converted from each data set in the heterogeneous data source into a line graph representation to obtain a corresponding query line graph and data line graph; A matching module is used to find all nodes that are semantically matched and network-structured with the query line graph nodes from all data line graphs according to the semantic information of each node in the query line graph and the network structure characteristics of the query line graph; A matching and merging module is used to obtain entities in the nodes that match the entities to be queried in the query graph based on nodes that satisfy both semantic matching and network structure matching, thereby forming candidate query results; A credibility initialization module is used to assign initial data source credibility and initial data credibility to each data set in the heterogeneous data source and each data in the candidate query result; The credibility iterative calculation module is used to iteratively estimate the credibility of the data source and the credibility of the data according to the candidate query results and the historical query results of each data set in the data source until the two converge; An output module is used to select the entity with the highest credibility among the candidate query results as the query result output according to the converged data credibility; The query graph input by the user and the data graph converted from each data set in the heterogeneous data source are converted into a line graph representation to obtain corresponding query line graphs and data line graphs, including: Create a one-to-one correspondence line graph for the query graph and the data graph, which correspond to the query line graph and the data line graph respectively. The created line graph is initially empty; Store the multi-tuples in the query graph and the data graph as nodes in the corresponding line graph; Consider two nodes in the line graph, which correspond to two tuples in the knowledge graph. If the same entity exists in the two tuples, an edge is formed between the two nodes in the line graph. Otherwise, no edge can be formed between the two nodes, thereby completing the edge between the nodes in the query line graph and the edge between the nodes in the data line graph.

7. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.

8. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

9. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Scientific and technological resource integration system based on multi-source database

    CN113312342A

  • Text matching method and device based on knowledge graph, equipment and storage medium

    CN113488165A