Knowledge graph alignment method and related device

By calculating the confidence of the paths in the knowledge graph and correcting the connection relationships, the problem of inconsistent quality in the fusion of knowledge graphs from different knowledge sources is solved, and higher quality knowledge graph fusion is achieved.

CN120632113APending Publication Date: 2025-09-12HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410284235.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-12
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When integrating multiple knowledge graphs, how to unify the knowledge quality of different knowledge sources and resolve the contradictions and knowledge conflicts that may arise in knowledge graphs from different sources.

Method used

By obtaining paths in different knowledge graphs, the confidence of the paths is calculated based on the relationship type of the connection relationship between entities, and the confidence of the connection relationship is corrected to achieve alignment of the relationships between entities and unify the quality.

Benefits of technology

The quality quantification performance of knowledge graphs after multi-source knowledge graph fusion is improved, ensuring the confidence consistency of the same or similar connection relationships in different knowledge graphs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632113A_ABST
    Figure CN120632113A_ABST
Patent Text Reader

Abstract

The invention relates to a knowledge graph alignment method, which can align the relationship confidence among entities and unify the knowledge quality of different knowledge sources. In the method, for a path between two same entities in different knowledge maps, based on a relationship type to which the connection relationship between the entities belongs, a correlation degree value between adjacent entities with the connection relationship on different paths is uniformly determined, and then the confidence coefficient of each path can be calculated. Therefore, on the basis of the confidence of different paths between the two same entities, confidence correction can be carried out on the connection relation between the entities on each path between the two entities, so that the relation confidence between the entities is aligned. According to the scheme, the association degree value between the entities is uniformly represented based on the relationship type to which the connection relationship between the adjacent entities belongs, so that the confidence degrees of different paths in the knowledge maps of different sources can be well evaluated from the same standard, and the quality quantification performance of the knowledge map after the multi-source knowledge maps are fused is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of knowledge graph technology, and in particular to a knowledge graph alignment method and related devices. Background Art

[0002] Knowledge graphs are a technology that graphically represents entities, relationships, and attributes. They organize entities, relationships, and attributes into nodes, edges, and labels, providing a structured, systematic, and standardized description of the world's knowledge.

[0003] With the rapid development of big data and artificial intelligence, knowledge graphs have garnered widespread attention and application in fields such as search engines, intelligent question-answering, recommendation systems, semantic analysis, medical diagnosis, and financial risk management. Knowledge graphs offer numerous advantages, including clear semantics, good interpretability, traceability, and strong domain expertise. These advantages enable humans to better understand and utilize the knowledge of the universe.

[0004] In practical applications, the diversity of knowledge sources and the significant differences in knowledge quality can lead to contradictory knowledge or even knowledge conflicts within knowledge graphs from different sources. Therefore, when integrating knowledge graphs from multiple sources, how to unify the knowledge quality of these different sources is an urgent issue to be addressed. Summary of the Invention

[0005] This application provides a knowledge graph alignment method that can align the relationship confidence between entities and unify the knowledge quality of different knowledge sources.

[0006] In a first aspect, the present application provides a method for aligning a knowledge graph, which is applied to quality alignment of multiple knowledge graphs from different sources. Specifically, the method for aligning a knowledge graph comprises: first, obtaining a first path in a first knowledge graph and a second path in a second knowledge graph, wherein the first path and the second path are both composed of a plurality of entities connected in sequence. The first path and the second path are different paths between a first entity and a second entity, respectively, that is, the first path and the second path are different paths between the same two entities. Moreover, the connection relationship between the entities on the first path and the second path has an initial confidence, which is used to indicate the degree of credibility of the connection relationship.

[0007] Then, a confidence level of the first path and a confidence level of the second path are determined.

[0008] Next, based on the confidence of the first path and the confidence of the second path, the initial confidence of the connection relationship between the entities on the first path and the second path is revised to obtain a revised confidence of the connection relationship between the entities on the first path and the second path. Since the first path and the second path are paths between the same two entities, after obtaining the confidence of these two paths, the confidence of the two paths can be combined to measure the reliability of the relationship between the two entities, thereby revising the confidence of the connection relationship between adjacent entities on the first path and the second path respectively.

[0009] Finally, after correcting the confidence of the connection relationship on the first path and the second path, the first knowledge graph and the second knowledge graph are fused to obtain a fused knowledge graph. The fused knowledge graph includes the first path and the second path, and in the fused knowledge graph, the confidence of the connection relationship between entities on the first path and the second path is the corrected confidence.

[0010] In this solution, for paths between the same two entities in different knowledge graphs, the confidence of each path is calculated based on the initial confidence of the connection relationship between the entities based on a unified standard. In this way, based on the confidence of different paths between the same two entities, the confidence of the connection relationship between the entities on each path between the two entities can be corrected, thereby aligning the confidence of the relationship between the entities, achieving the unification of the knowledge quality of different knowledge sources, and ultimately constraining the confidence of the same or similar connection relationships on different knowledge graphs to have a high degree of consistency, thereby improving the quality quantification performance of the knowledge graph after the fusion of multi-source knowledge graphs.

[0011] In one possible implementation, when determining the confidence of a first path and the confidence of a second path, the degree of association between the entities with a connection relationship in the first and second paths can be determined based on the relationship type to which the connection relationship between the entities in the first and second paths belongs. In the first and second paths, a unique corresponding relationship type can be found for the connection relationship between any two entities, i.e., the relationship type to which the connection relationship between the entities belongs is determined. Furthermore, for different relationship types, a corresponding degree of association value can be assigned to each relationship type, thereby representing the association between the entities under each relationship type. Therefore, based on the relationship type to which the connection relationship between the entities in the path belongs, the degree of association between the entities with a connection relationship in the path can be determined.

[0012] Next, the confidence of the first path and the confidence of the second path are determined based on the association degree values ​​and the initial confidences of the connection relationships between the entities. For example, the confidence of the entire path can be obtained by taking the weighted sum of the initial confidences of all connection relationships on the path, using the association degree values ​​of the connection relationships as weights.

[0013] In this solution, for the paths between the same two entities in different knowledge graphs, based on the relationship type to which the connection relationship between the entities belongs, the association degree values ​​between adjacent entities with connection relationships on different paths are uniformly determined, and then the confidence of each path can be calculated. In this way, based on the confidence of different paths between the same two entities, the confidence of the connection relationship between the entities on each path between the two entities can be corrected, thereby aligning the confidence of the relationship between the entities and achieving the uniform knowledge quality of different knowledge sources. Since this solution uniformly represents the association degree value between entities based on the relationship type to which the connection relationship between adjacent entities belongs, it can well evaluate the confidence of different paths in knowledge graphs from different sources from the same standard, and ultimately constrain the confidence of the same or similar connection relationships on different knowledge graphs to have a high consistency, thereby improving the quality quantification performance of the knowledge graph after the fusion of multi-source knowledge graphs.

[0014] In one possible implementation, the confidence of the first path can be obtained by summing, averaging, or weighted summing the initial confidences of each connection relationship on the first path; the confidence of the second path can also be obtained by summing, averaging, or weighted summing the initial confidences of each connection relationship on the second path. However, the method for calculating the confidence of the first path and the method for calculating the confidence of the second path must be consistent to ensure that the path confidence can be calculated based on a unified standard.

[0015] In one possible implementation, before determining the degree of association between entities having a connection relationship in the first path and the second path, a sorting result corresponding to a plurality of relationship types may be obtained first, and the sorting result is used to indicate the result after sorting the plurality of relationship types according to the degree of association corresponding to the relationship type. The plurality of relationship types include the relationship types to which the connection relationship between entities in the first knowledge graph and the second knowledge graph belongs. For example, assuming that the plurality of relationship types include family relationships, workplace relationships, and social relationships, then these plurality of relationship types are sorted in descending order according to the association procedure as follows: family relationships → social relationships → workplace relationships.

[0016] In this way, based on the sorting results, a degree of association value for each of the multiple relationship types can be determined. The degree of association value for each relationship type is used to determine the degree of association value between entities connected in the first path and the second path. For example, if the sorting results are sorted from high to low by degree of association, the degree of association value for each relationship type in the sorting results can be determined one by one in descending order of the degree of association value.

[0017] In this solution, by sorting the various relationship types appearing in different knowledge graphs according to the relative size of the association degree between the relationship types, and then determining the association degree value of each relationship type based on the sorting results, it is possible to determine the corresponding association degree value for each relationship type under the same standard, thereby ensuring the accuracy of the subsequent determination of the association degree value of the connection relationship in different knowledge graphs.

[0018] In one possible implementation, multiple relationship types may be input into a large language model to obtain a ranking result output by the large language model, wherein the large language model is used to rank different relationship types according to their corresponding association levels.

[0019] This solution, by combining the multi-domain knowledge and powerful semantic understanding capabilities of a large language model, can automatically sort various relationship types by their degree of association, improving efficiency and facilitating the subsequent rapid unification of knowledge quality across different knowledge sources. Furthermore, because the large language model's parameters incorporate extensive common sense, domain knowledge, and semantic understanding, it can provide a more objective relative ranking of the degree of association between relationship types.

[0020] In a possible implementation, in the first path, the association degree values ​​between entities having a connection relationship are used as weights, and the initial confidences of the connection relationships between the entities are weighted and summed to obtain the confidence of the first path.

[0021] In the second path, the association degree values ​​between entities with connection relationships are used as weights, and the initial confidences of the connection relationships between the entities are weighted and summed to obtain the confidence of the second path.

[0022] In this scheme, for the connection relationship on the path, the degree of association of different relationship types to which the connection relationship belongs is taken into account, and the degree of association of the connection relationship is used as a weight to combine the initial confidence of the connection relationship to calculate the path confidence, so that the calculated path confidence can more accurately represent the reliability of the relationship between the aligned entities.

[0023] In a possible implementation, in the process of correcting the confidence of the connection relationship between entities on the first path and the second path, a fusion confidence may be first determined based on the confidence of the first path and the confidence of the second path.

[0024] Then, based on the fused confidence, the confidence of the connection relationship between the entities on the first path and the second path is respectively revised. That is, the fused confidence is used as the actual confidence of the first path and the second path, and the confidence of the connection relationship between the entities on the first path and the second path is re-revised.

[0025] In this solution, after the confidence of the first path and the confidence of the second path are fused, the confidence of the connection relationship between entities on each path is corrected based on the fused confidence, which can achieve the fusion and weighing of the reliability of knowledge in different knowledge graphs, and thus facilitate the correction of the confidence of the connection relationship between entities in different knowledge graphs.

[0026] In a possible implementation, there are multiple methods for determining the fusion confidence based on the confidence of the first path and the confidence of the second path.

[0027] For example, between the confidence levels of the first and second paths, the confidence level of the target path is selected as the fused confidence level, where the target path is the path with the shortest or longest path length between the first and second paths. That is, when fusing the confidence levels of multiple paths, the confidence level of the shortest path can be considered the most reliable. Therefore, the confidence level of the shortest path is used as the fused confidence level, facilitating subsequent correction of the confidence levels of the connection relationships on each path based on the fused confidence level. Alternatively, the confidence level of the longest path can be used as the fused confidence level, facilitating subsequent correction of the confidence levels of the connection relationships on each path based on the unified fused confidence level.

[0028] Alternatively, the average of the confidence of the first path and the confidence of the second path is used as the fused confidence. That is, when fusing the confidences of multiple paths, the confidences of the multiple paths can be averaged and the obtained average is used as the fused confidence.

[0029] In general, the main purpose of fusing the confidence of different paths is to fuse and weigh the reliability of knowledge in different knowledge graphs, so that different paths can realize the correction of subsequent connection relationships based on the same fusion confidence, which is conducive to the subsequent constraints of the same or similar connection relationships in different knowledge graphs. The confidence has a high consistency, thereby improving the uniformity of knowledge graph quality evaluation.

[0030] In one possible implementation, the first path, the second path, the confidence of the first path, and the confidence of the second path can be input into a large language model to obtain a fused confidence. The large language model is used to determine the confidence of the existence of a relationship between two entities based on the different paths between the two entities.

[0031] In this solution, based on the multi-domain knowledge and powerful semantic understanding capabilities of the large language model, the connection relationships on different paths between two entities are automatically analyzed and aligned, thereby obtaining the confidence level of the existence of a relationship between the two entities, and using this confidence level as the fusion confidence level of multiple paths, which improves the accuracy and efficiency of obtaining the fusion confidence level and is conducive to the subsequent realization of the confidence level of the connection relationship on the correction path.

[0032] In one possible implementation, in the process of obtaining the fusion confidence based on the large language model, the first path, the second path, the confidence of the first path, and the confidence of the second path can be first input into the large language model to obtain the first probability and the second probability output by the large language model. The first probability indicates the probability that a relationship exists between the first entity and the second entity, and the second probability indicates the probability that no relationship exists between the first entity and the second entity.

[0033] Then, based on the first probability and the second probability, a fusion confidence is determined. For example, the first probability and the second probability are normalized, and then the normalized first probability is used as the fusion confidence.

[0034] In one possible implementation, in the process of respectively correcting the confidence of the connection relationship between entities on the first path and the second path, the confidence of the connection relationship between entities on the first path can be corrected based on the fusion confidence and the association degree value between entities with the connection relationship on the first path; and the confidence of the connection relationship between entities on the second path can be corrected based on the fusion confidence and the association degree value between entities with the connection relationship on the second path.

[0035] In this scheme, the confidence of each connection relationship on the path is corrected separately by combining the fusion confidence of the path and the correlation degree value of the connection relationship on the path. Starting from the relationship type to which the connection relationship belongs, it is possible to correct the confidence of the connection relationship on different knowledge graphs based on a unified standard, thereby ensuring that the quality of different knowledge graphs can be measured by a unified standard, which is conducive to the subsequent fusion of knowledge graphs.

[0036] The second aspect of the present application provides a knowledge graph alignment device, including: an acquisition module, used to acquire a first path in a first knowledge graph and a second path in a second knowledge graph, the first path and the second path are both composed of a plurality of entities connected in sequence, the first path and the second path are different paths between the first entity and the second entity, respectively, and the connection relationship between the entities on the first path and the second path has an initial confidence; a processing module, used to determine the confidence of the first path and the confidence of the second path; the processing module is also used to correct the initial confidence based on the confidence of the first path and the confidence of the second path to obtain a corrected confidence of the connection relationship between the entities on the first path and the second path; the processing module is also used to fuse the first knowledge graph and the second knowledge graph to obtain a fused knowledge graph, the fused knowledge graph includes the first path and the second path, and in the fused knowledge graph, the confidence of the connection relationship between the entities on the first path and the second path is a corrected confidence.

[0037] In one possible implementation, the processing module is further used to: determine the degree of association value between entities having a connection relationship in the first path and the second path based on the relationship type to which the connection relationship between the entities in the first path and the second path belongs; and determine the confidence of the first path and the confidence of the second path based on the degree of association value and the initial confidence of the connection relationship between the entities.

[0038] In one possible implementation, the acquisition module is further used to obtain sorting results corresponding to multiple relationship types, and the sorting results are used to indicate the results after sorting the multiple relationship types according to the degree of association corresponding to the relationship types. The multiple relationship types include the relationship types to which the connection relationships between entities in the first knowledge graph and the second knowledge graph belong; the processing module is also used to determine the degree of association value of each relationship type in the multiple relationship types based on the sorting results, and the degree of association value of each relationship type is used to determine the degree of association value between entities with connection relationships in the first path and the second path.

[0039] In one possible implementation, the acquisition module is further used to: input multiple relationship types into the large language model to obtain a ranking result output by the large language model; wherein the large language model is used to sort different relationship types according to the degree of association corresponding to the relationship types.

[0040] In one possible implementation, the processing module is further used to: in the first path, use the association degree values ​​between entities with a connection relationship as weights, perform weighted summation on the initial confidences of the connection relationships between the entities, and obtain the confidence of the first path; in the second path, use the association degree values ​​between entities with a connection relationship as weights, perform weighted summation on the initial confidences of the connection relationships between the entities, and obtain the confidence of the second path.

[0041] In a possible implementation, the processing module is further configured to: determine a fusion confidence based on the confidence of the first path and the confidence of the second path; and respectively correct the confidence of the connection relationship between entities on the first path and the second path based on the fusion confidence.

[0042] In one possible implementation, the processing module is further used to: select the confidence of the target path as the fusion confidence among the confidence of the first path and the confidence of the second path, where the target path is the path with the shortest or longest path length among the first path and the second path; or, take the average of the confidence of the first path and the confidence of the second path as the fusion confidence.

[0043] In one possible implementation, the processing module is further used to: input the first path, the second path, the confidence of the first path, and the confidence of the second path into a large language model to obtain a fused confidence, and the large language model is used to determine the confidence of the existence of a relationship between the two entities based on the different paths between the two entities.

[0044] In one possible implementation, the processing module is further used to: input the first path, the second path, the confidence of the first path, and the confidence of the second path into the large language model to obtain a first probability and a second probability output by the large language model, where the first probability is used to indicate the probability that there is a relationship between the first entity and the second entity, and the second probability is used to indicate the probability that there is no relationship between the first entity and the second entity; and determine the fusion confidence based on the first probability and the second probability.

[0045] In one possible implementation, the processing module is further used to: correct the confidence of the connection relationship between entities on the first path based on the fusion confidence and the association degree value between entities with connection relationship on the first path; correct the confidence of the connection relationship between entities on the second path based on the fusion confidence and the association degree value between entities with connection relationship on the second path.

[0046] The third aspect of the present application provides a knowledge graph alignment device, comprising: a processor and a memory; the memory is used to store computer instructions, and when the processor executes the instructions, the knowledge graph alignment device performs any of the methods described above.

[0047] In a fourth aspect, the present application provides a computer-readable storage medium having instructions stored therein. When the instructions are executed on a computer, the computer can execute any of the above methods.

[0048] A fifth aspect of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the methods described above.

[0049] In a sixth aspect, the present application provides a chip comprising a processor and a communication interface, wherein the communication interface is used to communicate with modules outside the chip, and the processor is used to run computer programs or instructions so that a device in which the chip is installed can execute any of the methods described above.

[0050] Among them, the technical effects brought about by any design method in the second to sixth aspects can refer to the technical effects brought about by different implementation methods in the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A schematic diagram of a knowledge graph provided in an embodiment of the present application;

[0052] Figure 2 A schematic diagram of a system architecture 100 provided in an embodiment of the present application;

[0053] Figure 3 A schematic diagram of a flow chart of a knowledge graph alignment method provided in an embodiment of the present application;

[0054] Figure 4 A schematic diagram of entity alignment between knowledge graphs provided in an embodiment of the present application;

[0055] Figure 5 A schematic diagram of a path in a knowledge graph provided in an embodiment of the present application;

[0056] Figure 6 A schematic diagram of determining the confidence of a path in a knowledge graph provided in an embodiment of the present application;

[0057] Figure 7 A schematic diagram of a knowledge graph fusion provided in an embodiment of the present application;

[0058] Figure 8 A schematic diagram of determining fusion confidence based on a large language model provided in an embodiment of the present application;

[0059] Figure 9 A schematic diagram of another method for calculating path confidence on a knowledge graph provided in an embodiment of the present application;

[0060] Figure 10 A schematic diagram of a process for realizing knowledge graph alignment provided in an embodiment of the present application;

[0061] Figure 11 A schematic diagram of the structure of a knowledge graph alignment device provided in an embodiment of the present application;

[0062] Figure 12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0063] Figure 13 A schematic diagram of the structure of a chip provided in an embodiment of the present application;

[0064] Figure 14 A schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of this application, rather than all embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0066] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchangeable where appropriate so that the embodiments can be implemented in a sequence other than that illustrated or described in this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. In actual application, there may be other division methods. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. Moreover, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed into multiple circuit units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this application.

[0067] To facilitate understanding, some technical terms involved in the embodiments of this application are first introduced below.

[0068] (1) Knowledge Graph

[0069] Knowledge graph is a technology that represents entities, relationships, and attributes in a graphical manner. Essentially, a knowledge graph is a structured semantic knowledge base that is used to quickly describe concepts and their relationships in the physical world.

[0070] Generally speaking, a knowledge graph is a graph-based data structure consisting of nodes and edges. Each node in a knowledge graph represents an entity, and each edge represents a relationship between entities. Essentially, a knowledge graph is a semantic network.

[0071] For example, see Figure 1 , Figure 1 This is a schematic diagram of a knowledge graph provided in an embodiment of the present application. Figure 1 As shown in the figure, the nodes in the knowledge graph represent different entities, such as Person A, Person B, Person C, Person D, Company A, and Company B. In addition, some nodes are connected by edges, which are used to represent the relationship between two entities. For example, the relationship between Person A and Task B is friends, and the relationship between Person C and Task D is colleagues.

[0072] (2) Entity

[0073] Entities in a knowledge graph refer to things or concepts in the real world, such as people, places, organizations, products, etc. Each entity has a unique identifier and some attributes. For example, a person entity can have attributes such as name, gender, and age.

[0074] (3) Attributes

[0075] In a knowledge graph, an attribute refers to a characteristic or description of an entity, such as the age or gender of a person entity. An attribute typically has a name and a value. For example, the name of the gender attribute is "gender," and its value can be "male" or "female."

[0076] (4) Relationship

[0077] In a knowledge graph, a relationship is a connection or link between entities, such as the work relationship between a person entity and an organization entity. A relationship typically has a name and two entities. For example, the name of the work relationship is "work," and it has a person entity and an organization entity.

[0078] It should be noted that, in the following text of this embodiment, the relationship between entities is also referred to as a connection relationship.

[0079] (5) Path

[0080] In a knowledge graph, a path is a series of entities. That is, starting from one entity in a knowledge graph and ending at another entity, these two entities and all other entities and relationships between them constitute a path.

[0081] (6) Large language model (LLM)

[0082] Large language models are deep learning models trained using large amounts of text data. They can generate natural language text or understand the meaning of text. Large language models can handle a variety of natural language tasks, such as text classification, question-answering, and conversation, and are an important path to artificial intelligence.

[0083] Specifically, large language models are a technology that has emerged in recent years. Because they undergo sophisticated data engineering and training processes, their parameters already incorporate a wealth of existing natural language processing knowledge. This knowledge can already replace humans in many language-related tasks, such as having large language models write code or perform text summarization.

[0084] With the rapid development of big data and artificial intelligence technologies, knowledge graph technology has received widespread attention and application in fields such as search engines, intelligent question-answering, recommendation systems, semantic analysis, medical diagnosis, and financial risk control. For example, in search engines, knowledge graphs can provide more accurate search results; in intelligent question-answering, knowledge graphs can provide more accurate answers; in recommendation systems, knowledge graphs can provide more personalized recommendations; in semantic analysis, knowledge graphs can provide deeper semantic understanding; in medical diagnosis, knowledge graphs can provide more accurate diagnostic results; and in financial risk control, knowledge graphs can provide more precise risk assessments.

[0085] However, the quality of knowledge graphs is limited by the information properties of the existing corpus, making it impossible to effectively judge the objectivity of solidified information. Furthermore, knowledge sources are diverse and their quality varies significantly, which can lead to contradictory knowledge or even knowledge conflicts in knowledge graphs from different sources.

[0086] Furthermore, strong or weak relationships exist between knowledge, and the strength of these relationships can affect the objective assessment of knowledge quality. This is especially true in multi-source knowledge graphs, where the same entities often have different relationships or paths connecting them that include different intermediate entities. Therefore, ensuring consistency in knowledge quality assessments of the same entities across multiple sources is a pressing issue.

[0087] Based on this, this embodiment provides a method for aligning knowledge graphs. For the paths between the same two entities in different knowledge graphs, based on the relationship type to which the connection relationship between the entities belongs, the degree of association between adjacent entities with connection relationships on different paths is uniformly determined, and then the confidence of each path can be calculated. In this way, based on the confidence of different paths between the same two entities, the confidence of the connection relationship between the entities on each path between the two entities can be corrected, thereby aligning the confidence of the relationship between the entities and achieving the uniform knowledge quality of different knowledge sources. Since this scheme uniformly represents the degree of association between entities based on the relationship type to which the connection relationship between adjacent entities belongs, it can well evaluate the confidence of different paths in knowledge graphs from different sources from the same standard, and ultimately constrain the confidence of the same or similar connection relationships on different knowledge graphs to have a high consistency, thereby improving the quality quantification performance of the knowledge graph after the fusion of multi-source knowledge graphs.

[0088] See also Figure 2 , Figure 2 A schematic diagram of a system architecture 100 provided in an embodiment of the present application. Figure 2 As shown, in the system architecture 100, the execution device 110 can be implemented by at least one computing instance of a physical host (computing device), a virtual machine, or a container. When the execution device 110 is implemented by a virtual machine or a container, the execution device 110 actually exists in the form of a cloud computing product and can provide cloud services.

[0089] Furthermore, the execution device 110 may be implemented by multiple computing instances of the same type. For example, the execution device 110 is implemented by multiple physical hosts, or by multiple virtual machines, or by multiple containers. It should be noted that multiple computing instances may be distributed in the same region or in different regions. Furthermore, multiple computing instances may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0090] Similarly, multiple compute instances can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0091] Optionally, the execution device 110 cooperates with other computing devices, such as data storage, routers, load balancers and other devices; the execution device 110 can be deployed on one physical site or distributed on multiple physical sites.

[0092] Optionally, in order to store data persistently, the system architecture 100 is further provided with a data storage system 120, which may be located outside the execution device 110 (e.g., Figure 2 As shown), the execution device 110 exchanges data with the execution device 110 through the network. Optionally, when the execution device 110 is a physical host, the data storage system 120 may also be located inside the execution device 110, such as when the data storage system 120 exchanges data with the processor through a bus. In this case, the data storage system 120 is represented by a hard disk. In the case of having the data storage system 120, the execution device 110 may use the data in the data storage system 120, or call the program code in the data storage system 120 to implement the knowledge graph alignment method provided in the embodiment of the present application.

[0093] Users can operate their respective user devices (such as local device 101 and local device 102) to interact with execution device 110. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a laptop computer, and a smart car.

[0094] Each user's local device can interact with the execution device 110 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0095] In one implementation, the execution device 110 is used to implement the knowledge graph alignment method provided in the embodiments of the present application, achieve alignment between knowledge graphs, and thus fuse a knowledge graph of uniform quality. Furthermore, when the local device 101 and the local device 102 need to use the knowledge graph to acquire knowledge, the execution device 110 processes the user's request based on the fused knowledge graph and then returns the corresponding processing results to the local device 101 and the local device 102.

[0096] In another implementation, the execution device 110 is used to implement the knowledge graph alignment method provided in the embodiment of the present application, and after fusing to obtain a knowledge graph of uniform quality, the obtained fused knowledge graph is sent to the local device 101 and the local device 102. In this way, the local device 101 and the local device 102 can locally deploy the fused knowledge graph to achieve rapid knowledge acquisition.

[0097] In another implementation, one or more aspects of the execution device 110 can be implemented by each local device. For example, the local device 101 can provide local data or feedback calculation results to the execution device 110, or execute the knowledge graph alignment method provided in the embodiment of the present application.

[0098] In general, the knowledge graph alignment method provided in the embodiments of the present application can be applied to electronic devices, such as the above-mentioned execution device 110, local device 101 or local device 102.

[0099] See also Figure 3 , Figure 3 A schematic diagram of a knowledge graph alignment method provided in an embodiment of the present application. Figure 3 As shown, the knowledge graph alignment method provided in the embodiment of the present application includes the following steps 301-305.

[0100] Step 301: Obtain a first path in the first knowledge graph and a second path in the second knowledge graph. The first path and the second path are both composed of multiple entities connected in sequence. The first path and the second path are different paths between the first entity and the second entity, respectively, and the connection relationship between the entities on the first path and the second path has an initial confidence.

[0101] In this embodiment, the first knowledge graph and the second knowledge graph can be two different knowledge graphs. For example, the sources of the first knowledge graph and the second knowledge graph are different, and the content recorded in the first knowledge graph and the content recorded in the second knowledge graph may be partially the same and partially different. For example, the first knowledge graph and the second knowledge graph both record two identical entities, but the other entities connected between the two identical entities in this knowledge graph are different. However, the first knowledge graph and the second knowledge graph belong to the same field or a similar field, and the first knowledge graph and the second knowledge graph can be integrated. For example, the first knowledge graph and the second knowledge graph are both knowledge graphs in the field of financial risk control; or, the first knowledge graph and the second knowledge graph are both knowledge graphs in the field of medical diagnosis.

[0102] For the first knowledge graph and the second knowledge graph, the entities in the first knowledge graph and the second knowledge graph can be aligned first to find entities that represent the same meaning in the first knowledge graph and the second knowledge graph. Specifically, different names may be used to name entities that represent the same meaning on different knowledge graphs. Therefore, the same entities on different knowledge graphs can be determined by aligning the entities. For example, on one knowledge graph, the name of an entity is table tennis; on another knowledge graph, the name of an entity is table tennis; but in fact, the entity named table tennis and the entity named table tennis represent the same meaning, so the two entities can be considered to be the same entity.

[0103] Specifically, in the process of entity alignment, the embedding vector of each entity in multiple knowledge graphs (such as the first knowledge graph and the second knowledge graph mentioned above) can be extracted by a feature extraction model, and then the similarity between the embedding vectors of entities in different knowledge graphs is calculated, such as calculating the Euclidean distance or cosine distance between the embedding vectors of the entities to represent the similarity between the embedding vectors. If the similarity between the embedding vectors of two entities in different knowledge graphs is greater than a certain threshold, the two entities can be considered to be the same entity, thereby achieving entity alignment. Among them, when the feature extraction model is used to extract the embedding vectors of entities in the knowledge graph, a corresponding embedding vector can be generated for each entity, and the embedding vector of each entity is represented in the form of a matrix, representing the feature representation of the entity.

[0104] For example, see Figure 4 , Figure 4 A schematic diagram of entity alignment between knowledge graphs provided in an embodiment of the present application. Figure 4 As shown, the first knowledge graph includes entity A, entity B, entity C, entity D and entity E; the second knowledge graph includes entity A', entity F, entity G and entity E'. By inputting the first knowledge graph and the second knowledge graph into the feature extraction model respectively, the embedding vector of each entity in the first knowledge graph and the second knowledge graph can be extracted respectively. Then, by calculating the similarity between the embedding vector of the entity in the first knowledge graph and the embedding vector of the entity in the second knowledge graph one by one, the entities with alignment relationship in the first knowledge graph and the second knowledge graph can be determined, that is, entities representing the same meaning. For example, in Figure 4In the example, if the similarity between the embedding vector of entity A in the first knowledge graph and the embedding vector of entity A' in the second knowledge graph is higher than a preset threshold, it is considered that entity A in the first knowledge graph and entity A' in the second knowledge graph have an alignment relationship, that is, entity A and entity A' are the same entity. Similarly, if the similarity between the embedding vector of entity E in the first knowledge graph and the embedding vector of entity E' in the second knowledge graph is higher than a preset threshold, entity E and entity E' are considered to be the same entity.

[0105] After aligning the entities in the first knowledge graph and the second knowledge graph, paths with the same starting point and end point can be found in the first knowledge graph and the second knowledge graph, such as the first path and the second path mentioned above. Specifically, for the first path in the first knowledge graph, and the second path in the second knowledge graph, the first path and the second path have the same starting point and the same end point. That is, the first entity on the first path is the same as the first entity on the second path (for example, both are first entities), and the last entity on the first path is the same as the last entity on the second path (for example, both are second entities). Moreover, the first path and the second path are both composed of at least two entities connected in sequence.

[0106] For example, see Figure 5 , Figure 5 This is a schematic diagram of a path in a knowledge graph provided in an embodiment of the present application. Figure 5 As shown, entity A in the first knowledge graph is aligned with entity A' in the second knowledge graph, and entity E in the first knowledge graph is aligned with entity E' in the second knowledge graph. Therefore, the path between entity A and entity E in the first knowledge graph can be determined as the first path, and the path between entity A' and entity E' in the second knowledge graph can be determined as the second path. The first path is specifically: entity A → entity B → entity C → entity D → entity E. The second path is specifically: entity A' → entity G → entity E'.

[0107] In addition, on the first path and the second path, there will be a connection relationship between adjacent entities. For example, in the first knowledge graph, there is a connection relationship r1 between entity A and entity B, and there is a connection relationship r2 between entity B and entity C; in the second knowledge graph, there is a connection relationship r5 between entity A' and entity G. Moreover, the connection relationship between adjacent entities will have a corresponding initial confidence, which is used to indicate the degree of credibility of the connection relationship, that is, the quality of the connection relationship in the knowledge graph. Generally speaking, the initial confidence of the connection relationship can be determined based on the knowledge source of the knowledge graph during the process of constructing the knowledge graph or after the knowledge graph is constructed. This embodiment does not limit the method for determining the initial confidence of the connection relationship.

[0108] Step 302 : Determine the association degree value between the entities having the connection relationship in the first path and the second path based on the relationship type of the connection relationship between the entities in the first path and the second path.

[0109] In the first path and the second path, the connection relationship between any two entities can find the corresponding relationship type, that is, determine the relationship type to which the connection relationship between the entities belongs. Moreover, for different relationship types in the knowledge graph, a corresponding association degree value can be assigned to each relationship type, thereby representing the association between entities under each relationship type. If the association degree value of the relationship type to which the connection relationship between two entities belongs is higher, it means that the association between the two entities is stronger, that is, the possibility of the two entities being associated is higher; if the association degree value of the relationship type to which the connection relationship between two entities belongs is lower, it means that the association between the two entities is weaker, that is, the possibility of the two entities being associated is lower.

[0110] For example, if the connection between two entities is a father-son relationship, a mother-son relationship, a father-daughter relationship, or a mother-daughter relationship, then the relationship type of the connection between the two entities can be determined to be a family relationship; if the connection between the two entities is a colleague relationship or a superior-subordinate relationship, then the relationship type of the connection between the two entities can be determined to be a workplace relationship; if the connection between the two entities is a friend relationship, a classmate relationship, or a teacher-student relationship, then the relationship type of the connection between the two entities can be determined to be a social relationship. For each of the following relationship types: family, workplace, and social, a corresponding degree of association value can be assigned to represent the association between entities within these relationship types. Entities with family relationships have the strongest association, so this relationship type can be assigned the highest degree of association value, such as 0.7; entities with social relationships have average association, so this relationship type can be assigned a medium degree of association value, such as 0.4; and entities with workplace relationships have the weakest association, so this relationship type can be assigned the lowest degree of association value, such as 0.2.

[0111] For example, assuming that the connection relationships between entities belong to the following relationship types: causal relationships, inclusion relationships, inductive relationships, and parallel relationships, then each of these relationship types can be assigned a corresponding degree of association value to represent the association between entities under these relationship types. Among them, the association between entities with inclusion relationships is the strongest, so the inclusion relationship type can be assigned the highest degree of association value, such as 0.8; the causal relationship type can be assigned the second highest degree of association value, such as 0.6; the inductive relationship type can be assigned an even lower degree of association value, such as 0.5; and finally, the parallel relationship type can be assigned the lowest degree of association value, such as 0.2.

[0112] In general, based on the relationship type of the connection relationship between any two adjacent entities on the path, the association degree value between the adjacent entities on the path can be determined to indicate the level of association between these adjacent entities.

[0113] Optionally, in order to determine the association degree value of each relationship type, the sorting results corresponding to the multiple relationship types involved in the first knowledge graph and the second knowledge graph can be obtained first. The sorting results of the multiple relationship types are used to indicate the results after the multiple relationship types are sorted according to the association degrees corresponding to the relationship types. For example, the sorting results are obtained by sorting the relationship types from high to low according to the association degree. For example, assuming that the multiple relationship types include family relationships, workplace relationships, and social relationships, then these multiple relationship types are sorted from high to low according to the association procedure: family relationships → social relationships → workplace relationships.

[0114] It should be noted that the multiple relationship types include all relationship types that can be found in the connection relationships between entities in the first and second knowledge graphs. For example, the multiple relationship types are all possible relationship types that can appear in the domains to which the first and second knowledge graphs belong. Therefore, any connection relationship between entities in the first and second knowledge graphs can be found in the multiple relationship types mentioned above.

[0115] In this way, based on the sorting results of multiple relationship types, the association degree value of each relationship type in the multiple relationship types can be determined. For example, when the sorting results are sorted from high to low by association degree, the association degree value of each relationship type in the sorting results can be determined one by one in descending order of association degree value, thereby achieving a unique corresponding association degree value for each relationship type under the same standard.

[0116] In this solution, by sorting the various relationship types appearing in different knowledge graphs according to the relative size of the association degree between the relationship types, and then determining the association degree value of each relationship type based on the sorting results, it is possible to determine the corresponding association degree value for each relationship type under the same standard, thereby ensuring the accuracy of the subsequent determination of the association degree value of the connection relationship in different knowledge graphs.

[0117] Optionally, to facilitate rapid sorting of multiple relationship types, multiple relationship types can be input into a large language model to obtain sorting results output by the large language model, wherein the large language model is used to sort different relationship types according to their corresponding degrees of association.

[0118] It is understandable that large language models are usually trained based on a huge number of learnable parameters and a large-scale corpus prepared in advance. They can automatically mine knowledge from massive amounts of data and implicitly embed knowledge in model parameters, with powerful knowledge expression and semantic understanding capabilities. Therefore, by combining the multi-domain knowledge and powerful semantic understanding capabilities of the large language model, multiple relationship types can be automatically sorted according to the degree of association, improving the efficiency of sorting multiple relationship types and facilitating the subsequent rapid unification of the knowledge quality of different knowledge sources. Moreover, because the large language model contains large-scale common sense, domain knowledge, and semantic understanding capabilities in its parameters, it can perform a more objective relative ranking of the degree of association of relationship types.

[0119] Step 303 : Determine the confidence of the first path and the confidence of the second path based on the association degree value and the initial confidence of the connection relationship between the entities.

[0120] In this embodiment, since the connection relationship between adjacent entities in each knowledge graph has a corresponding initial confidence to indicate the credibility of the connection relationship between the entities, in this embodiment, the weighted confidence on the path can be calculated as the confidence of the path based on the association degree value of the connection relationship calculated in step 302.

[0121] Specifically, because both the first and second paths are composed of multiple entities, they contain one or more connections, each with a corresponding degree of association and initial confidence. The confidence of the entire path is obtained by taking the weighted sum of the initial confidences of all connections along the path, using the degree of association as the weight.

[0122] Exemplarily, in the first path, the association degree values ​​between entities having a connection relationship are used as weights, and the initial confidences of the connection relationships between the entities are weighted and summed to obtain the confidence of the first path.

[0123] In the second path, the association degree values ​​between entities with connection relationships are used as weights, and the initial confidences of the connection relationships between the entities are weighted and summed to obtain the confidence of the second path.

[0124] See Figure 6 , Figure 6 This is a schematic diagram of determining the confidence of a path in a knowledge graph provided in an embodiment of the present application. Figure 6 As shown, in the first path of the first knowledge graph, the connection relationship r1 between entities A and B has an initial confidence of q1 and an association value of a1; the connection relationship r2 between entities B and C has an initial confidence of q2 and an association value of a2; the connection relationship r3 between entities C and D has an initial confidence of q3 and an association value of a3; and the connection relationship r4 between entities D and E has an initial confidence of q4 and an association value of a4. The confidence L1 of the first path is specifically expressed as follows:

[0125]

[0126] In the second path of the second knowledge graph, the initial confidence level for the connection r5 between entity A' and entity G is q5, and the association level is a5. The initial confidence level for the connection r6 between entity G and entity E' is q6, and the association level is a6. The confidence level L2 of the second path is expressed as follows:

[0127]

[0128] That is to say, when calculating the confidence of a path, the confidence of the path is represented by a fraction, where the numerator of the fraction is the weighted sum of the initial confidences of each connection relationship on the path, and the weight corresponding to each connection relationship is the association degree value of the connection relationship; the denominator of the fraction is the sum of the association degree values ​​of each connection relationship.

[0129] In this scheme, for the connection relationship on the path, the degree of association of different relationship types to which the connection relationship belongs is taken into account, and the degree of association of the connection relationship is used as a weight to combine the initial confidence of the connection relationship to calculate the path confidence, so that the calculated path confidence can more accurately represent the reliability of the relationship between the aligned entities.

[0130] Step 304 : Based on the confidence of the first path and the confidence of the second path, the confidence of the connection relationship between the entities on the first path and the second path is modified.

[0131] In this embodiment, since the first path and the second path are both paths between the same two entities, after obtaining the confidence levels of the two paths, the confidence levels of the two paths can be combined to measure the reliability of the relationship between the two entities, thereby respectively correcting the confidence levels of the connection relationships between adjacent entities on the first path and the second path.

[0132] For example, after obtaining the confidence of the first path and the confidence of the second path, the fusion confidence can be determined based on the confidence of the first path and the confidence of the second path. The fusion of the confidence of the first path and the confidence of the second path is mainly to integrate and weigh the reliability of knowledge in different knowledge graphs, thereby facilitating the correction of the confidence of the connection relationship between entities in different knowledge graphs.

[0133] Then, based on the fused confidence, the confidence of the connection relationship between the entities on the first path and the second path is respectively revised. That is, the fused confidence is used as the actual confidence of the first path and the second path, and the confidence of the connection relationship between the entities on the first path and the second path is re-revised.

[0134] Specifically, the confidence level of the connection between entities on the first path can be modified based on the fused confidence level and the association level between entities with a connection relationship on the first path. Similarly, the confidence level of the connection between entities on the second path can be modified based on the fused confidence level and the association level between entities with a connection relationship on the second path.

[0135] For example, the fusion confidence is distributed according to the association strength value of each connection relationship on the path, so as to obtain the confidence of each connection relationship, so as to achieve the alignment of the confidence of the connection relationships on different knowledge graphs.

[0136] by Figure 6 Taking the first path and the second path shown as an example, assuming that the fusion confidence obtained by fusing the confidence L1 of the first path and the confidence L2 of the second path is L, when correcting the connection relationship between the entities on the first path, the corrected confidence of the connection relationship between entity A and entity B on the first path can be specifically shown as the following formula.

[0137]

[0138] Among them, q′1 is the corrected confidence of the connection relationship between entity A and entity B; a1 is the association degree value of the connection relationship between entity A and entity B; a2 is the association degree value of the connection relationship between entity B and entity C; a3 is the association degree value of the connection relationship between entity C and entity D; a4 is the association degree value of the connection relationship between entity D and entity E.

[0139] Similarly, the revised confidence of the connection relationship between entity B and entity C on the first path may be specifically expressed as the following formula.

[0140]

[0141] Where q′2 is the corrected confidence of the connection relationship between entity B and entity C.

[0142] In general, when calculating the corrected confidence of the connection relationship between any two adjacent entities on a path, the corrected confidence can be obtained based on a fraction. The numerator of the fraction is the product of the correlation value of the connection relationship between the two adjacent entities and the fusion confidence, and the denominator is the sum of the correlation values ​​of the connection relationships between all adjacent entities on the path. That is, the fusion confidence of the path is assigned to the connection relationship between the entities according to the correlation value of the connection relationship between the entities on the path, as the confidence of the connection relationship between the entities. In this way, after the confidence of the connection relationship between adjacent entities on the path has been corrected, the sum of the corrected confidence of the connection relationship between the entities on the path is the fusion confidence of the entire path.

[0143] In this scheme, the confidence of each connection relationship on the path is corrected separately by combining the fusion confidence of the path and the correlation degree value of the connection relationship on the path. Starting from the relationship type to which the connection relationship belongs, it is possible to correct the confidence of the connection relationship on different knowledge graphs based on a unified standard, thereby ensuring that the quality of different knowledge graphs can be measured by a unified standard, which is conducive to the subsequent fusion of knowledge graphs.

[0144] In addition to steps 303-304 described above, the confidence of the first path and the confidence of the second path may also be calculated in other ways. For example, the confidence of the first path may be obtained by summing, averaging, or weighted summing the initial confidences of the various connection relationships on the first path; the confidence of the second path may also be obtained by summing, averaging, or weighted summing the initial confidences of the various connection relationships on the second path. However, the method for calculating the confidence of the first path and the method for calculating the confidence of the second path need to be consistent to ensure that the calculation of the path confidence can be implemented based on a unified standard.

[0145] Step 305: Fuse the first knowledge graph and the second knowledge graph to obtain a fused knowledge graph. The fused knowledge graph includes the first path and the second path, and in the fused knowledge graph, the confidences of the connection relationships between entities on the first path and the second path are both corrected confidences.

[0146] After correcting the confidence levels of the connections along the paths in the first and second knowledge graphs, the first and second knowledge graphs can be fused, resulting in a unified fused knowledge graph with a uniform knowledge quality evaluation. It should be noted that in the fused knowledge graph, the paths between the same two entities on different knowledge graphs are fused, and the confidence levels of the connections between the entities along these paths are corrected.

[0147] For example, see Figure 7 , Figure 7 This is a schematic diagram of a knowledge graph fusion provided in an embodiment of the present application. Figure 7 As shown, after the first knowledge graph and the second knowledge graph are fused, a fused knowledge graph is obtained. In the fused knowledge graph, the paths that exist on the first knowledge graph and the second knowledge graph are also present on the fused knowledge graph. In addition, the confidence on the path between the same two entities on the first knowledge graph and the second knowledge graph is corrected on the fused knowledge graph. Specifically, on the fused knowledge graph, the first path "entity A→entity B→entity C→entity D→entity E" in the first knowledge graph is retained; the second path "entity A'→entity G→entity E'" on the second knowledge graph is fused with some entities on the first path (i.e., entity A and entity E), thereby changing to the path "entity A→entity G→entity E". In addition, the confidence of the connection relationship between adjacent entities on the path "entity A→entity B→entity C→entity D→entity E" and the path "entity A→entity G→entity E" are corrected. For example, the confidence of the connection relationship between entity A and entity B is q'1; the confidence of the connection relationship between entity B and entity C is q'2; the confidence of the connection relationship between entity C and entity D is q'3... The confidence of the connection relationship between entity G and entity E is q'6.

[0148] In addition, since the path from entity A to entity F only appears on the first knowledge graph, the path from entity A to entity F is retained on the fused knowledge graph, and the confidence of the connection relationship between entities on the path from entity A to entity F is not corrected.

[0149] It should be noted that the above embodiment introduces the process of correcting the confidence of the connection relationship between entities on the first knowledge graph and the second knowledge graph during the fusion of the two knowledge graphs. In practical applications, two or more knowledge graphs can be fused, and the confidence of the connection relationship on different paths between the same two entities in these knowledge graphs can be corrected. In addition, there can be one or more paths between the same entities on a knowledge graph.

[0150] In short, for multiple different paths between the same two entities, the fusion confidence of these paths can be calculated according to the above method, and then the connection relationship on each path can be corrected based on the fusion confidence, thereby achieving the correction alignment of the path confidence.

[0151] In the above step 304 , there are many ways to fuse the confidences of different paths.

[0152] In a possible implementation, the confidence of the target path is selected as the fusion confidence between the confidence of the first path and the confidence of the second path, where the target path is the path with the shortest or longest path length between the first path and the second path.

[0153] That is to say, when fusing the confidences of multiple paths, the confidence of the shortest path can be considered to be the most reliable. Therefore, the confidence of the shortest path is used as the fusion confidence, so that the confidence of the connection relationship on each path can be corrected based on the fusion confidence.

[0154] In another possible implementation, an average of the confidence of the first path and the confidence of the second path is used as the fusion confidence.

[0155] That is, when fusing the confidences of multiple paths, the confidences of the multiple paths may be averaged, and the obtained average value may be used as the fused confidence.

[0156] In general, the main purpose of fusing the confidence of different paths is to fuse and weigh the reliability of knowledge in different knowledge graphs, so that different paths can realize the correction of subsequent connection relationships based on the same fusion confidence, which is conducive to the subsequent constraints of the same or similar connection relationships in different knowledge graphs. The confidence has a high consistency, thereby improving the uniformity of knowledge graph quality evaluation.

[0157] In another possible implementation, the first path, the second path, the confidence of the first path, and the confidence of the second path can be input into a large language model to obtain a fusion confidence. The large language model is used to determine the confidence of the existence of a relationship between the two entities based on the different paths between the two entities.

[0158] That is to say, in this embodiment, the different paths between the two entities and the confidence of these paths that have been obtained are input into the large language model, and the large language model is used to determine the confidence of the relationship between the two entities, and then the obtained confidence is used as the fusion confidence.

[0159] For example, data such as entities, entity attributes, connection relationships, relationship types, and initial confidence scores on multiple paths are arranged into linear text. This is then combined with the confidence scores of each path as input to a large language model for inference. The prediction results with the integrated confidence scores are then extracted from the output of the large language model.

[0160] In this solution, based on the multi-domain knowledge and powerful semantic understanding capabilities of the large language model, the connection relationships on different paths between two entities are automatically analyzed and aligned, thereby obtaining the confidence level of the existence of a relationship between the two entities, and using this confidence level as the fusion confidence level of multiple paths, which improves the accuracy and efficiency of obtaining the fusion confidence level and is conducive to the subsequent realization of the confidence level of the connection relationship on the correction path.

[0161] Optionally, after inputting the first path, the second path, the confidence score of the first path, and the confidence score of the second path into the large language model, a first probability and a second probability are output by the large language model. The first probability indicates the probability that a relationship exists between the first entity and the second entity, and the second probability indicates the probability that no relationship exists between the first entity and the second entity.

[0162] In this way, based on the first probability and the second probability, a fusion confidence can be determined. For example, the first probability and the second probability are normalized, and then the normalized first probability is used as the fusion confidence.

[0163] For example, see Figure 8 , Figure 8 This is a schematic diagram of determining fusion confidence based on a large language model provided in an embodiment of the present application. Figure 8 As shown, the first path in the first knowledge graph (including each entity on the first path, the attributes and connection relationships of the entities, the initial confidence and correlation value of the connection relationship), the confidence of the first path, the second path, and the confidence of the second path are input into the large language model, which processes them and outputs the first probability and the second probability. Among them, the first probability represents the probability that there is indeed a relationship between entity A and entity E (that is, the probability that the result is "true"); the second probability represents the probability that there is no relationship between entity A and entity E (that is, the probability that the result is "false"). In this way, after normalizing the first probability and the second probability, the normalized first probability (that is, the probability that the result is "true") is taken as the fusion confidence.

[0164] Furthermore, in step 303 above, the process of determining the path confidence level is described, where the association degree values ​​of the connection relationships along the path are used as weights, combined with the initial confidence levels of each connection relationship along the path. In some possible embodiments, some entities may be identical on different paths, but the connection relationships between these entities may be different. In other words, for two adjacent entities, the connection relationship between the two entities may be different on different paths.

[0165] For example, see Figure 9 , Figure 9 This is another schematic diagram of calculating the confidence of a path on a knowledge graph provided in an embodiment of the present application. Figure 9 As shown in the figure, the first path on the first knowledge graph is: "Entity A → Entity B → Entity C → Entity D → Entity E"; the second path on the second knowledge graph is: "Entity A' → Entity B' → Entity E'". Entity B on the first knowledge graph and entity B' on the second knowledge graph are also aligned entities. However, in the first knowledge graph, the connection relationship between entity A and entity B is r1; while in the second knowledge graph, the connection relationship between entity A' and entity B' is r5. In other words, the connection relationship between the same two entities is different on different knowledge graphs.

[0166] In this case, when calculating the path confidence, for adjacent entities with different connection relationships on different knowledge graphs, different connection relationships can be used to calculate the corresponding path confidence separately, so that multiple path confidences can be calculated for the same path. Then, from the multiple path confidences corresponding to the same path, one path confidence is selected to perform the subsequent path confidence fusion process (for example, the largest path confidence is selected); or multiple path confidences corresponding to the same path are used to perform the subsequent path confidence fusion process.

[0167] like Figure 9 As shown, for the first path, when the path confidence is calculated with the connection relationship between entity A and entity B as r1, the initial confidence between entity A and entity B is q1, and the association degree value is a1. Therefore, the corresponding path confidence L11 of the first path can be expressed by the following formula.

[0168]

[0169] For the first path, when calculating the path confidence with the connection relationship between entity A and entity B as r5, the initial confidence between entity A and entity B is q5, and the association degree value is a5. Therefore, the corresponding path confidence L12 of the first path can be expressed by the following formula.

[0170]

[0171] Similarly, for the second path, when the path confidence is calculated with the connection relationship between entity A' and entity B' as r5, the initial confidence between entity A' and entity B' is q5, and the association degree value is a5. Therefore, the corresponding path confidence L21 of the second path can be expressed by the following formula.

[0172]

[0173] For the second path, when calculating the path confidence with the connection relationship between entity A' and entity B' as r1, the initial confidence between entity A' and entity B' is q1, and the association degree value is a1, so the corresponding path confidence L22 of the second path can be expressed by the following formula.

[0174]

[0175] In this way, the first path corresponds to two path confidences, L11 and L12, respectively; the second path corresponds to two path confidences, L21 and L22, respectively. When calculating the fusion confidence, the calculation of the fusion confidence can be performed by selecting one path confidence from the two path confidences of the first path (for example, selecting the largest path confidence), and selecting one path confidence from the two path confidences of the second path. Alternatively, both the path confidences of the first path and the two path confidences of the second path can be used to calculate the fusion confidence, which is not specifically limited here.

[0176] In general, the process of achieving knowledge graph alignment in this embodiment can be divided into multiple stages. Figure 10 , Figure 10 A schematic diagram of a process for realizing the alignment of knowledge graphs provided in an embodiment of the present application. Figure 10 As shown in Figure 2, the process of aligning different knowledge graphs includes the following stages.

[0177] S1, obtains knowledge graphs from multiple sources.

[0178] Among them, multiple knowledge graphs from different sources can be constructed based on different knowledge sources, and these knowledge graphs from different sources all belong to the same field or similar fields, which can achieve the fusion of knowledge graphs.

[0179] S2, performs entity alignment on multiple knowledge graphs.

[0180] Entity alignment between different knowledge graphs is achieved by extracting the embedding vectors of entities in each knowledge graph and comparing the similarity between the embedding vectors of entities in different knowledge graphs. In other words, entities representing the same meaning are found in different knowledge graphs.

[0181] S3, estimates the path confidence between aligned entities in each knowledge graph.

[0182] When the connection between entities in each knowledge graph has an initial confidence, the association value corresponding to each connection can be determined based on the relationship type of the connection between the entities. Then, based on the association value and the initial confidence of the connection, the confidence of the path between the aligned entities can be determined. In other words, the confidence of the path between the same two entities in different knowledge graphs is calculated.

[0183] S4, fuse the confidences of different paths to obtain the fused confidence.

[0184] For different paths between the same two entities, the path confidences of these different paths can be fused to obtain a fused confidence.

[0185] S5, corrects the confidence of the connection relationship between entities in the knowledge graph based on the fusion confidence.

[0186] Finally, based on the path fusion confidence, the confidence of the connection relationship between entities on each path can be corrected, thereby achieving quality alignment of different knowledge graphs, that is, using the same standard to measure the quality of the connection relationship between entities in each knowledge graph.

[0187] The above describes in detail the method provided by the embodiment of the present application. Next, the device provided by the embodiment of the present application for executing the above method will be introduced.

[0188] See also Figure 11 , Figure 11 This is a schematic diagram of the structure of a knowledge graph alignment device provided in an embodiment of the present application. Figure 11As shown, the knowledge graph alignment device provided by the embodiment of the present application includes: an acquisition module 1101, used to obtain a first path in the first knowledge graph and a second path in the second knowledge graph, the first path and the second path are both composed of a plurality of entities connected in sequence, the first path and the second path are different paths between the first entity and the second entity respectively, and the connection relationship between the entities on the first path and the second path has an initial confidence; a processing module 1102, used to determine the confidence of the first path and the confidence of the second path; the processing module 1102 is also used to correct the initial confidence based on the confidence of the first path and the confidence of the second path to obtain a corrected confidence of the connection relationship between the entities on the first path and the second path; the processing module 1102 is also used to fuse the first knowledge graph and the second knowledge graph to obtain a fused knowledge graph, the fused knowledge graph includes the first path and the second path, and in the fused knowledge graph, the confidence of the connection relationship between the entities on the first path and the second path is a corrected confidence.

[0189] In one possible implementation, processing module 1102 is further used to: determine the degree of association between entities having a connection relationship in the first path and the second path based on the relationship type to which the connection relationship between the entities in the first path and the second path belongs; and determine the confidence of the first path and the confidence of the second path based on the degree of association value and the initial confidence of the connection relationship between the entities. In one possible implementation, acquisition module 1101 is further used to obtain sorting results corresponding to multiple relationship types, the sorting results being used to indicate the results after sorting multiple relationship types according to the degree of association corresponding to the relationship types, the multiple relationship types including the relationship types to which the connection relationship between entities in the first knowledge graph and the second knowledge graph belongs; processing module 1102 is further used to determine the degree of association value of each relationship type among the multiple relationship types based on the sorting results, the degree of association value of each relationship type being used to determine the degree of association value between entities having a connection relationship in the first path and the second path.

[0190] In a possible implementation, the acquisition module 1101 is further used to: input multiple relationship types into the large language model to obtain a ranking result output by the large language model; wherein the large language model is used to sort different relationship types according to the degree of association corresponding to the relationship types.

[0191] In one possible implementation, the processing module 1102 is further used to: in the first path, use the association degree values ​​between entities with a connection relationship as weights, perform weighted summation on the initial confidences of the connection relationships between the entities, and obtain the confidence of the first path; in the second path, use the association degree values ​​between entities with a connection relationship as weights, perform weighted summation on the initial confidences of the connection relationships between the entities, and obtain the confidence of the second path.

[0192] In a possible implementation, the processing module 1102 is further configured to: determine a fusion confidence based on the confidence of the first path and the confidence of the second path; and respectively correct the confidence of the connection relationship between entities on the first path and the second path based on the fusion confidence.

[0193] In one possible implementation, the processing module 1102 is further used to: select the confidence of the target path as the fusion confidence from the confidence of the first path and the confidence of the second path, where the target path is the path with the shortest or longest path length between the first path and the second path; or, take the average of the confidence of the first path and the confidence of the second path as the fusion confidence.

[0194] In one possible implementation, the processing module 1102 is further used to: input the first path, the second path, the confidence of the first path, and the confidence of the second path into a large language model to obtain a fusion confidence, and the large language model is used to determine the confidence of the existence of a relationship between the two entities based on the different paths between the two entities.

[0195] In one possible implementation, the processing module 1102 is further used to: input the first path, the second path, the confidence of the first path, and the confidence of the second path into the large language model to obtain a first probability and a second probability output by the large language model, where the first probability is used to indicate the probability that there is a relationship between the first entity and the second entity, and the second probability is used to indicate the probability that there is no relationship between the first entity and the second entity; and determine the fusion confidence based on the first probability and the second probability.

[0196] In one possible implementation, the processing module 1102 is further used to: correct the confidence of the connection relationship between entities on the first path based on the fusion confidence and the degree of association between entities with connection relationships on the first path; and correct the confidence of the connection relationship between entities on the second path based on the fusion confidence and the degree of association between entities with connection relationships on the second path.

[0197] See also Figure 12 , Figure 12 This is a structural diagram of an electronic device provided in an embodiment of the present application. The electronic device 1200 can be specifically manifested as a mobile phone, tablet, laptop, smart wearable device, server, etc., and is used to execute the knowledge graph alignment method provided in this embodiment, which is not limited here. Specifically, the electronic device 1200 includes: a receiver 1201, a transmitter 1202, a processor 1203 and a memory 1204 (wherein the number of processors 1203 in the electronic device 1200 can be one or more, Figure 12(taking one processor as an example), the processor 1203 may include an application processor 12031 and a communication processor 12032. In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203 and the memory 1204 may be connected via a bus or other means.

[0198] The memory 1204 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1203. A portion of the memory 1204 may also include non-volatile random access memory (NVRAM). The memory 1204 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0199] Processor 1203 controls the operation of the electronic device. In specific applications, the various components of the electronic device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0200] The method disclosed in the above embodiment of the present application can be applied to the processor 1203, or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 1203 or the instructions in the form of software. The above-mentioned processor 1203 can be a general-purpose processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0201] The processor 1203 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1204, and the processor 1203 reads the information in the memory 1204 and completes the steps of the above method in combination with its hardware.

[0202] Receiver 1201 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the electronic device. Transmitter 1202 can be used to output digital or character information through the first interface. Transmitter 1202 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group. Transmitter 1202 can also include a display device such as a display screen.

[0203] The electronic device provided in the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit. The processing unit may execute computer-executable instructions stored in the storage unit so that the chip in the electronic device executes the method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0204] For details, please refer to Figure 13 , Figure 13 This is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1300. NPU 1300 is mounted on the host CPU as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is the arithmetic circuit 1303, which is controlled by a controller 1304 to extract matrix data from memory and perform multiplication operations.

[0205] In some implementations, the arithmetic circuit 1303 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1303 is a two-dimensional systolic array. The arithmetic circuit 1303 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1303 is a general-purpose matrix processor.

[0206] For example, assume there are input matrix A, weight matrix B, and output matrix C. The computation circuit retrieves the corresponding data of matrix B from weight memory 1302 and caches it on each PE in the computation circuit. The computation circuit then retrieves the data of matrix A from input memory 1301 and performs a matrix operation on it with matrix B. The partial or final matrix result is stored in accumulator 1308.

[0207] Unified memory 1306 is used to store input and output data. Weight data is directly transferred to weight memory 1302 through the Direct Memory Access Controller (DMAC) 1305. Input data is also transferred to unified memory 1306 through the DMAC.

[0208] BIU stands for Bus Interface Unit 1310 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1309 .

[0209] The bus interface unit 1310 (BIU) is used for the instruction fetch memory 1309 to obtain instructions from the external memory, and is also used for the storage unit access controller 1305 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0210] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1306 or to move weight data to the weight memory 1302 or to move input data to the input memory 1301.

[0211] The vector calculation unit 1307 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1303, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0212] In some implementations, the vector calculation unit 1307 can store the processed output vector in the unified memory 1306. For example, the vector calculation unit 1307 can apply a linear function or a nonlinear function to the output of the operation circuit 1303, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1307 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1303, for example, for use in subsequent layers in a neural network.

[0213] An instruction fetch buffer 1309 connected to the controller 1304 is used to store instructions used by the controller 1304;

[0214] Unified memory 1306, input memory 1301, weight memory 1302, and instruction fetch memory 1309 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0215] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0216] See Figure 14 , Figure 14 This application also provides a computer-readable storage medium. In some embodiments, the above Figure 3 The disclosed methods may be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or articles of manufacture.

[0217] Figure 14 Schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.

[0218] In one embodiment, the computer readable storage medium 1400 is provided using a signal bearing medium 1401. The signal bearing medium 1401 may include one or more program instructions 1402, which when executed by one or more processors may provide the above-mentioned instructions for Figure 3 Describes the functionality or part of the functionality.

[0219] In some examples, the signal bearing medium 1401 may include a computer readable medium 1403 such as, but not limited to, a hard drive, a compact disk (CD), a digital video disk (DVD), a digital tape, a memory, a ROM or RAM, or the like.

[0220] In some embodiments, the signal-bearing medium 1401 may include a computer-recordable medium 1404, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, or the like. In some embodiments, the signal-bearing medium 1401 may include a communication medium 1405, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, or the like). Thus, for example, the signal-bearing medium 1401 may be communicated via a wireless form of the communication medium 1405 (e.g., a wireless communication medium conforming to the IEEE 802.X standard or other transmission protocol).

[0221] The one or more program instructions 1402 may be, for example, computer-executable instructions or logic-implemented instructions. In some examples, the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1402 communicated to the computing device via one or more of computer-readable media 1403, computer-recordable media 1404, and / or communication media 1405.

[0222] It should also be noted that the device embodiments described above are merely illustrative, in which the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0223] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods of each embodiment of the present application.

[0224] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0225] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training equipment or data center to another website, computer, training equipment or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training equipment, data center, etc. that includes one or more available media integrations. Available media can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).

Claims

1. A knowledge graph alignment method, characterized in that: include: Obtaining a first path in a first knowledge graph and a second path in a second knowledge graph, where the first path and the second path are both composed of a plurality of sequentially connected entities, the first path and the second path are different paths between the first entity and the second entity, and the connection relationship between the entities on the first path and the second path has an initial confidence; determining a confidence level of the first path and a confidence level of the second path; Based on the confidence of the first path and the confidence of the second path, revising the initial confidence to obtain a revised confidence of the connection relationship between the entities on the first path and the second path; The first knowledge graph and the second knowledge graph are fused to obtain a fused knowledge graph, wherein the fused knowledge graph includes the first path and the second path, and in the fused knowledge graph, the confidence of the connection relationship between entities on the first path and the second path is the corrected confidence.

2. The method according to claim 1, characterized in that The determining of the confidence of the first path and the confidence of the second path includes: determining, based on the relationship type of the connection relationship between the entities in the first path and the second path, a correlation degree value between the entities having the connection relationship in the first path and the second path; The confidence of the first path and the confidence of the second path are determined based on the association degree value and the initial confidence of the connection relationship between the entities.

3. The method according to claim 2, characterized in that Before determining the association degree values ​​between the entities having the connection relationship in the first path and the second path, the method further includes: Obtaining sorting results corresponding to a plurality of relationship types, the sorting results being used to indicate a result of sorting the plurality of relationship types according to the degree of association corresponding to the relationship types, the plurality of relationship types including relationship types to which connection relationships between entities in the first knowledge graph and the second knowledge graph belong; Based on the ranking result, an association degree value of each relationship type in the multiple relationship types is determined, and the association degree value of each relationship type is used to determine an association degree value between entities having a connection relationship in the first path and the second path.

4. The method according to claim 3, characterized in that The obtaining of sorting results corresponding to multiple relationship types includes: Inputting the plurality of relationship types into a large language model to obtain the ranking result output by the large language model; The large language model is used to sort different relationship types according to the degree of association corresponding to the relationship types.

5. The method according to any one of claims 1 to 4, characterized in that The determining of the confidence of the first path and the confidence of the second path based on the association degree value and the initial confidence of the connection relationship between the entities includes: In the first path, the association degree values ​​between entities with connection relationships are used as weights, and the initial confidence levels of the connection relationships between the entities are weighted and summed to obtain the confidence level of the first path; In the second path, the association degree values ​​between entities having a connection relationship are used as weights, and the initial confidences of the connection relationships between the entities are weighted and summed to obtain the confidence of the second path.

6. The method according to any one of claims 1 to 5, characterized in that The modifying the confidence of the connection relationship between entities on the first path and the second path based on the confidence of the first path and the confidence of the second path includes: Determining a fusion confidence based on the confidence of the first path and the confidence of the second path; Based on the fusion confidence, the confidences of the connection relationships between entities on the first path and the second path are respectively modified.

7. The method according to claim 6, characterized in that The determining of the fusion confidence based on the confidence of the first path and the confidence of the second path includes: Selecting the confidence of a target path from the confidence of the first path and the confidence of the second path as the fusion confidence, wherein the target path is the path with the shortest or longest path length between the first path and the second path; Alternatively, an average of the confidence of the first path and the confidence of the second path is used as the fusion confidence.

8. The method according to claim 6, characterized in that The determining of the fusion confidence based on the confidence of the first path and the confidence of the second path includes: The first path, the second path, the confidence of the first path, and the confidence of the second path are input into a large language model to obtain the fusion confidence. The large language model is used to determine the confidence of the existence of a relationship between the two entities based on different paths between the two entities.

9. The method according to claim 8, characterized in that The step of inputting the first path, the second path, the confidence of the first path, and the confidence of the second path into a large language model to obtain the fused confidence includes: Inputting the first path, the second path, the confidence of the first path, and the confidence of the second path into the large language model, obtaining a first probability and a second probability output by the large language model, wherein the first probability is used to indicate a probability that a relationship exists between the first entity and the second entity, and the second probability is used to indicate a probability that no relationship exists between the first entity and the second entity; The fusion confidence is determined based on the first probability and the second probability.

10. The method according to any one of claims 6 to 9, characterized in that: The step of respectively correcting the confidences of the connection relationships between entities on the first path and the second path based on the fusion confidences includes: Based on the fusion confidence and the association degree values ​​between the entities having the connection relationship on the first path, modifying the confidence of the connection relationship between the entities on the first path; Based on the fusion confidence and the association degree values ​​between the entities having the connection relationship on the second path, the confidence of the connection relationship between the entities on the second path is modified.

11. A knowledge graph alignment device, characterized in that: include: an acquisition module, configured to acquire a first path in a first knowledge graph and a second path in a second knowledge graph, wherein the first path and the second path are both composed of a plurality of sequentially connected entities, the first path and the second path are different paths between the first entity and the second entity, respectively, and the connection relationship between the entities on the first path and the second path has an initial confidence; a processing module, configured to determine a confidence level of the first path and a confidence level of the second path; The processing module is further configured to modify the initial confidence based on the confidence of the first path and the confidence of the second path to obtain a modified confidence of the connection relationship between the entities on the first path and the second path; The processing module is also used to fuse the first knowledge graph and the second knowledge graph to obtain a fused knowledge graph, wherein the fused knowledge graph includes the first path and the second path, and in the fused knowledge graph, the confidence of the connection relationship between entities on the first path and the second path is the corrected confidence.

12. The device according to claim 11, characterized in that The processing module is further configured to: determining, based on the relationship type of the connection relationship between the entities in the first path and the second path, a correlation degree value between the entities having the connection relationship in the first path and the second path; The confidence of the first path and the confidence of the second path are determined based on the association degree value and the initial confidence of the connection relationship between the entities.

13. The device according to claim 12, characterized in that The acquisition module is further used to obtain sorting results corresponding to multiple relationship types, where the sorting results indicate the results of sorting the multiple relationship types according to the degree of association corresponding to the relationship types, and the multiple relationship types include the relationship types to which the connection relationships between entities in the first knowledge graph and the second knowledge graph belong; The processing module is further configured to determine, based on the ranking result, an association degree value for each relationship type among the multiple relationship types, wherein the association degree value for each relationship type is used to determine an association degree value between entities having a connection relationship in the first path and the second path.

14. The device according to claim 13, characterized in that The acquisition module is further used to: Inputting the plurality of relationship types into a large language model to obtain the ranking result output by the large language model; The large language model is used to sort different relationship types according to the degree of association corresponding to the relationship types.

15. The device according to any one of claims 11 to 14, characterized in that The processing module is further configured to: In the first path, the association degree values ​​between entities with connection relationships are used as weights, and the initial confidence levels of the connection relationships between the entities are weighted and summed to obtain the confidence level of the first path; In the second path, the association degree values ​​between entities having a connection relationship are used as weights, and the initial confidences of the connection relationships between the entities are weighted and summed to obtain the confidence of the second path.

16. The device according to any one of claims 11 to 15, characterized in that The processing module is further configured to: Determining a fusion confidence based on the confidence of the first path and the confidence of the second path; Based on the fusion confidence, the confidences of the connection relationships between entities on the first path and the second path are respectively modified.

17. The device according to claim 16, characterized in that The processing module is further configured to: Selecting the confidence of a target path from the confidence of the first path and the confidence of the second path as the fusion confidence, wherein the target path is the path with the shortest or longest path length between the first path and the second path; Alternatively, an average of the confidence of the first path and the confidence of the second path is used as the fusion confidence.

18. The device according to claim 16, characterized in that The processing module is further configured to: The first path, the second path, the confidence of the first path, and the confidence of the second path are input into a large language model to obtain the fusion confidence. The large language model is used to determine the confidence of the existence of a relationship between the two entities based on different paths between the two entities.

19. The device according to claim 18, characterized in that The processing module is further configured to: Inputting the first path, the second path, the confidence of the first path, and the confidence of the second path into the large language model, obtaining a first probability and a second probability output by the large language model, wherein the first probability is used to indicate a probability that a relationship exists between the first entity and the second entity, and the second probability is used to indicate a probability that no relationship exists between the first entity and the second entity; The fusion confidence is determined based on the first probability and the second probability.

20. The device according to any one of claims 16 to 19, characterized in that The processing module is further configured to: Based on the fusion confidence and the association degree values ​​between the entities having the connection relationship on the first path, modifying the confidence of the connection relationship between the entities on the first path; Based on the fusion confidence and the association degree values ​​between the entities having the connection relationship on the second path, the confidence of the connection relationship between the entities on the second path is modified.

21. A knowledge graph alignment device, characterized in that: It includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the knowledge graph alignment device performs the method as described in any one of claims 1 to 10.

22. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 10.

23. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 10.