Local updating method, system, device and medium for knowledge graph
Patent Information
- Application Number
- CN202610973821.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
由于知识图谱通常具有节点数量庞大、关系结构复杂以及关联程度高等特点,上述更新方式往往需要对大量未发生变化的图谱内容进行重复计算,不可避免会产生较高的更新代价
[0009]采用本发明技术方案,首先对获取到的多源异构数据进行封装,得到当前事件及其对应的事件类型,并基于事件类型在待更新知识图谱中定位影响域,从而确定与当前事件相关的静态影响域。由于后续更新处理仅在静态影响域范围内进行,而无需对整个知识图谱执行全量扫描和全量更新,因此能够有效缩小更新范围,减少参与计算的节点、关系以及属性数据量。进一步地,在静态影响域范围内调用与事件类型对应的目标智能体执行抽取、对齐和融合处理,使得更新过程能够针对当前事件涉及的知识内容进行定向处理,避免对未受影响的图谱区域进行重复计算,从而降低知识抽取、实体对齐以及知识融合等环节的计算开销。此外,通过对候选变更集进行校验,并基于通过校验的候选变更子集生成局部知识图谱,再利用局部知识图谱对待更新知识图谱进行局部回写,仅对发生变化的局部图谱结构进行更新,而无需重新构建整个知识图谱。由此能够减少图谱重构过程中的资源消耗和数据处理量,降低知识图谱更新所需的计算资源、存储资源以及时间成本,进而实现降低知识图谱的更新代价。
Smart Images

Figure CN122817477A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge graph updating technology, and in particular to a method, system, device, and medium for local updating of a knowledge graph. Background Technology
[0002] A knowledge graph is a data organization structure that uses nodes to represent entities and edges to represent semantic relationships between entities. It is commonly used in scenarios such as interpersonal relationship mining, big data services for technological innovation, understanding of online audio and video content, enterprise knowledge management, and intelligent question answering. As business data continues to be generated, knowledge graphs need to be updated with newly added facts, deleted facts, and conflicting facts.
[0003] In existing technologies, knowledge graph updates typically employ a full reconstruction approach. This involves re-performing entity alignment, relation extraction, knowledge reasoning, and consistency checks on all global nodes and relationships within the knowledge graph after detecting data changes. Given the typically large number of nodes, complex relational structures, and high degree of interconnectivity in knowledge graphs, this update method often requires recalculating a large amount of unchanged graph content, inevitably resulting in high update costs. Summary of the Invention
[0004] This invention provides a method, system, device, and medium for local updating of knowledge graphs, which can solve at least one of the above-mentioned technical problems.
[0005] In a first aspect, embodiments of the present invention provide a method for local updating of a knowledge graph, comprising: The acquired multi-source heterogeneous data is encapsulated to obtain the current event and the event type corresponding to the current event; Based on the event type, the influence domain is located in the knowledge graph to be updated to obtain the static influence domain; Within the scope of the static influence domain, the target agent corresponding to the event type is invoked to perform extraction, alignment, and fusion processing on the current event to obtain a candidate change set; The candidate change set is verified, and a local knowledge graph is generated based on the verified candidate change subset. Based on the local knowledge graph, a partial write-back is performed on the knowledge graph to be updated to obtain the updated knowledge graph.
[0006] Secondly, embodiments of the present invention provide a local update system for a knowledge graph, comprising: The encapsulation module is used to encapsulate the acquired multi-source heterogeneous data to obtain the current event and the event type corresponding to the current event; The influence domain location module is used to locate the influence domain in the knowledge graph to be updated based on the event type, and obtain the static influence domain. The execution module is used to, within the scope of the static influence domain, invoke the target agent corresponding to the event type to perform extraction, alignment and fusion processing on the current event to obtain a candidate change set; The generation module is used to verify the candidate change set and generate a local knowledge graph based on the candidate change subset that passes the verification. The local write-back module is used to perform local write-back on the knowledge graph to be updated based on the local knowledge graph, so as to obtain the updated knowledge graph.
[0007] Thirdly, embodiments of the present invention also provide a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described in any one of the embodiments of the present invention.
[0008] Fourthly, embodiments of the present invention also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the method described in any one of the embodiments of the present invention.
[0009] The technical solution of this invention first encapsulates the acquired multi-source heterogeneous data to obtain the current event and its corresponding event type. Based on the event type, the influence domain is located in the knowledge graph to be updated, thereby determining the static influence domain related to the current event. Since subsequent update processing is only performed within the static influence domain, without needing to perform a full scan and full update of the entire knowledge graph, the update scope can be effectively narrowed, reducing the amount of nodes, relationships, and attribute data involved in the calculation. Furthermore, within the static influence domain, the target intelligent agent corresponding to the event type is invoked to perform extraction, alignment, and fusion processing, enabling the update process to perform targeted processing on the knowledge content involved in the current event, avoiding repeated calculations on unaffected graph regions, thereby reducing the computational overhead of knowledge extraction, entity alignment, and knowledge fusion. In addition, by verifying the candidate change set and generating a local knowledge graph based on the verified subset of candidate changes, the local knowledge graph is then used to perform a partial write-back of the knowledge graph to be updated, updating only the changed local graph structure without rebuilding the entire knowledge graph. This can reduce resource consumption and data processing during the knowledge graph reconstruction process, reduce the computing resources, storage resources, and time costs required for knowledge graph updates, and thus reduce the cost of updating knowledge graphs.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided for a better understanding of this solution and do not constitute a limitation of the invention. Wherein: Figure 1 This is a flowchart of a method for local updating of a knowledge graph according to an embodiment of the present invention; Figure 2 This is a structural block diagram of a knowledge graph local update system according to an embodiment of the present invention; Figure 3 This is a schematic block diagram of a computer device used to implement the methods of the embodiments of the present invention. Detailed Implementation
[0012] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0013] This invention provides a method, system, device, and medium for partial updating of a knowledge graph. The execution entity of this partial updating method can be the knowledge graph partial updating system provided in this invention, or a computer device integrating the knowledge graph partial updating system. The knowledge graph partial updating system can be implemented in hardware or software, and the computer device can be a terminal or a server.
[0014] Figure 1 This is a flowchart of a method for local updating of a knowledge graph according to an embodiment of the present invention.
[0015] like Figure 1 As shown, the local update method for this knowledge graph may include: S110, encapsulate the acquired multi-source heterogeneous data to obtain the current event and the event type corresponding to the current event; S120, Based on the event type, locate the influence domain in the knowledge graph to be updated to obtain the static influence domain; S130, within the scope of the static influence domain, the target agent corresponding to the event type is invoked to perform extraction, alignment and fusion processing on the current event to obtain a candidate change set; S140, validate the candidate change set, and generate a local knowledge graph based on the validated candidate change subset; S150: Based on the local knowledge graph, perform a local write-back on the knowledge graph to be updated to obtain the updated knowledge graph.
[0016] For example, multi-source heterogeneous data refers to data to be processed that differs in source, data structure, or data modality. Multi-source heterogeneous data may include one or more of the following: text data, structured records, image parsing results, and audio / video semantic results.
[0017] For example, this local update method for knowledge graphs can be applied to efficient management of scientific research results. In the scenario of updating knowledge graphs for university scientific research results, textual data can include paper abstracts, paper keywords, research project summaries, or academic activity introductions; structured records can include paper metadata, author information tables, institutional information tables, or research project records; image parsing results can include the first page of a paper, academic posters, or conference agendas obtained through optical character recognition; and audio-visual semantic results can include academic report text obtained through speech recognition or research topic tags obtained through video content recognition.
[0018] In this example, multi-source heterogeneous data can be obtained through open academic data interfaces, data interfaces of school scientific research information systems, database incremental logs, or file receiving interfaces.
[0019] For example, paper metadata such as paper identifiers, titles, authors, author affiliations, publication dates, and keywords can be obtained through open academic data interfaces; paper abstract text can be obtained through file receiving interfaces; and author and institution names from the paper's first page can be obtained through optical character recognition interfaces. For each piece of data obtained, its data source, data modality, and collection time can be recorded, and it can be treated as a data object to be encapsulated.
[0020] For example, a current event refers to a data object generated based on one or more related data objects to be processed, used to trigger this round of knowledge graph updates. A current event may include an entity identifier field, an attribute field, a timestamp field, and a modality type field.
[0021] The entity identifier field identifies the entity corresponding to the current event. For example, the entity identifier field can be a unique identifier for the paper, an author identifier, an institution identifier, or a research project identifier. The attribute field records the attribute content of the corresponding entity. For example, the attribute fields for a paper may include the paper title, author, author's institution, abstract text, keywords, publication date, retraction flag, or correction flag. The timestamp field records the time when the corresponding data was generated, updated, or collected. The modality type field records the data modality of the corresponding data, such as text type, structured record type, image parsing result type, or audio / video semantic result type.
[0022] For example, the event type can include new events, deletion events, and conflict events. A new event refers to an event where the entity identifier field of the current event does not have a corresponding record in the knowledge graph to be updated. For instance, if newly acquired paper metadata carries a unique paper identifier, and there is no paper node with that unique identifier in the knowledge graph to be updated, then the current event corresponding to that paper metadata can be identified as a new event.
[0023] A deletion event refers to an event whose attribute fields contain a deletion flag, retraction flag, termination flag, or invalidation flag. For example, if the paper status field in the paper's metadata contains a retraction flag, then the current event corresponding to that paper's metadata can be identified as a deletion event.
[0024] A conflict event refers to an event where the entity identifier field of the current event has a corresponding record in the knowledge graph to be updated, but the attribute fields of the current event are inconsistent with the existing attribute fields. For example, if a paper node corresponding to a unique identifier of a paper already exists in the knowledge graph to be updated, and the newly acquired paper metadata is different from the author's institution recorded in that paper node, then the current event corresponding to the paper metadata can be identified as a conflict event.
[0025] For example, if the entity identifier field of the current event is the same as the node identifier in the knowledge graph to be updated, the historical event corresponding to the current event and with the closest timestamp can be obtained. Based on the content difference and time difference between the current event and the historical event, the degree of change of the current event can be determined. If the degree of change of the current event is greater than the preset event judgment threshold, the current event is marked as a conflict event.
[0026] The degree of change in the current event can be calculated as follows: ;in, Indicates the current event Compared to historical events The degree of change; Indicates the current event; This indicates the historical event that has the same entity identifier field as the current event and has the closest timestamp. Indicates the current event With historical events Cross-modal similarity between them; Indicates the current event With historical events The normalized time difference between them; Indicates the weight of content differences; Indicates time difference weight. Content difference weight. The value range is [0,1].
[0027] The cross-modal similarity score ranges from [0,1]. The closer the cross-modal similarity score is to 1, the closer the current event is to a historical event; the closer the cross-modal similarity score is to 0, the more significant the difference in content between the two.
[0028] Cross-modal similarity can be calculated as follows: ;in, Indicates the current event The shared space vector obtained after modality coding; Indicating historical events The shared space vector obtained after modality coding; Represents a shared space vector With shared space vectors The normalized cosine similarity between them; Indicates the current event A collection of attribute fields; Representing historical events A collection of attribute fields; Represents a collection of attribute fields With attribute field collection The similarity between them to Jaccard; Represents vector similarity weights; Indicates the attribute similarity weight; .
[0029] For example, cosine similarity can be normalized to [0,1] as follows: ;in, Represents a shared space vector With shared space vectors The cosine similarity between them ranges from [-1, 1]; This represents the normalized cosine similarity.
[0030] For example, Jaccard similarity can be calculated as follows: ;in, Represents a collection of attribute fields With attribute field collection The number of identical attribute items; This indicates the number of all distinct attribute items in two attribute field sets.
[0031] For example, normalized time difference Used to characterize the time interval between current events and historical events. The normalized time difference can be calculated as follows: ;in, Indicates the current event timestamp; Representing historical events timestamp; This represents the preset time normalization window; This indicates taking the smaller value between the calculation result within the parentheses and 1. This process limits the normalized time difference to within [0,1].
[0032] For example, for paper metadata synchronized monthly, the time normalization window can be set to one month; for research results data summarized annually, the time normalization window can be set to one year.
[0033] For example, an event determination threshold can be preset. For example, a value of 0.6. If the attribute fields of the current event carry a retraction flag, deletion flag, termination flag, or invalidation flag, the current event will be preferentially identified as a deletion event. If the entity identifier field of the current event does not have a corresponding record in the knowledge graph to be updated, the current event will be identified as a new event. If the entity identifier field of the current event has a corresponding record in the knowledge graph to be updated, and Greater than or equal to the preset event determination threshold If so, the current event is identified as a conflict event. Less than the preset event detection threshold If so, the current event is determined to be an event that does not require triggering a graph update.
[0034] By combining cross-modal similarity and normalized time difference to determine the degree of change, we can reduce duplicate updates triggered by changes in field format or repeated collection of the same data within a short period of time, and enable subsequent candidate knowledge extraction, cross-modal alignment, evidence fusion and local write-back processes to focus on the current event that needs to be processed.
[0035] For example, the acquired multi-source heterogeneous data is encapsulated, that is, the data objects to be processed that have a relationship are merged based on the entity identifier field, and the current event is generated according to the preset field structure. The current event and event type obtained by encapsulation are used as inputs for subsequent static influence domain localization, candidate knowledge extraction, cross-modal alignment and evidence fusion processing.
[0036] For example, the preset field structure can be represented as: {Entity Identifier Field, Attribute Field, Timestamp Field, Modal Type Field, Data Source Field}. The Data Source Field records the source of the data object to be processed. For instance, for the current event of the paper's metadata formation, the unique identifier of the paper can be written to the Entity Identifier Field, the paper title, author, author's institution, abstract text, and keywords can be written to the Attribute Field, the update time of the paper's metadata can be written to the Timestamp Field, the structured record type and text type can be written to the Modal Type Field, and the open academic data interface and file receiving interface can be written to the Data Source Field.
[0037] In this example, the process involves: acquiring data objects to be processed from multi-source heterogeneous data; extracting entity identifier, attribute, timestamp, and modality type fields from these data objects; associating data objects with the same entity identifier field; encapsulating the associated data objects according to a preset field structure to obtain the current event; matching the entity identifier field of the current event with the node identifiers in the knowledge graph to be updated; and determining the event type corresponding to the current event based on the matching results and the attribute fields of the current event.
[0038] For example, metadata for a paper can be obtained through an open academic data interface, and the corresponding abstract text can be obtained through a file receiving interface. The paper metadata includes a unique identifier, title, authors, author affiliations, keywords, and update time. Based on the unique identifier, the paper metadata and abstract text can be associated; according to a preset field structure, the unique identifier is written to the entity identifier field, the title, authors, author affiliations, keywords, and abstract text are written to the attribute field, the update time is written to the timestamp field, and the structured record type and text type are written to the modal type field to obtain the current event.
[0039] If no paper node corresponding to the unique paper identifier exists in the knowledge graph to be updated, the event type of the current event is determined to be a new event. If the attribute field of the current event carries a retraction flag, the event type of the current event is determined to be a deletion event. If a paper node corresponding to the unique paper identifier exists in the knowledge graph to be updated, but the affiliation of the author in the current event is different from the affiliation of the author recorded in the paper node, the event type of the current event is determined to be a conflict event. In this way, the event type corresponding to the current event can be obtained.
[0040] According to the above implementation method, the acquired multi-source heterogeneous data is first encapsulated to obtain the current event and its corresponding event type. Then, based on the event type, the influence domain is located in the knowledge graph to be updated, and the static influence domain related to the current event is determined. Further, within the static influence domain, the target intelligent agent corresponding to the event type is invoked to perform knowledge extraction, entity alignment, and knowledge fusion processing to obtain a candidate change set. Subsequently, the candidate change set is verified, and a local knowledge graph is generated based on the verified subset of candidate changes. Finally, the local knowledge graph is used to perform a partial write-back of the knowledge graph to be updated to obtain the updated knowledge graph. By limiting the update scope to the influence domain corresponding to the event and performing local update processing on the knowledge content within the influence domain, the full reconstruction and repeated calculation of the entire knowledge graph are avoided, thereby reducing the computational resource consumption and update cost during the knowledge graph update process.
[0041] In one implementation, the static influence domain is obtained by locating the influence domain in the knowledge graph to be updated based on the event type. This includes: mapping the event type to a preset number of expansion hops to obtain the number of expansion hops; if there is a node in the knowledge graph to be updated that matches the current event, then the matching node is determined as an anchor node; if there is no node in the knowledge graph to be updated that matches the current event, then the nodes whose cross-modal similarity value between the current event and each node in the knowledge graph to be updated is greater than a preset similarity threshold are determined as anchor nodes; starting from the anchor node, the subgraph is expanded along the relation edges in the knowledge graph to be updated according to the number of expansion hops to obtain the static influence domain.
[0042] For example, a static influence domain refers to a local subgraph obtained by expanding along relational edges in the knowledge graph to be updated, starting from the anchor node. The static influence domain includes the nodes traversed during the expansion process and the relational edges between these nodes. The static influence domain is used to limit the processing scope of candidate knowledge extraction, cross-modal alignment, evidence fusion, verification, and local write-back in this round.
[0043] The static influence domain remains unchanged during this round of updates. In other words, once the static influence domain is obtained, subsequent processing will not dynamically expand or shrink it. This avoids the scope of subsequent processing continuously expanding as the number of candidate entities or relationships increases.
[0044] For example, an anchor node refers to a node in the knowledge graph to be updated that is used to determine the starting point for subgraph expansion. There can be one or more anchor nodes. If there are multiple anchor nodes, the subgraph is expanded starting from each anchor node, and the expanded local subgraphs are merged to obtain the static influence domain.
[0045] For example, the extended hop count is used to define the boundary of the static influence domain. A mapping relationship between event types and extended hop counts can be pre-defined. Different event types correspond to different influence domains, so the extended hop count can be configured separately for each. For example, a new event has an extended hop count of 1; a deleted event has an extended hop count of 2; and a conflict event has an extended hop count of 3.
[0046] For example, the mapping relationship between event types and extended hop counts can be stored in a mapping table. The mapping table may include an event type field and an extended hop count field. The event type field is used to record newly added events, deleted events, or conflicting events, and the extended hop count field is used to record the extended hop count for the corresponding event type.
[0047] For example, for new events, new papers typically need to establish relationships with authors, their institutions, and research topics. Therefore, the extended hop count can be configured to cover the paper node and its directly related nodes. For deletion events, deleting or retracting a paper may affect its relationships with authors, institutions, projects, and cited papers. Therefore, the extended hop count can be configured to cover the paper node and the relationship edges associated with it. For conflict events, if the conflict only involves the author's institution, the extended hop count can be configured to cover the author node, the institution node, and the relationship edges between them.
[0048] In this context, the graph pattern refers to predefined node types, relationship types, and various relationship connection rules. For example, the graph pattern of a university research achievement knowledge graph can include paper nodes, author nodes, institution nodes, project nodes, and research topic nodes, as well as relationships such as "author-authorship-paper," "author-affiliation-institution," "paper-related-project," and "paper-related-research topic." Determining the expansion hop count by event type allows the subgraph expansion range to match the impact range of the current event, avoiding a uniform full graph traversal approach and reducing the number of nodes and relationship edges processed subsequently.
[0049] For example, the corresponding node can be found in the knowledge graph to be updated based on the entity identifier field in the current event. If a node in the knowledge graph to be updated carries the same entity identifier field as the current event, then that node is identified as the node that matches the current event and is used as the anchor node.
[0050] For example, if the current event is a paper metadata update event, and the entity identifier field of the current event is a unique identifier for the paper, then nodes with the same unique identifier can be searched among the paper nodes in the knowledge graph to be updated. If a corresponding paper node is found, then that paper node is designated as the anchor node.
[0051] For example, if the current event is an author information update event, and the entity identifier field of the current event is the author identifier, then nodes with the same author identifier can be searched among the author nodes of the knowledge graph to be updated. If a corresponding author node is found, then that author node is designated as the anchor node.
[0052] Node matching based on entity identifier fields can directly determine the graph node corresponding to the current event, reduce the similarity calculation process, and reduce the impact of name abbreviations, authors with the same name, or aliases of institutions on node location results.
[0053] For example, some current events may not carry entity identifier fields that can be directly matched with nodes in the knowledge graph to be updated. For instance, the author name and institution name obtained by optical character recognition from the image of the first page of a paper may not carry author or institution identifiers. In this case, cross-modal similarity values can be calculated based on the current event and each node in the knowledge graph to be updated, and nodes that meet the matching conditions can be identified as anchor nodes.
[0054] For example, cross-modal similarity value refers to a numerical value used to characterize the semantic proximity between the current event and the graph node. The cross-modal similarity value can be calculated based on the text features, structured field features, or image parsing text features of the current event, and the name, alias, descriptive text, and attribute fields of the graph node.
[0055] In this example, feature encoding can be performed on the current event and the nodes in the knowledge graph to be updated to obtain event feature vectors and node feature vectors; the vector similarity between the event feature vectors and node feature vectors can be calculated; and the vector similarity can be determined as the cross-modal similarity value.
[0056] For example, cross-modal similarity values can be calculated using the following formula: In the formula, This represents the cross-modal similarity value between the current event x and node v; This represents the event feature vector obtained after feature encoding the current event x; This represents the node feature vector obtained after feature encoding of the attribute fields of node v; Represents the event feature vector With node feature vectors Cosine similarity between them.
[0057] For example, the preset similarity threshold can range from [0, 1], and in a preferred embodiment, it is set to 0.80. For instance, the cross-modal similarity value between the current event and the first author node is 0.92, the cross-modal similarity value with the second author node is 0.76, and the cross-modal similarity value with the third author node is 0.61. With a preset similarity threshold of 0.80, the first author node with a cross-modal similarity value of 0.92 is identified as the anchor node, while the second and third author nodes are not identified as anchor nodes.
[0058] For example, the current event originates from the optical character recognition (OCR) results of an academic poster image, which includes the author's name, institution abbreviation, and research topic, but not the author identifier. The OCR results can be text-encoded to obtain an event feature vector; the names, institution information, and research directions of each author node in the knowledge graph to be updated can be encoded to obtain node feature vectors; the cross-modal similarity values between the event feature vector and each node feature vector can be calculated; and author nodes with cross-modal similarity values greater than a preset similarity threshold can be identified as anchor nodes.
[0059] By determining anchor nodes through cross-modal similarity values, we can locate associated nodes for the current event when the current event lacks precise entity identifiers, thus avoiding directly writing the current event as an isolated node into the knowledge graph.
[0060] For example, after determining the anchor node and the number of hops to expand, a graph traversal approach can be used to expand the subgraph.
[0061] In this example, the anchor node can be added to the set of nodes to be visited and the set of nodes in the influence domain; the nodes in the set of nodes to be visited can be read and the relation edges connected to the nodes can be obtained; the relation edges can be added to the set of relations in the influence domain; the adjacent nodes connected by the relation edges and which have not yet exceeded the expansion hop count can be added to the set of nodes to be visited and the set of nodes in the influence domain; after completing the graph traversal corresponding to the expansion hop count, a static influence domain can be generated based on the set of nodes in the influence domain and the set of relations in the influence domain.
[0062] For example, the anchor node is the paper node "Paper A", the event type is a conflict event, and the conflict content is a change in the author's affiliation. The number of expansion hops can be determined based on the mapping relationship between the event type and the number of expansion hops, and expansion can be carried out starting from "Paper A" along the relationship edges such as "Author-Authorization-Paper" and "Author-Affiliation-Institution". The static influence domain obtained by expansion can include "Paper A", the corresponding author node, the original institution node, the new institution candidate node, and the relationship edges between the above nodes.
[0063] For example, if the anchor node is the author node "Author B" and the current event is the new academic achievement event, we can start from "Author B" and expand along the relationship edges such as "Author-Authorship-Paper", "Paper-Related-Research Topic", and "Paper-Related-Project" to obtain a static influence domain that includes author nodes, related paper nodes, research topic nodes, project nodes, and corresponding relationship edges.
[0064] By expanding the subgraph with a limited number of hops starting from the anchor node, subsequent candidate knowledge extraction, cross-modal alignment, evidence fusion, verification, and local write-back processes can be confined within the static influence domain. This avoids performing redundant processing on all nodes and all relation edges in the knowledge graph to be updated, thereby reducing the update cost of the knowledge graph and improving the response speed.
[0065] In one implementation, the target agent corresponding to the event type is invoked to perform extraction, alignment, and fusion processing on the current event to obtain a candidate change set. This includes: invoking the extraction agent in the target agent to extract a candidate entity set and a candidate relation set from the current event; invoking the alignment agent in the target agent to perform cross-modal feature mapping on the candidate entity set and the candidate relation set to obtain each shared space vector; calculating the alignment score based on each shared space vector and the attribute field of the current event, and merging homogeneous evidence based on the alignment score to obtain an edge-level evidence set; invoking the fusion agent in the target agent to perform hard conflict detection and weighted fusion processing on each candidate relation edge in the candidate relation set based on the edge-level evidence set to obtain the fusion confidence of each candidate relation edge; and determining the candidate change set based on the candidate entity set, the candidate relation set, and the fusion confidence of each candidate relation edge.
[0066] For example, a target agent refers to a set of data processing modules determined according to the event type corresponding to the current event, used to perform candidate knowledge extraction, cross-modal alignment, and evidence fusion processing. The target agent may include an extraction agent, an alignment agent, and a fusion agent. Each agent sequentially reads the output of the previous agent according to a preset processing order and writes the result of its own processing into the output.
[0067] For example, for new paper additions, the process can be carried out in the order of extracting agents, aligning agents, and fusing agents; for paper retractions, the nodes and relationships to be deleted can be determined in the order of extracting agents and aligning agents, and these nodes and relationships can be written into the candidate change set; for conflict events caused by changes in the author's affiliation, the extracting agents, aligning agents, and fusing agents can be invoked in sequence to process institutional relationship evidence from different sources.
[0068] For example, an extraction agent refers to a data processing module used to identify entities and relationships between entities from the current event. The input of the extraction agent is the current event, and the output is a set of candidate entities and a set of candidate relationships.
[0069] For example, the candidate entity set refers to the collection of candidate entities extracted from the attribute fields of the current event. Candidate entities may include paper entities, author entities, institution entities, research project entities, and research topic entities.
[0070] For example, a candidate relation set refers to the collection of candidate relation edges determined based on the candidate entities and attribute fields of the current event. Each candidate relation edge includes at least a head entity, a relation type, and a tail entity.
[0071] In this example, entity extraction rules and relation extraction rules can be pre-defined. Specifically, entity extraction rules refer to rules used to identify candidate entities based on field names, text fragments, or entity type labels. For example, for paper metadata, the paper corresponding to the paper title field can be identified as the paper entity, the various author names in the author field can be identified as author entities, the institution names in the author affiliation field can be identified as institution entities, and the keywords in the keyword field can be identified as research topic entities.
[0072] Relation extraction rules refer to rules used to determine candidate relation edges based on the associations between fields or semantic relationships in text. For example, based on the association between the paper title field and the author field, candidate relation edges for "author-authorship-paper" can be generated; based on the association between the author field and the author's affiliation field, candidate relation edges for "author-affiliation-institution" can be generated; and based on the association between the paper title field and the keyword field, candidate relation edges for "paper-related-research topic" can be generated. For example, the current event includes the paper title "A Method for Updating Knowledge Graphs for Multi-Source Heterogeneous Data", the author "Student A", the author's affiliation "a certain university", the keywords "knowledge graph" and "multi-source heterogeneous data".
[0073] Extracting the agent yields a candidate entity set: {Paper A, Student A, University, Knowledge Graph, Multi-source Heterogeneous Data}; and a candidate relation set: {Student A - Authorship - Paper A, Student A - Affiliation - University, Paper A - Involved - Knowledge Graph, Paper A - Involved - Multi-source Heterogeneous Data}. By transforming the current event into candidate entity and relation sets through the agent extraction, subsequent processing can revolve around structured entities and relation edges, avoiding direct graph writing to unstructured text.
[0074] For example, an alignment agent refers to a data processing module used to map candidate objects from different data sources or different data modalities to a shared semantic space and determine whether the candidate objects point to the same entity or the same relation. The input of the alignment agent is a set of candidate entities and a set of candidate relations, and the output is a shared space vector corresponding to each candidate object.
[0075] For example, a candidate object can be a candidate entity or a candidate relation edge.
[0076] In this example, corresponding encoding methods can be set for different data modalities. For example, text encoding can be performed on paper abstract text, keyword text, and optical character recognition text; structured field encoding can be performed on fields such as author, institution, and publication time in paper metadata; and image parsing text encoding can be performed on the recognition text of academic poster images. After encoding, each candidate object is mapped to a shared space vector of the same dimension.
[0077] For example, the institution name in the structured paper's metadata might be "University A," the institution name in the optical character recognition result of the paper's first page image might be "University A," and the institution abbreviation in the paper's abstract text might be "A School." The alignment agent can perform feature mapping on these candidate objects separately to obtain corresponding shared space vectors, thus determining whether the names refer to the same institution entity. By mapping candidate objects of different modalities to the shared semantic space, a unified comparison can be performed even with different data sources and representations, reducing the impact of differences in representation between institution abbreviations, English names, and image-recognized text on the entity merging results.
[0078] For example, the alignment score is a numerical value used to characterize whether two candidate objects refer to the same entity or the same relation. The alignment score can range from [0,1]. The closer the alignment score is to 1, the higher the semantic similarity between the two candidate objects.
[0079] In this example, the alignment score between candidate objects can be determined as follows: ; in, Indicates candidate objects With candidate Alignment score between them; Indicates candidate objects The corresponding shared space vector; Indicates candidate objects The corresponding shared space vector; Represents a shared space vector With shared space vectors Cosine similarity between them; Indicates candidate objects A collection of attribute fields; Indicates candidate objects A collection of attribute fields; Represents a collection of attribute fields With attribute field collection Similarity between sets; Represents vector similarity weights; This represents the attribute similarity weight.
[0080] For example, vector similarity weights This can be determined based on historical annotation samples. For example, in scenarios where the paper's metadata fields are relatively complete, it can be... Set it to 0.60, and set the attribute similarity weight to 0.40.
[0081] For example, an alignment threshold can be set. The alignment threshold is a threshold used to determine whether two candidate objects point to the same entity or the same relationship. The alignment threshold can be determined based on the alignment score distribution of candidate object samples with confirmed corresponding relationships and candidate object samples with confirmed non-corresponding relationships.
[0082] For example, the alignment threshold can be set to 0.80. If the vector similarity between candidate "a certain university" and candidate "UniversityA" is 0.93, the set similarity between attribute field sets is 0.85, and the vector similarity weight is 0.60, then the alignment score is: Since the alignment score of 0.898 is greater than the alignment threshold of 0.80, the candidate object "a certain university" and the candidate object "University A" can be merged into the same institutional entity.
[0083] For example, merging of evidence from the same source refers to combining evidence from different sources pointing to the same candidate relation edge into the same evidence set. An edge-level evidence set refers to a set of evidence organized according to the candidate relation edges. Each piece of edge-level evidence may include a data source identifier, data source credibility, and semantic support direction.
[0084] The data source identifier is used to record the source of evidence. For example, the data source identifier can be an open academic data interface, a school's research information system, the optical character recognition result of the first page of a paper, or the speech recognition result of an academic report.
[0085] Data source credibility is used to characterize the reliability of a corresponding data source. The value range of data source credibility can be [0,1]. Data source credibility can be determined based on the historical verification pass rate of the corresponding data source. For example, if the historical verification pass rate of the open academic data interface is 0.95, then the data source credibility corresponding to that data source can be set to 0.95; if the historical verification pass rate of the optical character recognition result on the first page of the paper is 0.82, then the data source credibility corresponding to that data source can be set to 0.82.
[0086] Semantic support direction is used to characterize whether the corresponding evidence supports the candidate relation edge. The semantic support direction corresponding to the candidate relation edge that supports it can be set to 1, the semantic support direction corresponding to the candidate relation edge that opposes it can be set to -1, and the semantic support direction that only supplements attributes without making a support or opposition judgment on the candidate relation edge can be set to 0.
[0087] For example, if evidence is provided for the candidate relation edge "Student A - Affiliation - A Certain University" through the open academic data interface, the optical character recognition results of the paper's first page, and the text introducing academic activities, the following edge-level evidence set can be generated: {{Open Academic Data Interface, 0.95, 1}, {Optical character recognition results on the first page of the paper: 0.82, 1} {Introduction text for academic activities, 0.76, 1}}.
[0088] By merging evidence from different sources into edge-level evidence sets, the evidence source, credibility, and semantic direction of each candidate relation edge can be preserved, providing structured input for subsequent conflict detection and weighted fusion.
[0089] In one implementation, hard conflict detection and weighted fusion processing are performed on each candidate relation edge in the candidate relation set based on the edge-level evidence set to obtain the fusion confidence of each candidate relation edge. This includes: parsing the data source credibility of each piece of evidence in the edge-level evidence set and the semantic support direction corresponding to each candidate relation edge. Based on the data source credibility and semantic support direction, hard conflict detection is performed on each candidate relation edge to obtain each conflict detection result. For each conflict detection result, if the conflict detection result indicates that the candidate relation edge has not triggered a hard conflict, then the data source credibility and semantic support direction are weighted and fused to obtain the fusion confidence of the candidate relation edge. Specifically, based on the credibility of each data source and each semantic support direction, hard conflict detection is performed on each candidate relation edge to obtain each conflict detection result, including: for each piece of evidence corresponding to the same candidate relation edge in the edge-level evidence set, it is determined whether there is a target evidence pair with opposite semantic support directions; if a target evidence pair exists, the credibility of the data source of the two pieces of evidence in the target evidence pair is determined to be greater than or equal to the median credibility of the corresponding edge-level evidence set; if both are greater than or equal to the median credibility, the candidate relation edge is determined to have triggered a hard conflict; if no target evidence pair exists, or if there is evidence in the target evidence pair with a data source credibility less than the median credibility, the candidate relation edge is determined not to have triggered a hard conflict.
[0090] For example, a fusion agent refers to a data processing module used to detect opposing evidence in a border-level evidence set and to perform weighted fusion of the various border-level pieces of evidence in the absence of hard conflicts. The input of the fusion agent is the border-level evidence set, and the output is the conflict detection result and the fusion confidence score.
[0091] In this example, the edge-level evidence set corresponding to the candidate relation edge is obtained; the data source identifier of each piece of evidence in the edge-level evidence set is extracted; based on the preset source credibility mapping table (the preset source credibility mapping table refers to the pre-configured correspondence table between data source identifiers and data source credibility, which includes at least data source identifiers and data source credibility), the data source credibility corresponding to each data source identifier is determined; the relation description field in each piece of evidence is parsed to obtain the semantic support direction corresponding to each piece of evidence.
[0092] Among them, the relationship description field refers to the field used to record the relationship between entities. For example, the relationship description field can be the author field, the author's institution field, the reference field, or the paper status field in the paper's metadata.
[0093] For example, for the candidate relation edge "Student A-Author-Paper A", if the author field of the first academic data interface includes "Student A", then the semantic support direction of the evidence is 1; if the correction field in the second paper metadata file records "Remove author Student A", then the semantic support direction of the evidence is -1.
[0094] By analyzing the data source credibility and semantic support direction of each piece of evidence, unstructured or semi-structured evidence from different sources can be converted into edge-level evidence with a unified field structure, providing an executable data foundation for subsequent hard conflict detection and weighted fusion processing.
[0095] For example, for each edge-level evidence corresponding to the same candidate relation edge, it can be determined whether there is a target evidence pair with opposite semantic support directions; if a target evidence pair exists, it can be determined whether the data source credibility of the two edge-level evidences in the target evidence pair is greater than or equal to the median credibility; if both are greater than or equal to the median credibility, it can be determined that the candidate relation edge triggers a hard conflict; otherwise (that is, there is no target evidence pair, or there is evidence in the target evidence pair with a data source credibility less than the median credibility), it can be determined that the candidate relation edge does not trigger a hard conflict.
[0096] For example, the median credibility refers to the value in the middle when the credibility of the data sources of all the edge-level evidence in the edge-level evidence set is arranged in numerical order. For instance, if the credibility of the data sources corresponding to the edge-level evidence set is {0.95, 0.82, 0.76}, then the median credibility is 0.82.
[0097] For example, for the candidate relation edge "Student A - Affiliated with - A Certain University", if the edge-level evidence set is: {{Open Academic Data Interface, 0.95, 1}, {University Research Information System, 0.93, -1}, {Optical Character Recognition Result of Paper's First Page, 0.82, 1}}, then the median confidence level is 0.93. Since the semantic support directions of the edge-level evidence corresponding to the Open Academic Data Interface and the edge-level evidence corresponding to the University Research Information System are opposite, and the confidence levels of their data sources are both greater than or equal to the median confidence level, this candidate relation edge is determined to have triggered a hard conflict. Candidate relation edges that trigger hard conflicts do not enter the normal weighted fusion process but instead enter a pending review state.
[0098] If no hard conflict is triggered by the candidate relation edges, the fusion confidence can be determined as follows: ;in, This represents the fusion confidence of candidate relation edge e; Let represent the set of edge-level evidence corresponding to candidate relation edge e; i represents a piece of edge-level evidence in the set of edge-level evidence. Indicates the credibility of the data source for the i-th edge-level evidence; This indicates the semantic support direction of the i-th edge-level evidence.
[0099] For example, for the candidate relation edge "Student A - Author - Paper A", the edge-level evidence set is: {{Open Academic Data Interface, 0.95, 1}, {Optical character recognition results on the first page of the paper: 0.82, 1} {Introduction text for academic activities, 0.76, 1}}; The fusion confidence of the candidate relation edge is then: .
[0100] For example, if the boundary-level evidence set is: {{Open Academic Data Interface, 0.95, 1}, {Optical character recognition results on the first page of the paper: 0.82, 1} {Introduction text for academic activities, 0.76, -1}}; If no hard conflict is triggered, then the fusion confidence of the candidate relation edge is: ; By performing hard conflict detection before weighted fusion, evidence with opposite directions and credible sources can be prevented from canceling each other out during the fusion process, allowing candidate relation edges that require further verification to be processed separately from normal candidate relation edges.
[0101] For example, a candidate change set refers to a collection of candidate entities, candidate relation edges, and candidate attribute records that are to be written, updated, or deleted. The candidate change set is not directly written into the knowledge graph to be updated before passing subsequent validation.
[0102] For example, a candidate change admission threshold can be set to determine whether a candidate relation edge should be included in the candidate change set. The candidate change admission threshold can be determined based on the fusion confidence distribution of correct and incorrect candidate relation edges in historical labeled samples.
[0103] For example, we can obtain candidate relationship edge samples that have been manually confirmed as correct and candidate relationship edge samples that have been manually confirmed as incorrect. If the fusion confidence of the correct candidate relationship edge samples is mainly distributed in [0.70, 1], and the fusion confidence of the incorrect candidate relationship edge samples is mainly distributed in [-1, 0.69], then the candidate change admission threshold can be set to 0.70.
[0104] In this example, for candidate relation edges that do not trigger hard conflicts and whose fusion confidence is greater than or equal to the candidate change admission threshold, the candidate relation edge, its corresponding head entity, its corresponding tail entity, and attribute records can be written into the candidate change set. For candidate relation edges that trigger hard conflicts, they can be written into the set to be reviewed. For candidate relation edges that do not trigger hard conflicts but whose fusion confidence is less than the candidate change admission threshold, they can be written into the set to be supplemented evidence.
[0105] For example, if the fusion confidence score of the candidate relationship edge "Student A - Authorship - Paper A" is 1, which is greater than the candidate change admission threshold of 0.70, then this candidate relationship edge and its associated entity are added to the candidate change set. If the fusion confidence score of the candidate relationship edge "Student A - Affiliation - A Certain University" is 0.399, which is less than the candidate change admission threshold of 0.70, then this candidate relationship edge is added to the supplementary evidence set and not directly added to the candidate change set. By filtering candidate relationship edges based on fusion confidence scores, the data entering the subsequent verification process has clear evidentiary support, and the number of candidate relationship edges that need to be processed in the subsequent verification process is reduced.
[0106] In one implementation, the verification of the candidate change set includes: determining compliance for each candidate relation edge in the candidate change set based on the importance weights of each constraint condition in a preset domain constraint set, and obtaining a structured score for each candidate relation edge; calculating a comprehensive verification score for each candidate relation edge based on the structured score and fusion confidence of each candidate relation edge; identifying candidate relation edges whose comprehensive verification score is greater than or equal to a preset comprehensive scoring threshold and do not violate the hard constraints in the domain constraint set as a subset of candidate changes that pass the verification; and identifying candidate relation edges whose comprehensive verification score is less than the comprehensive scoring threshold or violate the hard constraints in the domain constraint set as a subset of candidate changes that fail the verification.
[0107] For example, a domain constraint set refers to a set of constraints pre-set according to the graph pattern of the knowledge graph to be updated, used to verify whether candidate relation edges conform to graph structure rules and business data rules. A domain constraint set may include type constraints, cardinality constraints, temporal constraints, and evidence consistency constraints.
[0108] For example, a graph schema refers to pre-defined entity types, relationship types, and connection rules for various relationship edges. For instance, the graph schema of a university research achievement knowledge graph may include paper entities, student entities, teacher entities, institution entities, research project entities, and research topic entities, as well as relationships such as "student-authorship-paper," "paper-citation-paper," "paper-related-research topic," and "paper-associated-research project."
[0109] For example, type constraints are used to verify whether the head entity type, relation type, and tail entity type of a candidate relation edge conform to preset connection rules. For instance, for the candidate relation edge "Student A - Authorship - Paper A", the relation schema "Student Entity - Authorship Relation - Paper Entity" can be preset. If the head entity of the candidate relation edge is a student entity and the tail entity is a paper entity, then the candidate relation edge passes the type constraint verification. If the tail entity is an institution entity, then the candidate relation edge fails the type constraint verification.
[0110] For example, cardinality constraints are used to verify whether the number of relation edges connected to an entity under a corresponding relation type conforms to a preset number rule. For instance, it can be pre-set that each paper corresponds to at least one author. If paper A corresponds to at least one authorship relation edge, then the authorship relation corresponding to paper A satisfies the cardinality constraint. As another example, it can be pre-set that each paper's unique identifier corresponds to only one paper entity. If a candidate change would cause the same paper's unique identifier to correspond to multiple paper entities, then the candidate change fails the cardinality constraint verification.
[0111] For example, a temporal constraint is used to verify whether the temporal attribute carried by a candidate relation edge satisfies a preset temporal order rule with existing temporal attributes. For instance, it can be pre-set that the official publication date of a paper is no earlier than its submission date. If the submission date of paper A is the first time, and the official publication date is the second time, which is later than the first time, then the candidate relation edge corresponding to paper A passes the temporal constraint verification. If the official publication date is earlier than the submission date, then the corresponding candidate relation edge fails the temporal constraint verification.
[0112] For example, the evidence consistency constraint is used to verify whether there are unresolved hard conflicts in candidate relation edges. For instance, if the candidate relation edge "Student A-Author-Paper A" does not trigger a hard conflict during the evidence fusion processing stage, then the candidate relation edge passes the evidence consistency constraint verification. If the candidate relation edge has a target evidence pair with opposite semantic support directions and whose data source credibility meets the preset conflict conditions, then the candidate relation edge fails the evidence consistency constraint verification.
[0113] By validating candidate relation edges using domain constraint sets, candidate relation edges that do not meet the requirements of graph pattern, quantitative relationships, temporal order, or evidence consistency can be screened out before the local knowledge graph is generated, thus preventing graph structures that do not meet the constraints from being directly written into the knowledge graph to be updated.
[0114] For example, importance weight refers to a numerical value used to characterize the proportion of a corresponding constraint in the structured score calculation process. The sum of the importance weights of all constraints can be 1. Importance weights can be preset based on the degree of influence of constraints on the legality of the graph structure, or determined statistically based on the identification results of erroneous candidate relation edges under different constraints in historical labeled samples. For instance, based on the graph pattern of a university research achievement knowledge graph, the importance weight of type constraints can be set to 0.35, the importance weight of cardinality constraints to 0.20, the importance weight of temporal constraints to 0.20, and the importance weight of evidence consistency constraints to 0.25. The sum of all importance weights is 1.
[0115] The preset domain constraint set can be shown in Table 1 as follows: It should be noted that the importance weights and constraint types mentioned above are merely exemplary configurations. In practical applications, adjustments can be made based on the knowledge graph's schema, data sources, and historical validation results.
[0116] In this example, the verification result can be set to 1 when the candidate relation edge meets a certain constraint, and the verification result can be set to 0 when the candidate relation edge does not meet a certain constraint.
[0117] For example, the structured score of candidate relation edges can be determined as follows: ;in, Representing candidate relation edges Structured score; R represents the set of neighborhood constraints used to validate candidate relation edge e; r represents a constraint condition in the set of neighborhood constraints R. This represents the importance weight of constraint r; This represents the validation result of candidate relation edge e with respect to constraint condition r. If candidate relation edge e satisfies constraint condition r, The value is 1; when candidate relation edge e does not satisfy constraint condition r, It is 0.
[0118] For example, for the candidate relation edge "Student A - Authorship - Paper A", if the candidate relation edge passes the type constraint, cardinality constraint, and evidence consistency constraint, but fails the temporal constraint, then its structured score is: By weighting the verification results of various constraints, the degree of conformity of candidate relation edges with respect to the graph pattern can be converted into a unified value, which facilitates subsequent comprehensive verification by combining the fusion confidence score.
[0119] For example, the comprehensive verification score refers to the numerical value determined by combining the structured score and the fusion confidence of the candidate relation edge, used to characterize whether the candidate relation edge satisfies the local knowledge graph writing conditions. Since the fusion confidence ranges from [-1,1] and the structured score ranges from [0,1], the fusion confidence can be normalized before calculating the comprehensive verification score.
[0120] For example, the normalized fusion confidence level can be determined as follows: ;in, Represents the normalized fusion confidence of candidate relation edge e; This represents the fusion confidence of candidate relation edge e. After normalization, The value range is [0,1].
[0121] For example, the comprehensive verification score can be calculated as follows: ;in, The comprehensive verification score represents the candidate relation edge e; The structured score of candidate relation edge e is represented; Represents the normalized fusion confidence of candidate relation edge e; Indicates the structured score weights; This represents the normalized fusion confidence weight.
[0122] For example, structured score weights This can be determined based on historical labeled samples. For example, if historical verification results indicate that both graph pattern constraints and the degree of evidence support affect the identification results of erroneous relation edges, then the structured score weights can be adjusted. Set the normalized fusion confidence weight to 0.40, with a value of 0.60.
[0123] For example, if the structured score of the candidate relation edge "Student A-Author-Paper A" is 0.80 and the fusion confidence is 0.90, then the normalized fusion confidence is: With a structured score weight of 0.60, the overall verification score is: By combining structured scores and fusion confidence, we can simultaneously consider whether candidate relation edges conform to the graph pattern and whether they are supported by multi-source evidence, avoiding reliance on a single verification rule to determine whether candidate relation edges should be written into the local knowledge graph.
[0124] For example, the comprehensive scoring threshold can range from [0,1]. The comprehensive scoring threshold can be determined based on the distribution of the comprehensive verification scores of correct and incorrect candidate relation edges in the historical labeled samples. For instance, multiple candidate relation edge samples that have been manually confirmed as correct and multiple candidate relation edge samples that have been manually confirmed as incorrect can be obtained, and the comprehensive verification score of each sample can be calculated separately. If the comprehensive verification scores of the correct candidate relation edge samples are mainly distributed in [0.75,1], and the comprehensive verification scores of the incorrect candidate relation edge samples are mainly distributed in [0,0.74], then the comprehensive scoring threshold can be set to 0.75.
[0125] For example, a hard constraint refers to a constraint that a candidate relation edge must satisfy. If a candidate relation edge violates a hard constraint, even if its comprehensive verification score is greater than or equal to the comprehensive score threshold, the candidate relation edge will not be determined as a candidate relation edge that passes the verification.
[0126] For example, type constraints can be considered hard constraints. For the candidate relation edge "Paper A - Author - Student A", if the pre-defined authorship relation direction in the graph schema is "Student Entity - Author - Paper Entity", then the candidate relation edge does not satisfy the type constraint. Even if the comprehensive validation score of the candidate relation edge is 0.90, it will not be included in the validated candidate change subset.
[0127] For example, the cardinality constraint that "each unique paper identifier corresponds to only one paper entity" can be considered a hard constraint. If a candidate change would cause the same unique paper identifier to correspond to multiple paper entities, then that candidate change would be determined as a candidate change that fails the validation. By setting hard constraints, candidate relation edges that do not conform to the basic graph structure rules can be prevented from entering the local knowledge graph solely based on a high degree of evidence support.
[0128] For example, candidate relation edges whose comprehensive verification score is greater than or equal to a preset comprehensive score threshold and do not violate the hard constraints in the domain constraint set are identified as a subset of candidate changes that pass the verification. For instance, the candidate relation edge "Student A-Author-Paper A" has a comprehensive verification score of 0.86, which is greater than the comprehensive score threshold of 0.75 and does not violate the hard constraints. Therefore, this candidate relation edge is added to the subset of candidate changes that pass the verification.
[0129] Candidate relation edges whose comprehensive verification score is less than the preset comprehensive scoring threshold, or that violate the hard constraints in the domain constraint set, are included in the subset of candidate changes that failed the verification. For example, the comprehensive verification score of the candidate relation edge "Student B-Author-Paper B" is 0.61, which is less than the comprehensive scoring threshold of 0.75. Therefore, this candidate relation edge is added to the subset of candidate changes that failed the verification.
[0130] For example, the candidate relation edge "Paper C-Author-Student C" has a comprehensive verification score of 0.88, but its relation direction does not conform to the preset type constraint. Therefore, this candidate relation edge is written into the candidate change subset that failed the verification.
[0131] For example, candidate relation edges that fail validation can be processed according to the reason for failure. If a candidate relation edge violates a hard constraint, it can be deleted; if the relation direction of a candidate relation edge is deviated, the relation direction can be corrected and the validation can be resubmitted; if the number of evidences for a candidate relation edge is insufficient, it can be added to the set of evidence to be supplemented, and the validation can be resubmitted after obtaining new edge-level evidence.
[0132] For example, a local knowledge graph refers to a local graph structure that includes validated candidate relation edges, head entities, tail entities, and attribute records corresponding to each candidate relation edge. The local knowledge graph is used to write back the local subgraph corresponding to the static influence domain in the knowledge graph to be updated.
[0133] For example, if the candidate change subset that passes the verification includes: {Student A - Authorship - Paper A, Paper A - Involved in - Knowledge Graph, Paper A - Related to - Project A}, then Student A, Paper A, Knowledge Graph Subject Entity and Project A can be identified as a local node set, the above candidate relationship edges can be identified as a local relationship set, and the corresponding paper title, keywords and project name can be identified as local attribute records to generate a local knowledge graph.
[0134] By writing only candidate relation edges that meet the comprehensive scoring threshold and do not violate hard constraints into the local knowledge graph, graph structure verification can be completed before local write-back, reducing the impact of candidate relation edges that do not conform to the graph pattern or lack sufficient evidence on the updated knowledge graph.
[0135] In one implementation, a local write-back is performed on the knowledge graph to be updated based on a local knowledge graph to obtain an updated knowledge graph. This includes: extracting the target node set and the target relation set from the local knowledge graph; determining the local subgraph region of the knowledge graph to be updated based on the static influence domain; writing the target node set and the target relation set into the local subgraph region through a local update operator, while keeping the nodes and relations in the knowledge graph to be updated outside the local subgraph region unchanged, so as to obtain the updated knowledge graph.
[0136] For example, by writing back locally, the knowledge graph update process can be made to only apply to the local subgraph regions affected by the current event, avoiding overwriting or rebuilding all nodes and all relation edges in the knowledge graph to be updated, thereby reducing the update cost of the knowledge graph and improving the response speed.
[0137] For example, the target node set refers to the set of nodes in the local knowledge graph that need to be written into the knowledge graph to be updated. The target node set may include newly added nodes and nodes to be updated. A newly added node is a node that exists in the local knowledge graph but does not exist in the knowledge graph to be updated, and has the same entity identifier field. A node to be updated is a node that exists in both the local and the knowledge graph to be updated, and has the same entity identifier field, but the node attributes in the local knowledge graph are different from those in the knowledge graph to be updated.
[0138] For example, the target relation set refers to the set of relation edges in the local knowledge graph that need to be written into the knowledge graph to be updated. The target relation set may include newly added relation edges, relation edges to be updated, and relation edges to be deleted. Newly added relation edges are those that exist in the local knowledge graph but not in the knowledge graph to be updated, having the same head entity, relation type, and tail entity. Relation edges to be updated are those that exist in both the local and the knowledge graph to be updated, having the same head entity, relation type, and tail entity, but carrying different relation attributes. Relation edges to be deleted are those that need to be removed from the knowledge graph to be updated based on the deletion event corresponding to the current event, conflict verification results, or relation correction results.
[0139] For example, extracting the target node set and target relation set from a local knowledge graph may include: reading each local node and each local relation edge in the local knowledge graph; matching each local node with nodes in the knowledge graph to be updated based on the entity identifier field of each local node; identifying local nodes that are not matched as new nodes; identifying local nodes that are matched and whose node attributes have changed as nodes to be updated; matching each local relation edge with relation edges in the knowledge graph to be updated based on the head entity, relation type, and tail entity of each local relation edge; and determining new relation edges, relation edges to be updated, and relation edges to be deleted based on the relation edge matching results and the event type corresponding to the current event.
[0140] For example, a local knowledge graph includes a paper node "Paper A", an author node "Student A", a research topic node "Knowledge Graph Topic", and relationship edges "Student A-Authorization-Paper A" and "Paper A-Involved-Knowledge Graph Topic".
[0141] If the knowledge graph to be updated already contains the paper node "Paper A" and the author node "Student A", but does not contain the research topic node "Knowledge Graph Topic", then the research topic node "Knowledge Graph Topic" will be identified as a new node.
[0142] If the relationship edge "Paper A - Involved - Knowledge Graph Topic" does not exist in the knowledge graph to be updated, then this relationship edge will be identified as a new relationship edge.
[0143] If the relationship edge "Student A-Author-Paper A" already exists in the knowledge graph to be updated, but the author order attribute carried by this relationship edge changes in the local knowledge graph, then this relationship edge is identified as the relationship edge to be updated.
[0144] If the current event is a retraction event, and the local knowledge graph records that the association between the paper node "Paper B" and the project node "Project A" needs to be removed, then the relationship edge "Paper B-Association-Project A" is determined as the relationship edge to be deleted.
[0145] By distinguishing between newly added nodes, nodes to be updated, newly added relation edges, relation edges to be updated, and relation edges to be deleted, clear data input can be provided for subsequent local update operators, avoiding the generalization of local write-back as a global overwrite operation.
[0146] For example, a local subgraph region refers to the graph structure range in the knowledge graph to be updated that corresponds to the static influence domain and allows for writing, updating, or deleting nodes and relation edges during this round of updates.
[0147] In this example, the local subgraph region of the knowledge graph to be updated can be determined as follows: read the set of influence domain nodes and the set of influence domain relations in the static influence domain; identify the nodes in the knowledge graph to be updated that have the same entity identifier field as the set of influence domain nodes as region nodes; identify the relation edges connecting each region node in the knowledge graph to be updated, as well as the relation edges to be processed corresponding to the current event, as region relation edges; and generate the local subgraph region based on the region nodes and region relation edges.
[0148] For example, the set of nodes in the influence domain refers to the set of nodes obtained by expanding the subgraph along the relation edges according to the expansion hop count when determining the static influence domain. The set of relations in the influence domain refers to the set of relation edges traversed during the expansion process when determining the static influence domain.
[0149] For example, the current event is a change in the institution to which the author of Paper A belongs. After expanding the subgraph with Paper A as the anchor node, the static influence domain includes Paper A, author A, original institution A ("First University"), and new institution candidate A ("Second University"), as well as relation edges such as "Student A - Authorship - Paper A" and "Student A - Affiliation - First University".
[0150] Based on the static influence domain, local subgraph regions can be identified within the knowledge graph to be updated. These local subgraph regions include existing nodes with the same entity identifier field as the aforementioned nodes, as well as the relational edges connecting these nodes. By determining local subgraph regions based on the static influence domain, the scope of graph structure that can be modified in this round of updates can be limited, preventing local update operations from affecting nodes and relational edges unrelated to the current event.
[0151] For example, the execution process of the local update operator may include: for each newly added node in the target node set, writing the newly added node into the local subgraph region; for each node to be updated in the target node set, finding the corresponding node in the local subgraph region based on the entity identifier field, and writing the target attribute of the node to be updated into the corresponding node; for each newly added relation edge in the target relation set, writing the newly added relation edge into the local subgraph region; for each relation edge to be updated in the target relation set, finding the corresponding relation edge in the local subgraph region based on the head entity, relation type, and tail entity, and writing the target attribute of the relation edge to be updated into the corresponding relation edge; for each relation edge to be deleted in the target relation set, removing the corresponding relation edge from the local subgraph region; keeping the nodes and relation edges outside the local subgraph region in the knowledge graph to be updated unchanged, thus obtaining the updated knowledge graph.
[0152] For example, the processing rules for the local update operator can be represented by Table 2 as follows: For example, the local subgraph region includes the paper node "Paper A", the author node "Student A" and the institution node "First University", as well as the relationship edges "Student A-Author-Paper A" and "Student A-Affiliation-First University".
[0153] If the target node set includes the newly added node "Second University", and the target relation set includes the newly added relation edge "Student A - Affiliated with - Second University" and the relation edge to be deleted "Student A - Affiliated with - First University", then the local update operator can perform the following processing: Write the newly added node "Second University" into the local subgraph region; write the newly added relation edge "Student A - Affiliated with - Second University" into the local subgraph region; remove the relation edge "Student A - Affiliated with - First University" from the local subgraph region; keep the paper node "Paper A", author node "Student A", relation edge "Student A - Authorship - Paper A", and nodes and relation edges outside the local subgraph region unchanged.
[0154] For example, a partial write-back process can be represented as follows: ; in, This represents the updated knowledge graph; This indicates a knowledge graph that needs updating. Indicates the static influence domain; Represents the target set of nodes; Represents the target set of relations; This represents a local update operator used in the static influence domain. Write or update the target node set within the corresponding local subgraph region The nodes in the set, and the target relation set to write, update or delete. The relationship between the edges.
[0155] By performing write, update, or delete operations only on the local subgraph regions corresponding to the static influence domain, while keeping the graph structure outside the local subgraph regions unchanged, the number of nodes and relation edges that need to be accessed and processed during the knowledge graph update process can be reduced, thereby reducing the update cost of the knowledge graph and improving the response speed.
[0156] Figure 2 This is a structural block diagram of a knowledge graph local update system according to an embodiment of the present invention.
[0157] like Figure 2 As shown, the local update system for this knowledge graph may include: The encapsulation module 510 is used to encapsulate the acquired multi-source heterogeneous data to obtain the current event and the event type corresponding to the current event; The influence domain location module 520 is used to locate the influence domain in the knowledge graph to be updated based on the event type, and obtain the static influence domain. The execution module 530 is used to call the target agent corresponding to the event type to perform extraction, alignment and fusion processing on the current event within the scope of the static influence domain to obtain a candidate change set; The generation module 540 is used to verify the candidate change set and generate a local knowledge graph based on the candidate change subset that passes the verification. The local write-back module 550 is used to perform local write-back on the knowledge graph to be updated based on the local knowledge graph, so as to obtain the updated knowledge graph.
[0158] In one embodiment, the influence domain localization module includes: A mapping unit is used to map the event type to a preset number of extended hops to obtain the number of extended hops; The first determining unit is used to determine the matching node as the anchor node if there is a node in the knowledge graph to be updated that matches the current event. The second determining unit is used to determine the nodes whose cross-modal similarity value between the current event and each node of the knowledge graph to be updated is greater than a preset similarity threshold as anchor nodes if there is no node in the knowledge graph to be updated that matches the current event. The subgraph expansion unit is used to expand the subgraph along the relation edges in the knowledge graph to be updated, starting from the anchor node, according to the expansion hop count, to obtain the static influence domain.
[0159] In one embodiment, the execution module includes: An extraction agent unit is used to call the extraction agent in the target agent to extract a set of candidate entities and a set of candidate relationships from the current event; The alignment agent unit is used to call the alignment agent in the target agent to perform cross-modal feature mapping on the candidate entity set and the candidate relation set to obtain each shared space vector; The same-source evidence merging unit is used to calculate the alignment score based on the attribute fields of each of the shared space vectors and the current event, and to merge the same-source evidence based on the alignment score to obtain the edge-level evidence set; The fusion confidence calculation unit is used to call the fusion agent in the target agent, and perform hard conflict detection and weighted fusion processing on each candidate relation edge in the candidate relation set based on the edge-level evidence set to obtain the fusion confidence of each candidate relation edge; The candidate change set determination unit is used to determine the candidate change set based on the candidate entity set, the candidate relation set, and the fusion confidence of each candidate relation edge.
[0160] In one embodiment, the fusion confidence calculation unit includes: The parsing subunit is used to parse the data source credibility of each piece of evidence in the edge-level evidence set and the semantic support direction corresponding to each candidate relation edge; The hard conflict detection subunit is used to perform hard conflict detection on each of the candidate relation edges based on the credibility of each of the data sources and the semantic support directions, and to obtain each conflict detection result. The weighted fusion subunit is used to perform weighted fusion of the data source credibility and the semantic support direction for each of the conflict detection results. If the conflict detection result is that the candidate relation edge has not triggered a hard conflict, the fusion confidence of the candidate relation edge is obtained.
[0161] In one implementation, the hard collision detection subunit is specifically used for: For each piece of evidence corresponding to the same candidate relation edge in the edge-level evidence set, determine whether there is a target evidence pair with opposite semantic support directions; If the target evidence pair exists, then determine whether the data source credibility of the two pieces of evidence in the target evidence pair is greater than or equal to the median credibility of the edge-level evidence set. If all values are greater than or equal to the median confidence level, then the candidate relation edge is determined to have triggered a hard conflict. If the target evidence pair does not exist, or if the target evidence pair contains evidence whose data source credibility is less than the median credibility, then the candidate relation edge is determined not to have triggered a hard conflict.
[0162] In one implementation, the step of validating the candidate change set is specifically used for: Based on the importance weights of each constraint condition in the preset domain constraint set, compliance determination is performed on each candidate relation edge in the candidate change set to obtain the structured score of each candidate relation edge. Based on the structured score and fusion confidence of each candidate relation edge, the comprehensive verification score of each candidate relation edge is calculated; Each candidate relation edge whose comprehensive verification score is greater than or equal to the preset comprehensive score threshold and does not violate the hard constraints in the domain constraint set is determined as a candidate change subset that passes the verification. Candidate relation edges whose comprehensive verification score is less than the comprehensive score threshold, or that violate the hard constraints in the domain constraint set, are considered as a subset of candidate changes that fail the verification.
[0163] In one implementation, the local write-back module includes: The extraction unit is used to extract the set of target nodes and the set of target relationships in the local knowledge graph; The local subgraph region determination unit is used to determine the local subgraph region of the knowledge graph to be updated based on the static influence domain. The writing unit is used to write the target node set and the target relation set into the local subgraph region through a local update operator, while keeping the nodes and relations in the knowledge graph to be updated outside the local subgraph region unchanged, so as to obtain the updated knowledge graph.
[0164] The specific functions and examples of each module and submodule of the system in this embodiment of the invention can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0165] The acquisition, storage, and application of user personal information involved in the technical solution of this invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0166] This invention also provides a computer device, comprising: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described in any one of the embodiments of the present invention.
[0167] The beneficial effects of the computer device in this embodiment of the invention are equivalent to the beneficial effects of the local update method of the knowledge graph described above, and will not be repeated here.
[0168] This invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the method described in any one of the embodiments of this invention.
[0169] The beneficial effects of the storage medium of the present invention are equivalent to the beneficial effects of the above-described local update method for knowledge graphs, and will not be elaborated here.
[0170] Figure 3 A schematic block diagram of an example computer device 800 that can be used to implement embodiments of the present invention is shown. Computer device 800 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Computer device 800 may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0171] like Figure 3 As shown, the computer device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. The RAM 803 may also store various programs and data required for the operation of the computer device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0172] Multiple components in computer device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows computer device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0173] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the partial update method of a knowledge graph. For example, in some embodiments, the partial update method of a knowledge graph can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the computer device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the partial update method of the knowledge graph described above can be performed. Alternatively, in other embodiments, computing unit 801 may be configured to perform a local update method of the knowledge graph by any other suitable means (e.g., by means of firmware).
[0174] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0175] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0176] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0177] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0178] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0179] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0180] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0181] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for local updating of a knowledge graph, characterized in that, include: The acquired multi-source heterogeneous data is encapsulated to obtain the current event and the event type corresponding to the current event; Based on the event type, the influence domain is located in the knowledge graph to be updated to obtain the static influence domain; Within the scope of the static influence domain, the target agent corresponding to the event type is invoked to perform extraction, alignment, and fusion processing on the current event to obtain a candidate change set; The candidate change set is verified, and a local knowledge graph is generated based on the verified candidate change subset. Based on the local knowledge graph, a partial write-back is performed on the knowledge graph to be updated to obtain the updated knowledge graph.
2. The method according to claim 1, characterized in that, The process of locating the influence domain in the knowledge graph to be updated based on the event type to obtain the static influence domain includes: The event type is mapped to a preset number of extended hops to obtain the number of extended hops; If there is a node in the knowledge graph to be updated that matches the current event, then the matching node will be determined as the anchor node. If there is no node in the knowledge graph to be updated that matches the current event, then the nodes whose cross-modal similarity value between the current event and each node in the knowledge graph to be updated is greater than a preset similarity threshold are identified as anchor nodes. Starting from the anchor node, the subgraph is expanded along the relation edges in the knowledge graph to be updated according to the expansion hop count to obtain the static influence domain.
3. The method according to claim 1, characterized in that, The process involves invoking the target agent corresponding to the event type to perform extraction, alignment, and fusion processing on the current event, resulting in a candidate change set, including: Invoke the extraction agent in the target agent to extract the candidate entity set and candidate relation set from the current event; The alignment agent in the target agent is invoked to perform cross-modal feature mapping on the candidate entity set and the candidate relation set to obtain each shared space vector; Alignment scores are calculated based on the attribute fields of each of the shared space vectors and the current event, and homogeneous evidence is merged based on the alignment scores to obtain a side-level evidence set. The fusion agent in the target agent is invoked to perform hard conflict detection and weighted fusion processing on each candidate relation edge in the candidate relation set based on the edge-level evidence set, so as to obtain the fusion confidence of each candidate relation edge. The candidate change set is determined based on the fusion confidence of the candidate entity set, the candidate relation set, and each candidate relation edge.
4. The method according to claim 3, characterized in that, The process of performing hard conflict detection and weighted fusion processing on each candidate relation edge in the candidate relation set based on the edge-level evidence set to obtain the fusion confidence of each candidate relation edge includes: Analyze the data source credibility of each piece of evidence in the edge-level evidence set and the semantic support direction corresponding to each candidate relation edge; Based on the credibility of each data source and the semantic support direction, hard conflict detection is performed on each candidate relation edge to obtain each conflict detection result. For each of the conflict detection results, if the conflict detection result is that the candidate relation edge has not triggered a hard conflict, then the data source credibility and the semantic support direction are weighted and fused to obtain the fusion confidence of the candidate relation edge.
5. The method according to claim 4, characterized in that, Based on the credibility of each data source and the semantic support direction, hard conflict detection is performed on each candidate relation edge to obtain conflict detection results, including: For each piece of evidence corresponding to the same candidate relation edge in the edge-level evidence set, determine whether there is a target evidence pair with opposite semantic support directions; If the target evidence pair exists, then determine whether the data source credibility of the two pieces of evidence in the target evidence pair is greater than or equal to the median credibility of the edge-level evidence set. If all values are greater than or equal to the median confidence level, then the candidate relation edge is determined to have triggered a hard conflict. If the target evidence pair does not exist, or if the target evidence pair contains evidence whose data source credibility is less than the median credibility, then the candidate relation edge is determined not to have triggered a hard conflict.
6. The method according to claim 1, characterized in that, The verification of the candidate change set includes: Based on the importance weights of each constraint condition in the preset domain constraint set, compliance determination is performed on each candidate relation edge in the candidate change set to obtain the structured score of each candidate relation edge. Based on the structured score and fusion confidence of each candidate relation edge, the comprehensive verification score of each candidate relation edge is calculated; Each candidate relation edge whose comprehensive verification score is greater than or equal to the preset comprehensive score threshold and does not violate the hard constraints in the domain constraint set is determined as a candidate change subset that passes the verification. Candidate relation edges whose comprehensive verification score is less than the comprehensive score threshold, or that violate the hard constraints in the domain constraint set, are considered as a subset of candidate changes that fail the verification.
7. The method according to claim 1, characterized in that, The step of performing a partial write-back on the knowledge graph to be updated based on the local knowledge graph to obtain the updated knowledge graph includes: Extract the target node set and target relation set from the local knowledge graph; Based on the static influence domain, determine the local subgraph regions of the knowledge graph to be updated; The target node set and the target relation set are written into the local subgraph region by a local update operator, while keeping the nodes and relations in the knowledge graph to be updated outside the local subgraph region unchanged, so as to obtain the updated knowledge graph.
8. A local update system for a knowledge graph, characterized in that, include: The encapsulation module is used to encapsulate the acquired multi-source heterogeneous data to obtain the current event and the event type corresponding to the current event; The influence domain location module is used to locate the influence domain in the knowledge graph to be updated based on the event type, and obtain the static influence domain. The execution module is used to, within the scope of the static influence domain, invoke the target agent corresponding to the event type to perform extraction, alignment and fusion processing on the current event to obtain a candidate change set; The generation module is used to verify the candidate change set and generate a local knowledge graph based on the candidate change subset that passes the verification. The local write-back module is used to perform local write-back on the knowledge graph to be updated based on the local knowledge graph, so as to obtain the updated knowledge graph.
9. A computer device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7.