A knowledge graph construction method fusing multi-source heterogeneous open source intelligence
Patent Information
- Application Number
- CN202611213190.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-11
- Publication Date
- 2026-09-29
AI Technical Summary
然而,在多源异构开源情报场景中,不同来源对同一实体、同一事件或同一事实可能存在不同表述,甚至存在时间、地址、金额、负责人、合作关系等方面的矛盾信息
本发明通过将多源异构开源情报中的来源、文档、片段、实体、事件、时间、地点、证据和主题统一转换为异构节点,并根据发布、包含、提及、参与、发生、位于、支持、转载、冲突和同指候选关系建立异构边,使原本分散在不同公开渠道中的情报数据能够以异构图形式进行统一表达,避免仅依赖文本检索或普通三元组抽取导致的来源信息缺失、证据链断裂和事件关联不足的问题。本发明针对候选事实关系生成来源侧可信特征、证据侧可信特征和跨源一致性特征,能够区分官方公告、权威数据库、企业原始披露文件、转载内容、匿名内容等不同来源的可信程度,并结合原始文档定位、片段定位、附件页码定位以及多来源一致性情况,对候选事实进行可信度描述,从而减少单一来源、弱证据来源和长转载链来源对知识图谱构建结果的干扰。
Smart Images

Figure CN122838577A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of knowledge graph construction and open-source intelligence processing technology, specifically to a method for constructing a knowledge graph that integrates multi-source heterogeneous open-source intelligence. Background Technology
[0002] Open-source intelligence refers to intelligence data obtained from publicly available web pages, government announcements, corporate disclosures, bidding information, court judgments, news media, forum content, open databases, and other publicly accessible channels. With the continuous growth of online information, open-source intelligence has become an important data source for corporate risk identification, public safety analysis, business competition assessment, public opinion monitoring, and event correlation analysis. Unlike single-structured databases, open-source intelligence typically features diverse sources, inconsistent formats, significant differences in expression, rapid updates, fragmented evidence chains, and frequent factual conflicts. Relying solely on manual compilation or traditional keyword retrieval methods makes it difficult to form traceable, correlated, and sustainably updated structured knowledge.
[0003] Existing knowledge graph construction methods typically include steps such as entity recognition, relation extraction, entity alignment, and graph storage, which can achieve a certain degree of structured processing of textual information. However, in multi-source heterogeneous open-source intelligence scenarios, different sources may have different descriptions of the same entity, event, or fact, and there may even be contradictory information regarding time, address, amount, responsible person, and cooperative relationship. When handling conflicting information, some existing methods often use simple overwriting, manual screening, or rule filtering, which can easily lead to the loss of valid evidence and make it difficult to preserve the differences in evidence between different sources.
[0004] Meanwhile, existing graph fusion methods often focus on entity name similarity, text semantic similarity, or fixed rule matching, neglecting the credibility of sources, the completeness of evidence, reposting paths, cross-source consistency, and historical conflicts. When a fact originates from a non-original source with a long reposting chain, or when there are clear referencing relationships between multiple sources, treating it the same as original announcements, authoritative databases, or original corporate disclosures will affect the accuracy and credibility of the graph facts. Furthermore, while traditional graph neural networks are generally suitable for processing homogeneous or weakly heterogeneous graphs, they struggle to adequately distinguish the roles of different node types and edge types in knowledge dissemination when dealing with complex heterogeneous structures in open-source intelligence that simultaneously contain source nodes, document nodes, fragment nodes, entity nodes, event nodes, time nodes, location nodes, evidence nodes, and topic nodes.
[0005] Therefore, how to provide a knowledge graph construction method that integrates multi-source heterogeneous open source intelligence is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose a method for constructing a knowledge graph that integrates multi-source heterogeneous open-source intelligence. This invention identifies intelligence objects from multi-source heterogeneous open-source intelligence to obtain a set of open-source intelligence objects including sources, documents, fragments, entities, events, times, locations, evidence, and topics. Based on this set of open-source intelligence objects, an initial multi-source heterogeneous open-source intelligence graph is constructed, allowing different types of intelligence objects and their publication, inclusion, mention, participation, occurrence, location, support, reposting, conflict, and synonymous candidate relationships to be expressed in a heterogeneous graph form. Source-side credibility features, evidence-side credibility features, and cross-source consistency features are generated for candidate fact relationships and fused to obtain a candidate fact credibility description. The initial multi-source heterogeneous open-source intelligence graph, node types, edge types, time information, and candidate events are then integrated. The real credibility description is input into the HGT heterogeneous graph Transformer model. The heterogeneous attention channels of different edge types are adjusted through credibility gating values to obtain entity representations, event representations, evidence representations, and candidate fact representations. Cross-source same-reference judgment is performed based on the consistency of candidate fact representations and evidence neighborhoods. Contradictory candidate facts are retained, scored, and marked with status, generating an open-source intelligence knowledge graph that binds source chains, evidence chains, timestamps, and confidence labels. When new open-source intelligence is added, a local subgraph is determined based on the candidate facts affected by the new intelligence, the set of conflicting facts, and same-reference candidate relationships. This local subgraph is incrementally updated, thereby achieving credible fusion of multi-source heterogeneous open-source intelligence, conflict retention, evidence tracing, and dynamic knowledge graph construction.
[0007] A knowledge graph construction method integrating multi-source heterogeneous open-source intelligence according to an embodiment of the present invention includes the following steps: S1. Obtain multi-source heterogeneous open-source intelligence and identify intelligence objects to obtain a set of open-source intelligence objects including source, document, fragment, entity, event, time, location, evidence and topic; S2. Construct a heterogeneous intelligence graph based on the set of open source intelligence objects, convert different intelligence objects into different types of nodes, and establish different types of edges based on the relationships of publication, inclusion, mention, participation, occurrence, location, support, reprint, conflict, and same reference, to obtain the initial graph of multi-source heterogeneous open source intelligence. S3. Generate source-side credibility features, evidence-side credibility features, and cross-source consistency features for candidate fact relationships, and fuse them to obtain a candidate fact credibility description; S4. Input the initial graph of multi-source heterogeneous open source intelligence, node type, edge type, time information and candidate fact credibility description into the HGT heterogeneous graph Transformer model, perform credibility-aware heterogeneous intelligence propagation, and obtain entity representation, event representation, evidence representation and candidate fact representation; S5. Based on the candidate fact representation and the consistency of the evidence neighborhood, perform same-reference judgment on cross-source entity aliases, organization abbreviations, project codes and event names to obtain cross-source merge nodes; S6. Based on the candidate fact representation, candidate fact credibility description, and evidence node neighborhood, coexistence retention, confidence scoring, and state labeling are performed on contradictory candidate facts to obtain the fact state results. S7. Generate an open-source intelligence knowledge graph based on cross-source merged nodes and fact state results, and bind source chains, evidence chains, timestamps and confidence labels to entities, events, relationships and fact nodes; S8. When new open-source intelligence is added, a local subgraph is determined based on the candidate facts it affects, the set of conflicting facts, and the candidate relationships of the same reference, and incremental updates are performed to obtain the updated open-source intelligence knowledge graph.
[0008] Optionally, S2 includes the following steps: S21. Convert the source, document, fragment, entity, event, time, location, evidence, and topic into heterogeneous nodes respectively; S22. Establish heterogeneous edges based on the candidate relationships of publication, inclusion, mention, participation, occurrence, location, support, reposting, conflict, and reference; S23. Write node type identifiers for heterogeneous nodes and edge type identifiers for heterogeneous edges, and combine them to obtain the initial graph of multi-source heterogeneous open source intelligence.
[0009] Optionally, S3 includes the following steps: S31. Determine the source node, document node, fragment node, and evidence node corresponding to the candidate fact relationship to obtain the source evidence association structure; S32. Generate source-side credible features based on source type, publication time, and reposting path; S33. Generate credible features on the evidence side based on the supporting relationships between document nodes, fragment nodes, and evidence nodes; S34. Generate cross-source consistency features based on the degree of consistency in descriptions of the same candidate fact from multiple independent sources and historical conflict situations; S35. By integrating source-side credibility features, evidence-side credibility features, and cross-source consistency features, a credibility description of candidate facts is obtained.
[0010] Optionally, S35 includes the following steps: S351. Set the source-side credibility features corresponding to official announcements, authoritative databases and original corporate disclosure documents as high credibility features, and set anonymous content and reprinted content lacking original sources as low credibility features. S352. When candidate facts can be located to the original document, paragraph, table cell or attachment page number, the credibility of the evidence side is improved. S353. When candidate facts are supported by two or more independent sources and the entity, event, time and location descriptions are consistent, improve the cross-source consistency feature; S354. When a candidate fact is supported by only a single source, has a referencing relationship, or has a high frequency of historical conflicts, reduce the cross-source consistency feature. S355. Generate candidate fact credibility descriptions for adjusting HGT heterogeneous attention based on the above features.
[0011] 5. Optional, S4 includes the following steps: S41. Read the heterogeneous nodes, heterogeneous edges, node types, edge types, time information, and candidate fact credibility descriptions from the initial graph of multi-source heterogeneous open-source intelligence. S42. Set the type mapping parameters for source nodes, document nodes, fragment nodes, entity nodes, event nodes, time nodes, location nodes, evidence nodes, and topic nodes according to node type; S43. Set the attention parameters for publishing edges, containing edges, mentioning edges, participating edges, occurring edges, located edges, supporting edges, reprinting edges, conflicting edges, and same-pointing candidate edges according to edge type; S44. Generate credibility gating values based on the credibility descriptions of candidate facts, and adjust the heterogeneous attention channels of different edge types; S45. Aggregate information based on type mapping parameters, attention parameters, and adjusted attention channels, and output entity representation, event representation, evidence representation, and candidate fact representation.
[0012] Optionally, S44 includes the following steps: S441. Obtain the center node, adjacent nodes, and edge types participating in the current heterogeneous attention calculation; S442. Read the source-side credibility features, evidence-side credibility features, and cross-source consistency features of the candidate facts corresponding to adjacent nodes; S443. Generate a confidence gating value based on the above features; S444. When the edge type is evidence-supporting edge or co-pointing candidate edge, the credibility gating value is used to enhance the propagation contribution of high-credibility evidence nodes and high-consistency candidate nodes. S445. When the edge type is a source-reprinted edge or a candidate fact conflict edge, use the credibility gating value to suppress the propagation contribution of neighboring nodes with long reprinted chains, missing evidence, or serious historical conflicts. S446. Write the adjustment results into the heterogeneous attention channel of the corresponding edge type, so that the candidate fact representation carries heterogeneous neighborhood information, evidence support information and source credibility information.
[0013] Optionally, S5 includes the following steps: S51. Obtain the cross-origin entity alias, organization abbreviation, project code, or event name to be judged; S52. Generate initial same-reference judgment results based on name similarity, context fragment similarity, and candidate fact representation; S53. Obtain the evidence node, source node, related entity neighborhood, related event neighborhood, time information and location information corresponding to the object to be judged; S54. Correct the initial same-reference judgment result based on the neighborhood consistency of the evidence node and the source node to obtain the same-reference judgment score; S55. When the same index judgment score reaches the merging threshold, the objects to be judged are merged into the same graph node; otherwise, they are retained as different graph nodes and a suspected same-name relationship is established.
[0014] Optionally, S6 includes the following steps: S61. Detect whether there are candidate facts under the same entity or the same event that have different responsible persons, different addresses, different cooperative relationships, different times of occurrence, or different amounts; S62. When there are contradictory candidate facts, establish conflicting fact nodes and retain them in a coexisting manner. S63. Obtain the candidate fact representation, candidate fact credibility description and evidence node neighborhood for each conflicting fact node; S64. Calculate the confidence score based on the candidate fact representation, the candidate fact credibility description, and the evidence node neighborhood; S65. Based on the confidence score, mark the candidate facts as strongly credible facts, weakly credible facts, facts to be confirmed, or conflicting facts.
[0015] Optionally, S62 includes the following steps: S621. Create conflict fact nodes for the contradictory candidate facts and write the subject conflict, address conflict, relationship conflict, time conflict or numerical conflict identifiers. S622. Establish association edges between conflict fact nodes and corresponding entity nodes, event nodes, source nodes, document nodes, fragment nodes, and evidence nodes; S623. Write mutually contradictory conflicting fact nodes into the same conflicting fact set, and retain the source chain and evidence chain corresponding to each conflicting fact node; S624. When performing a graph query, return multiple candidate facts and their conflict types, source chains, evidence chains, and confidence scores.
[0016] Optionally, S8 includes the following steps: S81. Receive new open source intelligence and identify the source nodes, document nodes, entity nodes, event nodes, evidence nodes and candidate fact relationships involved. S82. Determine whether the newly added open-source intelligence affects existing candidate facts, conflicting fact sets, or identical candidate relationships; S83. When there is an impact, determine the local subgraph based on the affected node and its one-hop or two-hop neighborhood. S84. Perform HGT representation update on the local subgraph to obtain the locally updated node representation and candidate fact representation; S85. Recalculate the fact confidence and conflict status based on the locally updated candidate fact representation; S86. Write the updated results back to the open-source intelligence knowledge graph, while keeping the unaffected graph regions unchanged.
[0017] The beneficial effects of this invention are: This invention unifies the sources, documents, fragments, entities, events, times, locations, evidence, and topics in multi-source heterogeneous open-source intelligence into heterogeneous nodes. It then establishes heterogeneous edges based on candidate relationships such as publication, inclusion, mention, participation, occurrence, location, support, reposting, conflict, and attribution. This allows intelligence data originally scattered across different public channels to be uniformly represented in a heterogeneous graph form, avoiding the problems of missing source information, broken evidence chains, and insufficient event associations caused by relying solely on text retrieval or ordinary triple extraction. This invention generates source-side credibility features, evidence-side credibility features, and cross-source consistency features for candidate fact relationships. It can distinguish the credibility levels of different sources such as official announcements, authoritative databases, original corporate disclosures, reposted content, and anonymous content. Combining original document location, fragment location, attachment page number location, and multi-source consistency, it describes the credibility of candidate facts, thereby reducing the interference of single-source, weak-evidence sources, and long reposting chains on the knowledge graph construction results.
[0018] This invention inputs the credibility description of candidate facts into the HGT heterogeneous graph Transformer model and adjusts the heterogeneous attention channels of different edge types through credibility gating values. This allows evidence-supporting edges, same-reference candidate edges, source-reprinted edges, and candidate fact conflict edges to have different weight adjustment methods during intelligence dissemination. The resulting entity representation, event representation, evidence representation, and candidate fact representation not only reflect the heterogeneous neighborhood structure but also the source credibility, evidence support, and cross-source consistency, improving the accuracy of multi-source open-source intelligence fusion mapping. In the cross-source same-reference judgment process, this invention combines candidate fact representation and evidence neighborhood consistency, judging not only based on name similarity and context similarity but also further considering evidence nodes, source nodes, related entity neighborhoods, related event neighborhoods, time information, and location information. This can reduce erroneous merging caused by homonyms, aliasing, confusion of organizational abbreviations, and ambiguity of project codes.
[0019] This invention employs a mechanism of coexistence retention, confidence scoring, and state marking for contradictory candidate facts. Instead of directly deleting or overwriting conflicting facts, it establishes conflicting fact nodes and conflicting fact sets, preserving the source chain and evidence chain corresponding to each conflicting fact. This allows graph queries to return multiple candidate facts supported by different sources, along with their conflict types, evidence bases, and confidence scores, improving the interpretability and traceability of the open-source intelligence knowledge graph. Furthermore, when new open-source intelligence is added, this invention determines a local subgraph based on the candidate facts affected by the new intelligence, the conflicting fact set, and the corresponding candidate relationships. It only updates the HGT representation and fact confidence levels in the affected areas, avoiding redundant construction of the entire knowledge graph. This adapts to application scenarios with continuous updates of open-source intelligence, improving graph update efficiency. Attached Figure Description
[0020] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of the knowledge graph construction method that integrates multi-source heterogeneous open source intelligence proposed in this invention; Figure 2 This is a schematic diagram of the trust-aware heterogeneous intelligence propagation based on HGT proposed in this invention; Figure 3 This is a schematic diagram illustrating the coexistence of conflicting facts and local incremental updates proposed in this invention. Detailed Implementation
[0021] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0022] refer to Figures 1-3 A method for constructing a knowledge graph that integrates heterogeneous open-source intelligence from multiple sources includes the following steps: S1. Obtain multi-source heterogeneous open-source intelligence and identify intelligence objects to obtain a set of open-source intelligence objects including source, document, fragment, entity, event, time, location, evidence and topic; S2. Construct a heterogeneous intelligence graph based on the set of open source intelligence objects, convert different intelligence objects into different types of nodes, and establish different types of edges based on the relationships of publication, inclusion, mention, participation, occurrence, location, support, reprint, conflict, and same reference, to obtain the initial graph of multi-source heterogeneous open source intelligence. S3. Generate source-side credibility features, evidence-side credibility features, and cross-source consistency features for candidate fact relationships, and fuse them to obtain a candidate fact credibility description; S4. Input the initial graph of multi-source heterogeneous open source intelligence, node type, edge type, time information and candidate fact credibility description into the HGT heterogeneous graph Transformer model, perform credibility-aware heterogeneous intelligence propagation, and obtain entity representation, event representation, evidence representation and candidate fact representation; S5. Based on the candidate fact representation and the consistency of the evidence neighborhood, perform same-reference judgment on cross-source entity aliases, organization abbreviations, project codes and event names to obtain cross-source merge nodes; S6. Based on the candidate fact representation, candidate fact credibility description, and evidence node neighborhood, coexistence retention, confidence scoring, and state labeling are performed on contradictory candidate facts to obtain the fact state results. S7. Generate an open-source intelligence knowledge graph based on cross-source merged nodes and fact state results, and bind source chains, evidence chains, timestamps and confidence labels to entities, events, relationships and fact nodes; S8. When new open-source intelligence is added, a local subgraph is determined based on the candidate facts it affects, the set of conflicting facts, and the candidate relationships of the same reference, and incremental updates are performed to obtain the updated open-source intelligence knowledge graph.
[0023] In S1, the system acquires multi-source heterogeneous open-source intelligence and identifies intelligence objects, resulting in a set of open-source intelligence objects including sources, documents, fragments, entities, events, times, locations, evidence, and topics.
[0024] Specifically, the system first parses the raw data from different sources, but this parsing does not simply convert the text into plain text. Instead, it preserves the original location information based on the source structure of the open-source intelligence.
[0025] For web page text, the system parses the web page title, body paragraphs, publication time, publication subject, page links, and page hierarchy structure; For government announcements and corporate disclosure documents, the system parses the announcement title, body paragraphs, attachment names, attachment page numbers, table areas, and seal or signature information; For PDF documents, the system locates pages by page number, paragraph coordinates, table cells, and chapter titles; For bidding announcements, the system parses the purchaser, supplier, project name, project number, amount, time, address, and attachment links; For court judgments, the system analyzes the case number, parties involved, court name, trial date, judgment result, and document paragraphs. For open database records, the system parses the field name, field value, record number, update time, and reference path.
[0026] After completing the structural analysis, the system identifies multiple types of intelligence objects from the intelligence content.
[0027] The source object is used to indicate the entity that releases or reprints the intelligence, such as government websites, corporate websites, news websites, industry platforms, public databases, forum accounts, or social media accounts.
[0028] Document objects are used to represent specific web pages, announcements, documents, news articles, posts, table records, PDF attachments, or open database records.
[0029] Fragment objects are used to represent headings, paragraphs, sentences, table cells, field values, attachment page numbers, or locationable text areas in a document.
[0030] Entity objects are used to represent enterprises, institutions, personnel, projects, products, equipment, departments, organizations, qualifications, licenses, and contract objects.
[0031] Event objects are used to represent events such as winning bids, penalties, collaborations, litigation, changes, disclosures, tenders, procurements, public opinion events, and announcements.
[0032] Time objects are used to represent the time of intelligence release, the time of event occurrence, the time of announcement release, the time of contract signing, the time of judgment, or the time of update.
[0033] Location objects are used to represent registered addresses, project locations, event locations, organizational locations, or areas covered by announcements.
[0034] Evidence objects are used to represent the specific original location that supports a candidate fact, such as a webpage link, PDF page number, announcement paragraph number, table cell coordinates, database field number, or attachment location.
[0035] Subject objects are used to represent the business subject to which the intelligence belongs, such as corporate risk, bidding, litigation, administrative penalties, public opinion clues, project cooperation or industry relations.
[0036] Each identified intelligence object is assigned a unique object number, and a correspondence is established between the object and its source, document, or fragment.
[0037] For example, if a company name is identified from the third paragraph on page 2 of a government announcement, then that entity is simultaneously associated with the corresponding source node, document node, page number fragment node, and evidence node.
[0038] For example, if a winning bid amount is identified from a table cell, then the relevant candidate facts for that amount will be associated with the document containing the table, the location of the table cell, and the source of the announcement.
[0039] In this way, the system retains the evidence location information required for subsequent credibility judgment during the object recognition stage, avoiding the inability to trace the source of facts after the knowledge graph is generated.
[0040] In S2, the system constructs an initial graph of multi-source heterogeneous open-source intelligence based on a set of open-source intelligence objects.
[0041] Specifically, the system converts sources, documents, fragments, entities, events, times, locations, evidence, and topics into heterogeneous nodes, and writes a node type identifier for each type of node.
[0042] Each node includes at least a node number, node type, source number, document number, fragment number, timestamp, text summary, object attributes, and status field. The node number is used to uniquely identify the node in the graph. Node types are used to distinguish different nodes such as Source, Document, Segment, Entity, Event, Time, Location, Evidence, and Topic; Source ID is used to trace the source of intelligence; document ID is used to locate the original document; fragment ID is used to locate a specific text or table area. Timestamps are used to record the publication time, event time, or graph creation time of the object corresponding to the node; Text summaries are used to store a brief description of the content corresponding to a node; Object properties are used to store attributes such as entity name, event type, location name, time value, and evidence path; The status field is used to record whether a node is a newly created node, a merged node, a suspected node with the same name, or a node related to a conflict.
[0043] The system further establishes heterogeneous edges based on the relationships between different intelligence objects.
[0044] A publishing edge is established between the source node and the document node to indicate that a document is published by a certain source; An containment edge is established between document nodes and fragment nodes to indicate the location where the document contains a title, paragraph, table cell, or attachment; A mention edge is established between fragment nodes and entity nodes to indicate that a fragment mentions a certain enterprise, person, project or organization; Participation edges are established between entity nodes and event nodes to indicate that an entity participates in an event; An occurrence edge is established between the event node and the time node to indicate when the event occurs or the announcement is published. An edge is established between the event node and the location node to represent the location where the event occurred or the project location; Support edges are established between evidence nodes and candidate fact nodes to indicate that a piece of evidence supports a candidate fact; Reprint edges are established between source nodes to indicate that content from one source originates from another source or references another source; Conflict edges are established between candidate fact nodes to indicate that two candidate facts point to the same object but have different values; Establish candidate edges between entity nodes to indicate that two entities may be the same object.
[0045] Each heterogeneous edge includes at least the following fields: edge number, starting node number, ending node number, edge type, candidate fact number, evidence number, source number, direction information, establishment time, and edge status.
[0046] Edge types are used to distinguish between publishing edges, containing edges, mentioning edges, participating edges, occurring edges, located edges, supporting edges, reprinting edges, conflicting edges, and same-reference candidate edges.
[0047] The candidate fact number is used to associate the edge with a fact to be judged. The evidence number is used to associate the edge with the location of the original evidence.
[0048] The edge status field is used to identify whether the edge is a new edge, a verified edge, an edge pending confirmation, or a conflict edge.
[0049] For example, if a bidding announcement states "Company A won the bid for Project B, with an amount of 5 million yuan", the system can establish publishing edges from the announcement source node to the announcement document node, inclusion edges from the announcement document node to the relevant paragraph nodes, mention edges from the paragraph nodes to the entity nodes of Company A and Project B, participation edges from the entity node of Company A to the winning bid event node, and association edges from the winning bid event node to the amount fact node, time node, and location node, and establish supporting edges from the corresponding evidence nodes.
[0050] Thus, the system obtains an initial graph of multi-source heterogeneous open-source intelligence that can simultaneously express the relationships between sources, evidence, entities, events, time, location, and candidate facts.
[0051] In this embodiment, to facilitate subsequent processing of the HGT heterogeneous graph Transformer model, the system encodes the node type and edge type.
[0052] Node type identifiers can use integer numbers, such as source node number 1, document node number 2, fragment node number 3, entity node number 4, event node number 5, time node number 6, location node number 7, evidence node number 8, topic node number 9, and candidate fact node number 10.
[0053] Edge types can also use integer numbers, for example, publishing edge number 1, containing edge number 2, mentioning edge number 3, participating edge number 4, occurring edge number 5, located edge number 6, supporting edge number 7, reposting edge number 8, conflicting edge number 9, and same-pointing candidate edge number 10.
[0054] The above encoding is only for the example. In other examples, string type, enumeration type or database dictionary table can also be used for representation.
[0055] In S3, the system generates source-side credible features, evidence-side credible features, and cross-source consistency features for candidate fact relationships in the initial graph of multi-source heterogeneous open source intelligence, and fuses them to obtain a candidate fact credibility description.
[0056] Candidate factual relationships refer to factual relationships extracted from intelligence fragments that are to be written into the graph, such as "Company A participates in Project B", "Company A wins the bid for Project B", "Personnel C serves as the person in charge of Company A", "Event D occurs at time T", "Company A's registered address is location L", "Project B's amount is value M", "Organization E punishes Company A", etc.
[0057] For each candidate fact, the system first determines its corresponding source node, document node, fragment node, and evidence node to obtain the source evidence association structure.
[0058] The source evidence association structure is used to describe which source the candidate fact is obtained from, which document it is carried by, which segment it is located in, and which evidence node it is supported by, and is used for subsequent credibility calculation and HGT attention adjustment.
[0059] Source-side credibility features are used to represent the credibility of the source corresponding to the candidate fact.
[0060] In this embodiment, the source-side trust features include source level, publication time freshness, and reposting path level.
[0061] Official announcements, authoritative databases, and original corporate disclosure documents are classified as Level 3 sources. The source level for news media, industry websites, and publicly released corporate press releases is set to Level 2. Forum content, anonymous content, and reprinted content lacking the original source are assigned a source level of 1.
[0062] The freshness of the release time is determined by the interval between the current map update time and the intelligence release time. An interval of no more than 30 days is set as high freshness, an interval of more than 30 days but no more than 180 days is set as medium freshness, and an interval of more than 180 days is set as low freshness.
[0063] The reposting path level is determined according to the propagation hierarchy from the original source to the current source. The length of the reposting chain directly from the original source is set to 0, the length of the chain arriving through one reposting source is set to 1, and the length of the chain arriving through two or more reposting sources is set to level 2 or above.
[0064] When the system cannot identify the original source, it will mark the reposting path level as unknown and write an incomplete source identifier in the candidate fact credibility description.
[0065] The credibility feature of evidence is used to indicate whether the evidence corresponding to the candidate fact is complete, locatable, and directly supports the fact.
[0066] In this embodiment, the credible characteristics of the evidence include the completeness of the evidence location and the type of evidence support.
[0067] When candidate facts can be located to the original web page address, PDF page number, announcement paragraph, table cell or database record field, the evidence side credibility feature is set to complete; When the document can only be located but not a specific paragraph, page number, or field, set the evidence-side credibility feature to partially complete. When only the paraphrased text is available but the original path is missing, set the evidence side credibility feature to incomplete.
[0068] Evidence can be categorized into direct evidence, indirect evidence, and weak evidence.
[0069] Direct evidence refers to evidence in which the subject, relationship, and object of the candidate fact are clearly presented; Indirect evidence refers to evidence fragments that only contain part of the content and require inference from other fragments. Weak evidence refers to evidence that only contains vague descriptions or reprinted summaries, making it impossible to directly confirm the candidate facts.
[0070] Cross-source consistency features are used to indicate whether multiple sources describe the same candidate fact consistently, and whether related sources frequently conflict during historical graph construction.
[0071] In this embodiment, cross-source consistency features include the number of independent sources, the degree of consistency description, and the level of historical conflict.
[0072] When two or more independent sources provide consistent descriptions of entities, events, times and locations in the same candidate fact, the cross-source consistency feature is set to high consistency. When two sources have the same content but there is a clear referencing relationship, they are not considered as two independent sources, but rather counted as the same source chain; When there is only a single source to support it, or when the independence between multiple sources cannot be determined, the cross-source consistency feature is set to low consistency. When the same source generates conflicting facts multiple times during the historical mapping process, the historical conflict level corresponding to that source is increased.
[0073] For example, if a source generates more than ten candidate facts that are marked as conflicting facts among the most recent one hundred candidate facts, then the historical conflict level of that source is set to high. If the number of conflicting facts is between three and ten, then it is set to medium. If there are fewer than three, then set it to low. The above quantities are only values used in this example and can be adjusted according to the data scale in different business scenarios.
[0074] The system generates candidate fact credibility descriptions based on source-side credibility features, evidence-side credibility features, and cross-source consistency features.
[0075] Specifically, when the source level is level 3, the evidence side credibility feature is complete, and the cross-source consistency feature is highly consistent, the candidate fact credibility description is marked as highly credible. When the source level is 2 or the evidence side credibility feature is partially complete and there is no obvious conflict, the candidate fact credibility description is marked as medium credibility; When the source level is 1 and the evidence side credibility feature is incomplete or only a single source supports it, the candidate fact credibility description is marked as low credibility; When a candidate fact directly conflicts with an existing highly credible fact, the credibility description of the candidate fact is marked as conflict pending judgment.
[0076] The credibility description of candidate facts may include fields such as credibility level, source level, evidence completeness, source independence, reprint path level, historical conflict level, number of supporting evidence, and conflict identifier.
[0077] The above-mentioned level generation rules are not limited to fixed rules. In other embodiments, the level division can be adjusted according to the results of manual verification, the authority of industry sources, and application scenarios.
[0078] In this embodiment, the candidate fact credibility description is not only used for the final fact confidence score, but also serves as input information for the HGT heterogeneous graph Transformer model to propagate credibility perception.
[0079] In other words, the system does not simply score facts after the HGT model outputs them. Instead, it adjusts the attention channels of different edge types during the heterogeneous information propagation process by utilizing source-side credibility features, evidence-side credibility features, and cross-source consistency features. This allows the entity representations, event representations, evidence representations, and candidate fact representations generated by the model to naturally carry source credibility and evidence support information.
[0080] This process makes the credibility description of candidate facts an integral part of the graph representation learning process, rather than an additional label independent of the graph structure.
[0081] In S4, the system inputs the initial graph of multi-source heterogeneous open-source intelligence, node types, edge types, time information, and candidate fact credibility descriptions into the HGT heterogeneous graph Transformer model to perform credibility-aware heterogeneous intelligence propagation, resulting in entity representations, event representations, evidence representations, and candidate fact representations.
[0082] The HGT Heterogeneous Graph Transformer model is used to handle heterogeneous graph structures where different node types and different edge types coexist.
[0083] Unlike ordinary isomorphic graph neural networks, the HGT model can set different type mapping parameters according to the node type and different attention parameters according to the edge type, thereby distinguishing the different roles of source nodes, document nodes, fragment nodes, entity nodes, event nodes, time nodes, location nodes, evidence nodes, topic nodes, and candidate fact nodes in graph propagation.
[0084] Specifically, the system first reads the heterogeneous nodes, heterogeneous edges, node type identifiers, edge type identifiers, time information, and candidate fact credibility descriptions from the initial graph of multi-source heterogeneous open-source intelligence.
[0085] Then, set the type mapping parameters according to the node type to map different types of nodes to a unified representation space.
[0086] For example, the initial characteristics of a source node can consist of source type, source level, historical conflict level, and source text summary; The initial characteristics of a document node can consist of the document title, publication time, document type, and source number; The initial features of a fragment node can consist of fragment text, fragment position, document number, and fragment type; The initial characteristics of an entity node can consist of entity name, entity category, alias, number of associated sources, and historical merge status; The initial characteristics of an event node can consist of event type, event time, event location, and participating entities; The initial features of an evidence node can consist of the evidence path, the completeness of evidence location, and the type of evidence support. The model uses type mapping parameters to map these node features with different dimensions and meanings into a node representation with a unified dimension.
[0087] The system also sets edge type attention parameters according to edge type, so that different edge types have different propagation channels.
[0088] For example, publishing edges are primarily used to transmit source information between a source and a document; Edges are included to convey document structure information; Reference edges are used to convey semantic information between text fragments and entities; Participating edges are used to convey information about the relationships between entities and events; The occurrence edge and the location edge are used to transmit time and location information; Supporting edges are used to convey supporting information about the candidate facts; The reprinted edge is used to convey information about the source and dissemination path; Conflict edges are used to convey information about the differences between contradictory facts; Candidate edges that point to the same point are used to convey merging information between objects that may be the same.
[0089] By using the aforementioned edge type attention parameters, the HGT model is able to distinguish the semantic roles of different relations during propagation.
[0090] In this embodiment, the hidden dimension of the HGT model is set to 128, the number of attention heads is set to 4, the number of stacking layers is set to 2, and the number of node representation update rounds is set to 2.
[0091] The above parameters are selectable values for this embodiment. When the data scale is large, the number of nodes is large, or the relationship type is more complex, the hidden dimension can be set to 256, the number of attention heads can be set to 8, and the number of stacking layers can be set to 3 or 4.
[0092] In scenarios with small data scale and high real-time requirements, the hidden dimension can be set to 64, the number of attention heads can be set to 2, and the number of stacking layers can be set to 1.
[0093] The above parameter adjustments do not affect the core mechanism of this invention, which utilizes the credibility of candidate facts to describe and adjust the heterogeneous attention channel of HGT.
[0094] For each central node, the model reads its neighboring nodes, the type of the central node, the type of the neighboring nodes, the edge type between them, and the credibility description of the candidate facts associated with the neighboring nodes, and generates a credibility gating value based on the source-side credibility feature, the evidence-side credibility feature, and the cross-source consistency feature.
[0095] The credibility gating value is used to adjust the heterogeneous attention channels of different edge types, rather than as a separate post-processing score.
[0096] Specifically, when the edge type is an evidence-supporting edge, the credibility gating value is preferentially adjusted by the credibility features of the evidence side, and evidence nodes that can locate the original document, original paragraph, table unit or attachment page number have a higher propagation contribution. The contribution of evidence nodes that are incomplete or merely paraphrases to the dissemination of evidence is reduced.
[0097] When the edge type is a source-retransfer edge, the credibility gating value is preferentially adjusted by the retransfer path level and the source level, and the propagation contribution of the original source or short retransfer chain source is higher than that of the long retransfer chain source. When the source level is low and the original source cannot be located, the propagation contribution of the reprinted edge is suppressed.
[0098] When the edge type is a candidate fact conflict edge, the credibility gating value is preferentially adjusted by the cross-source consistency feature and the historical conflict level. The propagation contribution of adjacent nodes with high historical conflict frequency, lack of evidence, or contradiction with highly credible facts is reduced.
[0099] When the edge type is an entity-related candidate edge, the credibility gating value is primarily adjusted by the consistency of the evidence neighborhood, the independence of the source, and the similarity of the candidate fact representation. The propagation contribution of nodes with similar names but inconsistent evidence sources, event times, or related objects is reduced.
[0100] In a specific propagation process, if the central node is a candidate fact node and its adjacent nodes include evidence nodes, source nodes, entity nodes, and conflicting fact nodes, the model will assign different propagation weights according to different edge types.
[0101] A complete evidence node from an evidence-supporting edge can enhance the evidence-related information in the candidate fact node representation; The propagation contribution of source nodes in long reprint chains originating from source edges will decrease; Low-credibility conflict fact nodes from conflict edges do not directly overwrite the candidate fact node representation, but instead enter the representation as conflict information; Highly consistent candidate entity nodes from entity-symmetric candidate edges can enhance the entity semantic stability of candidate fact nodes.
[0102] In the above manner, the model obtains the credibility-aware attention weights, and aggregates information from heterogeneous nodes based on the type mapping parameters, edge type attention parameters, and credibility-aware attention weights, outputting entity representations, event representations, evidence representations, and candidate fact representations.
[0103] In S5, the system performs same-reference judgment on cross-source entity aliases, organization abbreviations, project codes, and event names based on the candidate fact representation and the consistency of evidence neighborhood, and obtains cross-source merge nodes.
[0104] It should be noted that in this embodiment, name similarity and context fragment similarity are only used as candidate recall conditions to filter candidate nodes that may represent the same object, and are not used as the final merging basis.
[0105] In other words, even if two entity names are highly similar, the system will not directly merge them based solely on name similarity. Instead, it will further combine the candidate fact representations, evidence node consistency, source node independence, related event consistency, time information, and location information output by the HGT model to make a judgment.
[0106] Specifically, the system first obtains the object to be judged, such as the abbreviation of a company, the alias of an organization, the abbreviation of a project, the alias of a person, or the variant of an event name, and generates a set of candidate names based on name similarity and context fragment similarity.
[0107] Subsequently, the system obtains the evidence node, source node, related entity neighborhood, related event neighborhood, time information, location information, and candidate fact representation for each candidate object.
[0108] If the evidence nodes corresponding to two candidate objects point to the same or mutually corroborating original evidence, the source nodes are independent of each other, the related events and related entities are highly consistent, and the candidate fact representations output by HGT are similar, then the same-reference judgment score will be increased.
[0109] If two candidate objects have similar names but significantly different addresses, responsible persons, project numbers, event times, or sources of evidence, their similarity judgment scores will be reduced, and they will be retained as different graph nodes.
[0110] In this embodiment, the merging threshold is set to 0.80, and the suspected threshold is set to 0.60.
[0111] When the same-name judgment score reaches 0.80 or above, the objects to be judged are merged into the same graph node; when the same-name judgment score is below 0.60, they are retained as different graph nodes; when the same-name judgment score is between 0.60 and 0.80, a suspected same-name relationship is established, and the judgment is updated when new information is added in the future.
[0112] The threshold values mentioned above are those used in this embodiment. In other embodiments, they can be adjusted based on the number of manually labeled samples, the number of open-source intelligence sources, business fault tolerance requirements, and the results of manual review.
[0113] For example, in risk assessment scenarios, to avoid incorrect merging, the merging threshold can be increased to 0.85; In the scenario of recalling public opinion clues, in order to improve the recall rate, the suspected threshold can be reduced to 0.55.
[0114] In S6, the system performs coexistence retention, confidence scoring, and state labeling on contradictory candidate facts based on candidate fact representations, candidate fact credibility descriptions, and evidence node neighborhoods, thereby obtaining the fact state results.
[0115] Specifically, the system detects whether there are candidate facts with different fact values under the same entity or the same event. The different fact values include different responsible persons, different registered addresses, different cooperative relationships, different occurrence times, different project amounts, different penalty results, different event statuses, or different announcement versions.
[0116] When contradictory candidate facts are detected, the system does not directly delete either candidate fact, nor does it use the later-entered fact to directly overwrite the earlier-entered fact. Instead, it creates conflict fact nodes for each contradictory candidate fact and writes these conflict fact nodes into the same set of conflict facts.
[0117] Each conflict fact node includes at least the conflict fact number, associated entity number, associated event number, conflict type, fact value, candidate fact credibility description, confidence score, source chain number, evidence chain number, first discovery time, and most recent update time.
[0118] Conflict types can include subject conflicts, address conflicts, relationship conflicts, time conflicts, value conflicts, state conflicts, or version conflicts.
[0119] The system establishes association edges between conflicting fact nodes and corresponding entity nodes, event nodes, source nodes, document nodes, fragment nodes, and evidence nodes, enabling each conflicting fact to be traced back to its corresponding source and evidence.
[0120] For example, if the same company's registered address has two factual values, location A and location B, the system will establish address conflict fact nodes separately, and associate the fact node corresponding to location A with the company's public documents, original field evidence, and company entity nodes, and associate the fact node corresponding to location B with reprinted news, reprinted page fragments, and news source nodes.
[0121] For each conflicting fact node, the system obtains its candidate fact representation, candidate fact credibility description, and evidence node neighborhood, and calculates the fact confidence score.
[0122] The fact confidence score is used to indicate the degree of credibility of the conflict fact within the current open-source intelligence graph.
[0123] In this embodiment, the confidence score is represented in the range of 0 to 1.
[0124] Candidate facts with a score greater than or equal to 0.80 are marked as strongly credible facts; Candidate facts with a score greater than or equal to 0.60 and less than 0.80 are marked as weakly credible facts; Candidate facts with a score greater than or equal to 0.40 and less than 0.60 are marked as facts to be confirmed; Candidate facts with a score less than 0.40 or that directly contradict highly credible facts are marked as conflicting facts.
[0125] The above scoring range is the value taken in this embodiment. In other embodiments, it can be adjusted according to the tolerance for false positives and false negatives in the business scenario.
[0126] For example, for applications with a high proportion of highly credible sources such as judicial and administrative penalties, the threshold for strongly credible facts can be increased; for applications that discover public opinion clues, the threshold for facts to be confirmed can be decreased in order to retain more clues.
[0127] It should be noted that the coexistence and retention of conflicting facts does not mean that all candidate facts have the same level of credibility, but rather that the system does not directly delete weak facts when there is a lack of sufficient evidence.
[0128] The system distinguishes different candidate facts by fact state, confidence score, source chain, and evidence chain.
[0129] When a user queries a specific entity or event, the system can return strongly credible facts, while also displaying existing weakly credible facts, facts to be confirmed, or conflicting facts, and showing the corresponding sources of evidence.
[0130] This avoids the loss of evidence due to simple overwriting, and also avoids directly writing conflicting facts as established facts into the main graph result.
[0131] In S7, the system generates an open-source intelligence knowledge graph based on cross-source merged nodes and fact state results, and binds source chains, evidence chains, timestamps, and confidence labels to entities, events, relationships, and fact nodes in the graph.
[0132] Specifically, the system writes the merged entity nodes, event nodes, and topic nodes into the graph master database, and writes strongly credible facts, weakly credible facts, facts to be confirmed, and conflicting facts into the fact relation database according to their different fact states.
[0133] A fact node must include at least the fact number, subject node number, object node number, relationship type, fact state, confidence label, source chain number, evidence chain number, first discovery time, and most recent update time. If the fact involves an event, it may further include the event type, event time, event location, participating entities, amount field, status field, and version field.
[0134] Source chains are used to record the propagation path of facts from the original source to the source that reproduced or cited them.
[0135] For example, if a fact first comes from a government announcement, is then cited by a news website, and subsequently reprinted by an industry platform, the source chain records the source node of the government announcement, the source node of the news website, the source node of the industry platform, and the reprint edges between them.
[0136] A chain of evidence is used to record the document, fragment, table cell, attachment page number, or database field location corresponding to the facts.
[0137] For example, the evidence chain for a certain winning bid amount can record the URL of the winning bid announcement, page 3 of the PDF, row 2, column 5 of the table, and the amount field. Timestamps are used to record the intelligence release time, event occurrence time, system collection time, first entry into the data map time, and most recent update time.
[0138] Confidence labels are used to indicate the level of credibility of candidate facts in the current graph state, such as high credibility, medium credibility, low credibility, pending confirmation, or conflict.
[0139] When querying the graph, the system returns not only entities and relationships, but also fact states, source chains, evidence chains, and confidence labels.
[0140] For example, when a user queries "whether company A won the bid for project B", the system returns the bidding relationship between company A and project B, as well as the source of the bidding announcement, the paragraph of the announcement, the amount field, the release time, and the confidence score corresponding to this relationship.
[0141] If there are multiple versions of the same project amount, the system returns multiple candidate amount facts and marks their respective source chains, evidence chains and conflict types.
[0142] Through this graph generation method, the resulting open-source intelligence knowledge graph is no longer just a simple set of entity relationships, but a traceable graph that includes the source of facts, the location of evidence, the state of trust, and the state of conflict.
[0143] In S8, when new open-source intelligence is added, the system determines a local subgraph based on the candidate facts affected by the new intelligence, the set of conflicting facts, and the candidate relationships of the same reference, and performs incremental updates on the local subgraph to obtain the updated open-source intelligence knowledge graph.
[0144] Specifically, after receiving new open-source intelligence, the system identifies the source nodes, document nodes, entity nodes, event nodes, evidence nodes, and candidate fact relationships involved, and determines whether the new open-source intelligence hits existing candidate facts, conflicting fact sets, or identical candidate relationships.
[0145] If the newly added open-source intelligence only involves entirely new entities and events, then new local nodes and edges are created and written into the graph; If the newly added open-source intelligence hits an existing candidate fact, then update the source chain, evidence chain, number of supporting sources, and confidence score of that candidate fact; If the newly added open-source intelligence hits an existing set of conflicting facts, the confidence score and fact status of each conflicting fact node in that set of conflicting facts will be updated first. If newly added open-source intelligence hits a candidate relationship with the same name, the same-name judgment score is recalculated, and a decision is made based on the new score whether to merge or continue to retain the suspected same-name relationship. If newly added open-source intelligence provides contradictory evidence to existing strongly credible facts, then a new conflicting fact node is created and written into the corresponding conflicting fact set, instead of directly overwriting existing facts.
[0146] The system determines the local subgraph based on the affected nodes and their one-hop or two-hop neighborhoods.
[0147] The affected nodes include new source nodes, new document nodes, new evidence nodes, new candidate fact nodes, entity nodes associated with new facts, event nodes, time nodes, location nodes, as well as related conflicting fact nodes and related candidate nodes.
[0148] One-hop neighborhood is used to obtain nodes that are directly connected to the affected node, while two-hop neighborhood is used to obtain evidence, source, or event nodes that are further connected to the directly adjacent node.
[0149] The system only re-executes the HGT representation update on the local subgraph to obtain the locally updated node representation and candidate fact representation, and recalculates the fact confidence, conflict state and same-point judgment score based on the locally updated representation.
[0150] For graph regions that are not affected by newly added open-source intelligence, the system maintains their node representations, fact states, and confidence labels unchanged, thereby avoiding the repeated construction of the entire knowledge graph.
[0151] Example 1: To verify the feasibility of this invention in practice, it was applied to a risk intelligence analysis scenario for public projects in a high-tech zone of a certain city. This scenario requires continuous fusion and analysis of publicly available information regarding target company A's participation in public construction project B. The analysis aims to determine whether target company A is genuinely involved in the project, whether the project amount is accurate, whether project abbreviations from different public sources refer to the same project, and whether the addition of new public information alters the credibility of existing facts. The main problem in this scenario is that there are many sources of intelligence in public channels. Information about the same company or project may appear in public resource trading platforms, government procurement announcements, corporate credit information, publicly available news on corporate websites, news webpages, industry websites, and public databases. However, the descriptions of project names, project amounts, publishing entities, company abbreviations, and project status from different sources are not entirely consistent. Some public reports abbreviate the full project name to "Security Data Platform Project," some industry websites use approximate project amounts, some reprinted content lacks links to the original announcements, and some forum content provides unverified descriptions of project progress. If we rely directly on keyword retrieval or ordinary triple extraction, it is easy to split the same project into multiple project nodes, and it may also write the incorrect amount in the reprinted report into the knowledge graph.
[0152] In this embodiment, the system collected a total of 138 pieces of heterogeneous open-source intelligence from multiple sources, focusing on target company A and public construction project B. These included 19 government procurement announcements, 14 pieces of information from public resource trading platforms, 11 pieces of enterprise credit information, 9 pieces of news published on the company's official website, 27 news web pages, 18 articles from industry websites, 8 court judgments, 5 administrative penalty announcements, 17 public database records, and 10 public forum posts. Of the above data, 78 were HTML web pages, 24 were PDF announcements and attachments, 19 were structured database records, 11 were web page tables, and 6 were unstructured text. Each piece of data, upon entering the system, recorded its source address, source type, publication order, collection batch, document number, and original content summary. Project winning bid announcements published on public resource trading platforms are recorded as official announcement sources, with document number set as DOC-A001; project interpretation articles published on an industry website are recorded as industry website sources, with document number set as DOC-A014; reprinted reports published on a news website are recorded as news media sources, with document number set as DOC-A021; project payment records in a public database are recorded as authoritative database sources, with document number set as DOC-A087.
[0153] In this scenario, the system performs object identification on publicly available intelligence. For government procurement announcements and public resource trading platform pages, the system extracts project name, project number, purchaser, winning bidder, winning bid amount, announcement batch, and attachment page number; for enterprise credit information, the system extracts enterprise name, unified social credit code, registered address, legal representative, and registration status; for news webpages and industry website articles, the system extracts title, body paragraphs, reprint relationships, and related enterprises, projects, amounts, and locations; for publicly available database records, the system extracts field names, field values, update batches, and record numbers. After identification, the system obtained a total of 138 source objects, 138 document objects, 742 fragment objects, 326 entity objects, 104 event objects, 156 time-related objects, 61 location objects, 472 evidence objects, and 31 subject objects. Among them, 47 entity objects, 19 event objects, and 86 evidence objects are directly related to target enterprise A and public construction project B.
[0154] In practical applications, the system transforms these objects into a heterogeneous intelligence graph. Sources, documents, fragments, entities, events, time-related objects, locations, evidence, and topics are treated as different types of nodes, and different types of edges are established based on candidate relationships such as publication, inclusion, mention, participation, occurrence, location, support, reposting, conflict, and attribution. In this scenario, an initial heterogeneous graph with 2137 nodes and 4216 edges is generated. Specifically, there are 138 source nodes, 138 document nodes, 742 fragment nodes, 326 entity nodes, 104 event nodes, 156 time-related nodes, 61 location nodes, and 472 evidence nodes. Instead of directly writing the extracted facts into the final knowledge graph, the system first forms candidate fact nodes. For example, if the announcement on the public resource trading platform states that "Target Company A won the bid for public construction project B with a bid amount of 4.968 million yuan", the system will record this fact as candidate fact F-001, with the subject being target company A, the relationship being winning the bid, the object being public construction project B, the amount being 4.968 million yuan, and the evidence location being the 5th paragraph on page 2 of the PDF attachment and the 6th column of the 4th row of the attachment table.
[0155] For candidate facts, the system generates source-side credibility features, evidence-side credibility features, and cross-source consistency features. Official announcements, public resource trading platforms, and public database records are assigned a source level of 3; news media, industry websites, and corporate website news are assigned a source level of 2; and forum posts, anonymous comments, and reprinted content lacking an original source are assigned a source level of 1. Taking candidate fact F-001 as an example, its source is the public resource trading platform, with a source level of 3, belonging to the original announcement source, and a reprint chain length of 0. The evidence for this fact can be located in the PDF attachment page number, announcement paragraph, and table cell, and the evidence-side credibility feature is marked as complete. The system also finds descriptions of target company A participating in public construction project B in government procurement announcements, public resource trading platform records, and corporate website project news. Since the three sources are independent, the cross-source consistency feature is marked as highly consistent. The system marks the credibility description of candidate fact F-001 as highly credible.
[0156] Within the same batch of data, the system also identified candidate facts for news reposts (F-014) stating the project amount was "over 8 million yuan"; candidate facts for industry websites (F-021) stating the project amount was "approximately 5 million yuan"; and candidate facts for forum posts (F-038) stating the project "has not yet been confirmed as awarded." Further investigation revealed that F-014 could not be located at the original announcement amount field, only at the 6th paragraph of the news article, and this news content was cited from another industry website; F-021 could be located at a paragraph in an industry website article, but lacked a project number and announcement attachments; and F-038 originated solely from a forum post, without supporting contracts, announcements, or database records. Therefore, the system marked F-014 as a conflict pending determination, F-021 as moderately credible, and F-038 as a low-credibility clue, rather than directly deleting these facts or allowing later news content to overwrite the official announcement amount.
[0157] The system inputs heterogeneous intelligence graphs, node types, edge types, time-related information, and candidate fact credibility descriptions into the HGT heterogeneous graph Transformer model. In this embodiment, the model has 128 hidden dimensions, 4 attention heads, 2 stacking layers, and 2 rounds of node representation updates. The system generates credibility gating values based on the candidate fact credibility descriptions and writes these gating values into heterogeneous attention channels for different edge types. For evidence-supporting edges, complete evidence and direct evidence correspond to a high gating range of 0.80 to 1.00; partially complete evidence corresponds to a medium gating range of 0.50 to 0.79; and weak evidence lacking the original path corresponds to a low gating range of 0.10 to 0.49. For source-reprinted edges, the original source and short reprinted chain sources have higher propagation contributions, while sources with reprinted chain lengths exceeding level 2 and whose original source cannot be identified have reduced propagation contributions. For candidate fact conflict edges, the system does not simply suppress all conflicting facts but determines their re-evaluation value based on source level, evidence completeness, and cross-source consistency. In this way, official announcements and complete evidence have a strong influence on the representation of candidate facts, and content with long reposting chains, incomplete evidence, and high frequency of historical conflicts will not be mistakenly amplified during the dissemination process.
[0158] In cross-source same-reference judgment, the system encountered a typical problem: the public resource trading platform uses the full name of the project "Public Construction Project B", the company's official website news writes "Security Data Fusion Platform Project", and the news reprint content writes "City Security Platform Project". If only name matching is performed, the three names are similar but not completely identical, and can easily be split into multiple project nodes. The system uses name similarity only as a candidate recall condition, and further compares the project number, procuring unit, target company A, announcement batch, project location, and candidate fact representation output by HGT. The system found that the full project name and the company's official website abbreviation are associated with the same procuring unit, the same target company A, the same project number, and the same project location. The candidate fact representation similarity is 0.88, and the same-reference judgment score is 0.86, which is higher than the merging threshold of 0.80. Therefore, the two are merged into the same project node. For the "City Security Platform Project" in the news reprint, since the project number is missing, but the associated company, location, and project semantics are consistent, the candidate fact representation similarity is 0.74, and the same-reference judgment score is 0.67. The system retains it as a suspected same-name relationship instead of directly merging it.
[0159] Regarding the handling of conflicting facts, the system established a set of conflicting facts based on the project's monetary amount. The official announcement's figure of 4.968 million yuan corresponds to conflicting fact node C-001, the news reprint's "over 8 million yuan" corresponds to conflicting fact node C-002, and the industry website's "approximately 5 million yuan" corresponds to conflicting fact node C-003. The system writes C-001, C-002, and C-003 into the same set of conflicting facts and binds them to the source chain and evidence chain, respectively. Calculations show that C-001 has a source level of 3, complete evidence, high cross-source consistency, and a confidence score of 0.94, thus being marked as a strongly credible fact; C-002 has a source level of 2, but incomplete evidence and conflicts with the highly credible fact, with a confidence score of 0.33, thus being marked as a conflicting fact; C-003 has a source level of 2, partially complete evidence, the amount is an approximation and close to 4.968 million yuan, with a confidence score of 0.66, thus being marked as a weakly credible fact. When querying the project amount using graph analysis, the system prioritizes returning 4.968 million yuan, while retaining two other candidate amounts along with their source chains, evidence chains, and conflict types.
[0160] To demonstrate the beneficial effects of this invention, the system was statistically analyzed in this scenario. Using a combination of standard keyword retrieval and manual processing, it took an average of 7.5 hours for two business personnel to process the same batch of 138 publicly available information items, with approximately 3.2 hours spent on locating the original source and verifying monetary conflicts. After adopting the method of this invention, the entire process from data collection to generating a knowledge graph with source chains, evidence chains, and conflict states took 46 minutes. This included 11 minutes for intelligence object identification, 8 minutes for heterogeneous graph construction, 13 minutes for HGT credibility perception and propagation, 9 minutes for same-reference judgment and conflict fact processing, and 5 minutes for graph writing and verification. The final data generated 281 confirmed entity nodes, 88 confirmed event nodes, and 296 candidate facts, including 132 strongly credible facts, 61 weakly credible facts, 38 facts to be confirmed, 34 conflicting facts, 31 suspected same-name relationships, 19 conflicting fact sets, 214 source chains, and 267 evidence chains.
[0161] Regarding accuracy, 120 facts were manually sampled and verified. The conventional triplet sampling method, without considering source credibility and the chain of evidence, incorrectly included "over 8 million yuan" in the main fact result from news reprints and split "Security Data Fusion Platform Project" and "Public Construction Project B" into two project nodes, achieving a fact accuracy rate of 86.7%. Using the method of this invention, the system marked the official announcement amount of 4.968 million yuan as a strongly credible fact, marked "over 8 million yuan" as a conflicting fact, and correctly merged the project abbreviations from the company's official website into the same project node. Manual verification yielded 114 correct facts, achieving a fact accuracy rate of 95.0%. Regarding traceability, the conventional method could only provide original source links for 67 of the 120 sampled facts, with an evidence locatability rate of 55.8%. Using the method of this invention, 112 of the 120 sampled facts could be located to the original webpage, PDF page number, announcement paragraph, table cell, or database field, achieving an evidence locatability rate of 93.3%.
[0162] Regarding incremental updates, the system newly collected project payment information released by the fiscal disclosure platform. This information shows that the project number matches the announcement, the payment recipient is target company A, and the payment amount is 4.968 million yuan. The system identified that this new intelligence hit the existing project nodes, company nodes, and set of conflicting amount facts, and only updated the affected one-hop and two-hop local subgraphs. This local subgraph contains 1 new source node, 1 new document node, 3 new evidence nodes, 1 existing company node, 1 existing project node, 3 existing conflicting amount fact nodes, and 46 related source and evidence nodes, instead of recalculating the full 2137 nodes and 4216 edges. The local update took 2 minutes and 18 seconds. The confidence score of the official amount fact C-001 increased from 0.94 to 0.97, the confidence score of the news reprinted amount fact C-002 decreased from 0.33 to 0.27, and the confidence score of the industry website approximate amount fact C-003 was adjusted from 0.66 to 0.62. This demonstrates that the present invention can quickly update the affected factual state after new authoritative evidence is added, while keeping the unaffected spectral regions unchanged.
[0163] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for constructing a knowledge graph that integrates heterogeneous open-source intelligence from multiple sources, characterized in that, Includes the following steps: S1. Obtain multi-source heterogeneous open-source intelligence and identify intelligence objects to obtain a set of open-source intelligence objects including source, document, fragment, entity, event, time, location, evidence and topic; S2. Construct a heterogeneous intelligence graph based on the set of open source intelligence objects, convert different intelligence objects into different types of nodes, and establish different types of edges based on the relationships of publication, inclusion, mention, participation, occurrence, location, support, reprint, conflict, and same reference, to obtain the initial graph of multi-source heterogeneous open source intelligence. S3. Generate source-side credibility features, evidence-side credibility features, and cross-source consistency features for candidate fact relationships, and fuse them to obtain a candidate fact credibility description; S4. Input the initial graph of multi-source heterogeneous open source intelligence, node type, edge type, time information and candidate fact credibility description into the HGT heterogeneous graph Transformer model, perform credibility-aware heterogeneous intelligence propagation, and obtain entity representation, event representation, evidence representation and candidate fact representation; S5. Based on the candidate fact representation and the consistency of the evidence neighborhood, perform same-reference judgment on cross-source entity aliases, organization abbreviations, project codes and event names to obtain cross-source merge nodes; S6. Based on the candidate fact representation, candidate fact credibility description, and evidence node neighborhood, coexistence retention, confidence scoring, and state labeling are performed on contradictory candidate facts to obtain the fact state results. S7. Generate an open-source intelligence knowledge graph based on cross-source merged nodes and fact state results, and bind source chains, evidence chains, timestamps, and confidence labels to entities, events, relationships, and fact nodes; S8. When new open-source intelligence is added, a local subgraph is determined based on the candidate facts it affects, the set of conflicting facts, and the candidate relationships of the same reference, and incremental updates are performed to obtain the updated open-source intelligence knowledge graph.
2. The knowledge graph construction method for integrating multi-source heterogeneous open-source intelligence according to claim 1, characterized in that, S2 includes the following steps: S21. Convert the sources, documents, fragments, entities, events, times, locations, evidence, and topics in the open-source intelligence object set into heterogeneous nodes respectively; S22. Establish basic intelligence association edges based on the publication relationship between the source and the document, the inclusion relationship between the document and the fragment, and the reference relationship between the fragment and the entity; S23. Based on the participation relationship between entities and events, the occurrence relationship between events and time, and the location relationship between events and places, establish semantic association edges for events; S24. Based on the supporting relationships between evidence and candidate facts, the reprinting relationships between sources, the conflicting relationships between candidate facts, and the co-referenced candidate relationships between entities, establish credibility association edges; S25. Write node type identifiers for heterogeneous nodes and edge type identifiers for heterogeneous edges. Combine heterogeneous nodes, basic intelligence-related edges, event semantic-related edges, and credibility-related edges to obtain the initial graph of multi-source heterogeneous open-source intelligence.
3. The knowledge graph construction method for integrating multi-source heterogeneous open-source intelligence according to claim 1, characterized in that, S3 includes the following steps: S31. Determine the source node, document node, fragment node and evidence node corresponding to the candidate fact relationship to obtain the source evidence association structure; S32. Generate source-side credible features based on source type, publication time, and reposting path; S33. Generate credible features on the evidence side based on the supporting relationships between document nodes, fragment nodes, and evidence nodes; S34. Generate cross-source consistency features based on the degree of consistency in descriptions of the same candidate fact from multiple independent sources and historical conflict situations; S35. By integrating source-side credibility features, evidence-side credibility features, and cross-source consistency features, a credibility description of candidate facts is obtained.
4. The knowledge graph construction method for integrating multi-source heterogeneous open-source intelligence according to claim 3, characterized in that, S35 includes the following steps: S351. Set the source-side credibility features corresponding to official announcements, authoritative databases and original corporate disclosure documents as high credibility features, and set anonymous content and reprinted content lacking original sources as low credibility features. S352. When candidate facts can be located to the original document, paragraph, table cell or attachment page number, the credibility of the evidence side is improved. S353. When candidate facts are supported by two or more independent sources and the entity, event, time and location descriptions are consistent, improve the cross-source consistency feature; S354. When a candidate fact is supported by only a single source, has a referencing relationship, or has a high frequency of historical conflicts, reduce the cross-source consistency feature. S355. Generate candidate fact credibility descriptions for adjusting HGT heterogeneous attention based on the above features.
5. The knowledge graph construction method for integrating multi-source heterogeneous open-source intelligence according to claim 1, characterized in that, S4 includes the following steps: S41. Read the heterogeneous nodes, heterogeneous edges, node types, edge types, time information, and candidate fact credibility descriptions from the initial graph of multi-source heterogeneous open-source intelligence. S42. Set the type mapping parameters for source nodes, document nodes, fragment nodes, entity nodes, event nodes, time nodes, location nodes, evidence nodes, and topic nodes according to node type; S43. Set the attention parameters for publishing edges, containing edges, mentioning edges, participating edges, occurring edges, located edges, supporting edges, reprinting edges, conflicting edges, and same-pointing candidate edges according to edge type; S44. Generate credibility gating values based on the credibility descriptions of candidate facts, and adjust the heterogeneous attention channels of different edge types; S45. Aggregate information based on type mapping parameters, attention parameters, and adjusted attention channels, and output entity representation, event representation, evidence representation, and candidate fact representation.
6. The knowledge graph construction method for integrating multi-source heterogeneous open-source intelligence according to claim 5, characterized in that, S44 includes the following steps: S441. Obtain the center node, adjacent nodes, and edge types participating in the current heterogeneous attention calculation; S442. Read the source-side credibility features, evidence-side credibility features, and cross-source consistency features of the candidate facts corresponding to adjacent nodes; S443. Generate a confidence gating value based on the above features; S444. When the edge type is evidence-supporting edge or co-pointing candidate edge, the credibility gating value is used to enhance the propagation contribution of high-credibility evidence nodes and high-consistency candidate nodes. S445. When the edge type is a source-reprinted edge or a candidate fact conflict edge, use the credibility gating value to suppress the propagation contribution of neighboring nodes with long reprinted chains, missing evidence, or serious historical conflicts. S446. Write the adjustment results into the heterogeneous attention channel of the corresponding edge type, so that the candidate fact representation carries heterogeneous neighborhood information, evidence support information and source credibility information.
7. The knowledge graph construction method for integrating multi-source heterogeneous open-source intelligence according to claim 1, characterized in that, S5 includes the following steps: S51. Obtain the cross-origin entity alias, organization abbreviation, project code, or event name to be judged; S52. Generate initial same-reference judgment results based on name similarity, context fragment similarity, and candidate fact representation; S53. Obtain the evidence node, source node, related entity neighborhood, related event neighborhood, time information and location information corresponding to the object to be judged; S54. Correct the initial same-reference judgment result based on the neighborhood consistency of the evidence node and the source node to obtain the same-reference judgment score; S55. When the same index judgment score reaches the merging threshold, the objects to be judged are merged into the same graph node; otherwise, they are retained as different graph nodes and a suspected same-name relationship is established.
8. The method for constructing a knowledge graph that integrates multi-source heterogeneous open-source intelligence according to claim 1, characterized in that, S6 includes the following steps: S61. Detect whether there are candidate facts under the same entity or the same event that have different responsible persons, different addresses, different cooperative relationships, different times of occurrence, or different amounts; S62. When there are contradictory candidate facts, establish conflicting fact nodes and retain them in a coexisting manner. S63. Obtain the candidate fact representation, candidate fact credibility description and evidence node neighborhood for each conflicting fact node; S64. Calculate the confidence score based on the candidate fact representation, the candidate fact credibility description, and the evidence node neighborhood; S65. Based on the confidence score, mark the candidate facts as strongly credible facts, weakly credible facts, facts to be confirmed, or conflicting facts.
9. The knowledge graph construction method for integrating multi-source heterogeneous open-source intelligence according to claim 8, characterized in that, S62 includes the following steps: S621. Create conflict fact nodes for the contradictory candidate facts and write the subject conflict, address conflict, relationship conflict, time conflict or numerical conflict identifiers. S622. Establish association edges between conflict fact nodes and corresponding entity nodes, event nodes, source nodes, document nodes, fragment nodes, and evidence nodes; S623. Write mutually contradictory conflicting fact nodes into the same conflicting fact set, and retain the source chain and evidence chain corresponding to each conflicting fact node; S624. When performing a graph query, return multiple candidate facts and their conflict types, source chains, evidence chains, and confidence scores.
10. The method for constructing a knowledge graph that integrates multi-source heterogeneous open-source intelligence according to claim 1, characterized in that, S8 includes the following steps: S81. Receive new open source intelligence and identify the source nodes, document nodes, entity nodes, event nodes, evidence nodes and candidate fact relationships involved. S82. Determine whether the newly added open-source intelligence affects existing candidate facts, conflicting fact sets, or identical candidate relationships; S83. When there is an impact, determine the local subgraph based on the affected node and its one-hop or two-hop neighborhood. S84. Perform HGT representation update on the local subgraph to obtain the locally updated node representation and candidate fact representation; S85. Recalculate the fact confidence and conflict status based on the locally updated candidate fact representation; S86. Write the updated results back to the open-source intelligence knowledge graph, while keeping the unaffected graph regions unchanged.