Multi-source heterogeneous fair competition review knowledge graph construction method and system
Patent Information
- Application Number
- CN202611123753.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-28
- Publication Date
- 2026-09-25
AI Technical Summary
一方面,不同来源记录对同一事项的表述可能不一致,若直接合并,容易削弱图谱结果的可靠性;另一方面,新增审查数据持续进入后,若仍采用全量重建方式,不仅处理开销较大,而且不利于保持图谱结构的连续性和质量状态的一致性
[0021]本发明的有益效果在于:本发明针对公平竞争审查数据来源分散、格式不一、表述冲突和持续更新等问题,先将多源异构审查数据统一组织为带来源约束的证据对象,再依次完成语义规整、实体归并、关系组织、审查逻辑链构建、事实分层处理和多层知识图谱构建,使不同来源的数据能够在统一结构下完成关联表达。通过将来源可信权重、完整度、结构纯度、链支撑度和质量状态贯穿于图谱构建全过程,既保留了事实来源和冲突差异,又提高了知识组织结果的一致性、可追溯性和可维护性。进一步地,结合受影响子图判定和局部重建机制,可在新增数据进入时完成增量更新,减少全量重构带来的重复处理。
Smart Images

Figure CN122817486A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer information processing and knowledge graph construction technology, specifically involving a method and system for constructing a multi-source heterogeneous fair competition review knowledge graph. Background Technology
[0002] Data related to fair competition reviews is typically scattered across review documents, announcements, statements, forms, and other business documents. Data from different sources exhibits significant differences in format, field organization, expression, and granularity. The same review matter often corresponds to multiple records from different sources, at different times, and with different content focuses, including both structured fields and semi-structured or unstructured text. Existing processing methods largely rely on manual organization, keyword retrieval, or extraction methods targeting single data sources. These methods struggle to clean, align, merge, and correlate multi-source data under a unified standard, leading to low efficiency in subsequent knowledge organization and issues such as fragmented facts, difficulty in eliminating duplicate records, and challenges in tracing the source.
[0003] Meanwhile, existing knowledge graph construction methods typically focus more on entity recognition and relation extraction, neglecting the common challenges in review scenarios such as source differences, conflicting evidence, temporal changes, and the need for partial updates. On the one hand, records from different sources may offer inconsistent descriptions of the same matter; directly merging these could weaken the reliability of the graph results. On the other hand, with the continuous influx of new review data, using a full-scale reconstruction approach not only incurs significant processing overhead but also undermines the continuity of the graph structure and the consistency of its quality. Summary of the Invention
[0004] This invention provides a method and system for constructing a multi-source heterogeneous fair competition review knowledge graph, which solves the technical problems in the background art.
[0005] This invention provides a method for constructing a multi-source heterogeneous fair competition review knowledge graph, comprising the following steps:
[0006] Step 1: Obtain multi-source heterogeneous review data in the field of fair competition review, perform unified text processing and cleaning on the multi-source heterogeneous review data, construct a source-constrained evidence package set, and configure source credibility weights for the source-constrained evidence packages in the source-constrained evidence package set;
[0007] Step 2: Semantically segment the source constraint evidence package set to generate a set of structured semantic units, and generate a set of source constraint features based on the source credibility weight and the completeness and structural purity of each structured semantic unit;
[0008] Step 3: Based on the set of structured semantic units and the set of source constraint features, generate a set of candidate entity clusters and a set of directed relation units;
[0009] Step 4: Based on the candidate entity cluster set and the directed relation unit set, construct the review logic chain and determine the set of review logic chains that can be merged;
[0010] Step 5: Merge and classify the relational facts in the set of fusionable review logic chains to generate a core fact set, a supporting fact set, and a conflict record set;
[0011] Step 6: Based on the core fact set, the circumstantial fact set, the conflict record set, and the source constraint evidence package set, construct a multi-layer knowledge graph with quality status;
[0012] Step 7: Obtain the newly added source constraint evidence package set. Based on the newly added source constraint evidence package set and the multi-layer knowledge graph with quality state, determine the affected subgraph. Re-execute steps 2 to 6 on the affected subgraph to obtain the incrementally updated knowledge graph.
[0013] This invention also provides a multi-source heterogeneous fair competition review knowledge graph construction system, including:
[0014] The evidence encapsulation module is used to acquire multi-source heterogeneous review data in the field of fair competition review, perform unified text processing and cleaning on the multi-source heterogeneous review data, construct a source-constrained evidence package set, and configure source credibility weights for the source-constrained evidence packages in the source-constrained evidence package set.
[0015] The semantic regularization module is used to perform semantic segmentation on the source constraint evidence package set, generate a set of structured semantic units, and generate a set of source constraint features based on the source credibility weight and the completeness and structural purity of each structured semantic unit.
[0016] The entity relationship module is used to generate a candidate entity cluster set and a directed relationship unit set based on the set of structured semantic units and the set of source constraint features;
[0017] The link construction module is used to construct review logic chains based on the candidate entity cluster set and the directed relation unit set, and to determine the set of review logic chains that can be merged.
[0018] The fact classification module is used to merge and classify the relational facts in the set of fusionable review logic chains, and generate a core fact set, a supporting fact set, and a conflict record set.
[0019] The graph quality control module is used to construct a multi-layered knowledge graph with quality status based on the core fact set, the circumstantial fact set, the conflict record set, and the source constraint evidence package set.
[0020] The incremental update module is used to obtain the newly added source constraint evidence package set, determine the affected subgraph based on the newly added source constraint evidence package set and the multi-layer knowledge graph with quality state, and re-execute steps 2 to 6 on the affected subgraph to obtain the incrementally updated knowledge graph.
[0021] The beneficial effects of this invention are as follows: Addressing the issues of fragmented data sources, inconsistent formats, conflicting expressions, and continuous updates in fair competition review data, this invention first unifies multi-source heterogeneous review data into evidence objects with source constraints. Then, it sequentially completes semantic normalization, entity merging, relation organization, review logic chain construction, factual hierarchical processing, and multi-layered knowledge graph construction, enabling data from different sources to achieve correlated expression under a unified structure. By integrating source credibility weight, completeness, structural purity, chain support, and quality status throughout the entire graph construction process, it preserves the differences in factual sources and conflicts while improving the consistency, traceability, and maintainability of the knowledge organization results. Furthermore, by combining affected subgraph determination and local reconstruction mechanisms, incremental updates can be completed when new data enters, reducing redundant processing caused by full reconstruction. Attached Figure Description
[0022] Figure 1 This is a flowchart of the multi-source heterogeneous fair competition review knowledge graph construction method of the present invention;
[0023] Figure 2 This is a flowchart illustrating the process of generating a candidate entity cluster set and a directed relation unit set according to the present invention.
[0024] Figure 3 This is a schematic diagram of the modules of the multi-source heterogeneous fair competition review knowledge graph construction system of the present invention. Detailed Implementation
[0025] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.
[0026] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of the present invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in one or more embodiments of the present invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0027] like Figure 1 As shown, the method for constructing a multi-source heterogeneous fair competition review knowledge graph includes the following steps:
[0028] Step 1: Obtain multi-source heterogeneous review data in the field of fair competition review, perform unified text processing and cleaning on the multi-source heterogeneous review data, construct a source-constrained evidence package set, and configure source credibility weights for the source-constrained evidence packages in the source-constrained evidence package set;
[0029] Step 2: Semantically segment the source constraint evidence package set to generate a set of structured semantic units, and generate a set of source constraint features based on the source credibility weight and the completeness and structural purity of each structured semantic unit;
[0030] Step 3: Based on the set of structured semantic units and the set of source constraint features, generate a set of candidate entity clusters and a set of directed relation units;
[0031] Step 4: Based on the candidate entity cluster set and the directed relation unit set, construct the review logic chain and determine the set of review logic chains that can be merged;
[0032] Step 5: Merge and classify the relational facts in the set of fusionable review logic chains to generate a core fact set, a supporting fact set, and a conflict record set;
[0033] Step 6: Based on the core fact set, the circumstantial fact set, the conflict record set, and the source constraint evidence package set, construct a multi-layer knowledge graph with quality status;
[0034] Step 7: Obtain the newly added source constraint evidence package set. Based on the newly added source constraint evidence package set and the multi-layer knowledge graph with quality state, determine the affected subgraph. Re-execute steps 2 to 6 on the affected subgraph to obtain the incrementally updated knowledge graph.
[0035] In one embodiment of the present invention, multi-source heterogeneous review data in the field of fair competition review is first acquired. The multi-source heterogeneous review data refers to a collection of data with different sources, formats, and structural levels, but all related to fair competition review matters. It includes at least announcement texts, review documents, situation reports, rule documents, tabular records, and other electronic text records. Because different data differ in source, format, and arrangement, directly entering the subsequent knowledge graph construction process could easily lead to problems such as difficulty in aligning the same matter, difficulty in referencing sources, and inconsistencies in subsequent semantic processing. Therefore, this embodiment first uses a single record as the processing object, splitting the multi-source heterogeneous review data into individual original review data, and determining the source category label and format label corresponding to each original review data.
[0036] The source category tag is used to characterize the source category of the original review data, and the format tag is used to characterize the carrier format of the original review data. Subsequently, a unified textualization process is performed on each piece of original review data according to the format tag, resulting in unified text content and position mapping information. The unified text content refers to the standardized text result after converting original content in different formats into a unified text expression; the position mapping information refers to the correspondence between text segments in the unified text content and corresponding positions in the original review data. For example, when the original review data is a tabular record, the position mapping information can correspond to the row and column positions in the table; when the original review data is a paragraph-style document, the position mapping information can correspond to the chapter or paragraph positions.
[0037] After obtaining the unified text content and location mapping information, the unified text content is cleaned. Cleaning refers to standardizing duplicate records, abnormal characters, irrelevant formatting residues, inconsistent numbering, and inconsistent text separation to obtain cleaned unified text content. Subsequently, the cleaned unified text content, along with source category markers, record time, record number, format markers, and location mapping information, is packaged to obtain a source-constrained evidence package. The source-constrained evidence package refers to a unified packaged object formed based on a single piece of original review data, which at least includes the text body, source identification information, time identification information, and location reference information.
[0038] In other words, the source-constrained evidence package is not simply a text record, but rather a data object that incorporates the text content, source information, and location tracing information required for subsequent knowledge graph construction. For example, two review records with similar content but different sources may have highly similar texts after cleaning, but due to differences in their source category labels, record times, or record numbers, they will still form different source-constrained evidence packages, thereby preventing data from different sources from being mistakenly identified as the same source of evidence.
[0039] After the source constraint evidence package is formed, a source credibility weight is further assigned to each source constraint evidence package. Specifically, the credibility of the source category is determined based on the source category label and the preset correspondence between source categories; the completeness of the metadata is determined based on the number of valid items in the record time, record number, and format label; and the location citation degree is determined based on the proportion of valid text fragments in the location mapping information that can be referenced back to the original review data to the total number of valid text fragments. The source category credibility reflects the credibility level of the source category itself, the metadata completeness reflects the completeness of the information in terms of time, number, and format of the source constraint evidence package, and the location citation degree reflects the sufficiency of the text content in the source constraint evidence package being able to be referenced back to the location of the original review data.
[0040] Then, the source category credibility, metadata completeness, and location backlinking are weighted to obtain the source credibility weight. This source credibility weight is a fundamental parameter directly invoked when generating the subsequent source constraint feature set. Its function is to embed the source layer's credibility constraints into the knowledge graph construction process in advance, rather than waiting until the graph is built before external correction. For example, if a source constraint evidence package has a clear source category, complete record time and record number, and most text fragments can be backlinked to the original review data through location mapping information, then the source credibility weight corresponding to this source constraint evidence package is high; conversely, if the source category is unstable, there is a lot of missing metadata, and location backlinking is insufficient, then its source credibility weight is relatively low.
[0041] After the above processing, the multi-source heterogeneous review data in the field of fair competition review is organized into a set of source-constrained evidence packages, and unified textualization, cleaning, encapsulation, and source credibility weight configuration are completed at the object level. This data organization method enables the subsequent processes of generating structured semantic units, entity merging, relation construction, and multi-layer knowledge graph construction to revolve around the same set of source-constrained evidence packages, reducing ambiguity between records from different sources, in different formats, and in different locations, and enhancing data consistency, reproducibility, and quality control capabilities in the knowledge graph construction process.
[0042] In one embodiment of the present invention, after constructing the source constraint evidence package set and configuring source credibility weights for each source constraint evidence package, the source constraint evidence package set is further semantically segmented to generate a set of structured semantic units. A set of source constraint features is then generated based on the source credibility weights and the completeness and structural purity of each structured semantic unit. The core of this process lies in transforming the source constraint evidence packages formed in the previous stage from evidence encapsulation objects oriented towards original records into unified semantic objects oriented towards knowledge extraction and relation construction. This ensures that subsequent candidate entity cluster generation, directed relation unit generation, and review logic chain construction are all based on a consistent data structure.
[0043] Specifically, the source constraint evidence packages within the source constraint evidence package set are first semantically segmented according to paragraph boundaries, title boundaries, table boundaries, and action phrase boundaries to obtain a set of semantic fragments. This semantic segmentation does not involve arbitrarily dividing the text into segments, but rather dividing the text content within the same source constraint evidence package into multiple separately processable semantic fragments based on boundary positions within the text that have independent semantic meaning. Paragraph boundaries are used to identify natural separation positions of continuous narrative content, title boundaries to identify topic transitions, table boundaries to identify structured record blocks, and action phrase boundaries to identify semantic breakpoints directly related to behavioral expression. Through this process, each semantic fragment maintains a relatively complete semantic direction, preventing multiple different matters from being mixed in the same processing unit. For example, in a review document, if the first paragraph describes the basic situation of the review object and the second paragraph describes the specific processing conclusion, the two paragraphs will form different semantic fragments after segmentation.
[0044] After obtaining the set of semantic fragments, the subject information, action information, object information, result information, basis information, time information, and region information are extracted from each semantic fragment. These are then encapsulated in the order of subject information, action information, object information, result information, basis information, time information, and region information to generate a set of structured semantic units. The structured semantic unit refers to a unified structured object that organizes the key semantic elements originally scattered in the semantic fragments according to a preset field order. Specifically, subject information represents the subject that initiates or carries out the action; action information represents the specific content of the action; object information represents the object on which the action affects; result information represents the conclusion or state produced by the action; basis information represents the basis supporting the action or conclusion; time information represents the corresponding time element; and region information represents the corresponding geographical element.
[0045] After the set of structured semantic units is formed, the number of valid information items in the subject information, action information, object information, result information, basis information, time information, and region information of each structured semantic unit is further counted. The completeness of each structured semantic unit is determined by the ratio of the number of valid information items to the total number of the seven information items. A valid information item refers to an information item that has identifiable, attributable, and non-empty content in the corresponding field. Completeness is used to characterize the degree of completeness of a single structured semantic unit in terms of semantic element coverage. When all seven information items in a structured semantic unit are identified and their content is clear, its completeness is high; when only some information items are identified, its completeness decreases accordingly. For example, if a semantic segment only contains subject information, action information, and object information, but lacks result information, basis information, time information, and region information, then the completeness of the structured semantic unit corresponding to this semantic segment is lower than that of a structured semantic unit containing all seven complete information items.
[0046] Subsequently, content tags that can be uniquely categorized into one of the following: subject information, action information, object information, result information, basis information, time information, and region information, are identified as valid content tags. The structural purity is determined by the ratio of the number of valid content tags to the total number of content tags. A content tag refers to the smallest unit of content that can participate in field attribution judgment after semantic fragments are segmented and normalized; a valid content tag refers to a content tag whose attribution field is unique and does not have cross-field ambiguity. Structural purity characterizes the clarity of content allocation within a structured semantic unit. When most of the content in a structured semantic unit can be stably categorized into a unique field, its structural purity is high; when a large amount of content is difficult to uniquely categorize, or when the same content competes for attribution in multiple fields, its structural purity decreases accordingly. For example, if the phrase "a decision was made on a certain day based on a certain rule" in a text can be stably split into basis information, time information, and result information, then the structural purity of this structured semantic unit is high; if related content is mixed and difficult to uniquely categorize into fields, then the structural purity is low.
[0047] After obtaining the completeness and structural purity of each structured semantic unit, the source credibility weight, completeness, and structural purity are combined according to their correspondence to generate a set of source constraint features that correspond one-to-one with each structured semantic unit. The correspondence refers to the source credibility weight corresponding to a structured semantic unit being taken from its respective source constraint evidence package; its completeness and structural purity are taken from the calculation results of the structured semantic unit itself. Thus, each structured semantic unit corresponds to a source constraint feature, which simultaneously characterizes the source credibility, semantic completeness, and internal structural clarity of the structured semantic unit. Through this processing, subsequent steps in merging candidate entity clusters and generating directed relation units can not only call the semantic fields of the structured semantic unit itself but also synchronously call its source constraint features, thereby unifying source-level constraints and semantic-level constraints into the same processing chain.
[0048] After the above processing, the source constraint evidence package set is further organized into a set of structured semantic units and a corresponding set of source constraint features. This result enables the subsequent intelligent knowledge graph construction process to perform entity recognition, relation organization, and graph structure generation at a unified semantic granularity, reducing ambiguity in field extraction, semantic attribution, and quality judgment of records from different sources, and enhancing data regularization capabilities and graph construction consistency in multi-source heterogeneous data fusion scenarios.
[0049] In one embodiment of the present invention, such as Figure 2 After generating the set of structured semantic units and the set of source constraint features, the process further generates a set of candidate entity clusters and a set of directed relation units based on these sets. The core of this process lies in transforming the structured semantic units formed in the previous stage from field-based semantic objects into entity organization results and relation organization results that can be used for knowledge graph modeling. This ensures that the subsequent review logic chain construction no longer directly deals with scattered fields, but rather with entity clusters and relation units that have already undergone merging constraints.
[0050] Specifically, firstly, the source credibility weight, completeness, and structural purity of each source constraint feature in the source constraint feature set are weighted to determine the confidence level of each structured semantic unit. The confidence level of a structured semantic unit refers to a unified evaluation metric obtained by comprehensively representing a single structured semantic unit in terms of source credibility, information completeness, and internal structural clarity. Subsequently, entity candidates corresponding to subject information, object information, basis information, and region information are extracted from each structured semantic unit. These entity candidates are field-based semantic items that can serve as the basis for subsequent entity merging, retaining their original slot source. In other words, at this stage, all entity content is not directly treated as independent entities at the same level, but rather maintained as entity candidates with slot semantic background, thereby avoiding premature merging of content with the same name in different semantic positions.
[0051] After obtaining entity candidates, for entity candidates with the same slot category, determine the degree of name consistency, context consistency, slot category consistency, and structured semantic unit confidence coupling. Name consistency reflects the similarity in the names of two entity candidates; context consistency reflects the similarity in adjacent semantic content; slot category consistency reflects whether two entity candidates belong to the same field category; and structured semantic unit confidence coupling is determined based on the smaller value of the structured semantic unit confidence scores for the two entity candidates. The smaller value is used to determine the structured semantic unit confidence coupling because the stable merging of two entity candidates is more significantly constrained by the weaker side. For example, if one source has high credibility but the other has low credibility, the merging credibility should be limited by the lower side, not solely by the higher side.
[0052] Subsequently, a weighted sum is calculated based on name consistency, context consistency, slot category consistency, and the coupling degree of structured semantic unit confidence to determine the merging score. Entity candidates with a merging score not lower than a preset merging threshold are grouped into the same candidate entity cluster, while the remaining entity candidates are placed into different candidate entity clusters, resulting in a candidate entity cluster set. A candidate entity cluster refers to a set of entities that, after merging, are deemed to point to the same entity object. This process does not simply merge entities based on identical names; instead, it incorporates name, context, slot category, and upstream source constraints into the same merging decision process. For example, the same name may appear in both subject information and object information; if the slot categories are different, even if the names are the same, they will not be directly grouped into the same candidate entity cluster. Through this process, entity merging no longer relies on single text similarity but is based on the combined effect of structured semantics and source constraints.
[0053] After the candidate entity cluster set is formed, the head entity cluster identifier and tail entity cluster identifier are determined according to the candidate entity cluster to which the subject information and object information in each structured semantic unit belong. Then, a directed relation unit set is generated by combining action information, result information, basis information, time information, and region information. The confidence level of the structured semantic unit corresponding to each directed relation unit is determined as the relation support level of each directed relation unit. The directed relation unit refers to a relational expression object with the head entity cluster identifier and tail entity cluster identifier as its two ends, action information as its relation core, and attached result information, basis information, time information, and region information. Since the subject information and object information have already completed the merging of candidate entity clusters, the relation generated by the same structured semantic unit no longer directly depends on the original text fields, but rather on the merged head entity cluster identifier and tail entity cluster identifier.
[0054] Finally, a unique correspondence check is performed on the head entity cluster identifier and the tail entity cluster identifier, and directed relation units that pass the check are retained. The unique correspondence check means confirming that both the head and tail entity cluster identifiers can find a unique and stable belonging result in the candidate entity cluster set. If the subject information or object information in a structured semantic unit corresponds to multiple candidate entity clusters, then that directed relation unit is not included in the retained results. Through this process, the final set of directed relation units can maintain a unified reference relationship with the candidate entity cluster set, avoiding the problem of uncertain relation endpoints in the subsequent review logic chain construction stage.
[0055] After the above processing, the set of structured semantic units and the set of source constraint features are further organized into a set of candidate entity clusters and a set of directed relation units. This processing enables the subsequent intelligent knowledge graph construction process to continue based on the entity merging results and relation merging results, reducing ambiguity in entity identification and relation connection of multi-source heterogeneous review data, and enhancing the consistency and computability of knowledge organization results.
[0056] In one embodiment of the present invention, after generating the candidate entity cluster set and the directed relation unit set, a review logic chain is further constructed based on the candidate entity cluster set and the directed relation unit set, and a set of review logic chains that can be merged is determined. This part organizes the discrete directed relation units formed in the previous stage into a continuously readable chain structure according to unified entity connection relationships and constraints, so that subsequent relation fact merging and conflict classification processing are no longer oriented towards isolated relation units, but towards review logic chain objects that already have contextual continuity.
[0057] Specifically, firstly, based on the candidate entity cluster set and the directed relation unit set, the tail entity cluster identifier of the preceding directed relation unit is the same as the head entity cluster identifier of the following directed relation unit, and the directed relation units with consistent information, time information, and region information are assembled sequentially to obtain the review logic chain set. The review logic chain refers to a chain-like relation object formed by connecting multiple directed relation units end-to-end according to the entity connection order. The tail entity cluster identifier is the same as the head entity cluster identifier of the following directed relation unit, ensuring that the relation units within the chain can be continuously connected at the entity level; the consistency of information, time information, and region information ensures that the relation units within the chain belong to the same review context. In other words, this embodiment does not mechanically chain all connectable directed relation units, but only retains connections that are simultaneously valid in terms of entity connection and contextual constraints. For example, although two directed relation units can be connected in entity clusters, if their corresponding time information is different, they will not be included in the same review logic chain.
[0058] After the review logic chain set is formed, the number of valid items in the starting entity cluster identifier, relation type sequence, ending entity cluster identifier, result information sequence, basis information, time information, and region information of each review logic chain is further counted. The number of valid items is used as the numerator, and the total number of the seven items is used as the denominator to obtain the logical completeness. The logical completeness is used to characterize the completeness of a single review logic chain in its chain structure. Valid items refer to items that have clear content at their corresponding positions and can be stably assigned. A review logic chain has a high logical completeness when it simultaneously possesses the starting entity cluster identifier, relation type sequence, ending entity cluster identifier, result information sequence, basis information, time information, and region information; when some items are missing, its logical completeness decreases accordingly.
[0059] Subsequently, the number of directed relation units consistent with the assembly direction in each review logic chain is counted. This number is used as the numerator, and the total number of directed relation units is used as the denominator to obtain the direction consistency. The direction consistency is used to characterize the degree of uniformity of the directed relation units within a review logic chain in the connection direction. At the same time, the relation support of each directed relation unit is summed as the numerator, and the total number of directed relation units is used as the denominator to obtain the chain support.
[0060] The chain support can be expressed as: ,in, Indicates the first The chain support of the review logic chain, Indicates the first A set of directed relational units linked by a logical chain of checks. Represents a directed relation unit Relationship support Indicates the first The total number of directed relation units associated with a review logic chain. The chain support is used to characterize the average support strength of directed relation units within a review logic chain. That is, directional consistency reflects the stability of the chain's connection direction, while chain support reflects whether the relation units within the chain have sufficient overall support. For example, although a review logic chain may have complete entity connections, if some directed relation units have directions opposite to the assembly direction, its directional consistency decreases; similarly, although another review logic chain may have consistent directions, if the relation support of several directed relation units is low, its chain support will also decrease accordingly.
[0061] After obtaining the logical completeness, directional consistency, and chain support, the logical completeness is multiplied by the preset weight corresponding to the logical completeness, the directional consistency is multiplied by the preset weight corresponding to the directional consistency, and the chain support is multiplied by the preset weight corresponding to the chain support. The three multiplications are then summed to obtain the fusion judgment value.
[0062] The fusion determination value can be expressed as: ,in, Indicates the first The fusion judgment value of the review logic chain, Indicates the first The logical integrity of each review logic chain. Indicates the first The consistency of the direction of the review logic chain. This represents the preset weight corresponding to the logical completeness. This represents the preset weight corresponding to the directional consistency. This represents the preset weight corresponding to the chain support level. The fusion judgment value is a unified result after comprehensively evaluating a single review logic chain in terms of structural completeness, directional consistency, and relational support. Then, review logic chains with fusion judgment values not lower than the preset fusion threshold are determined as the set of fusionable review logic chains, and the remaining review logic chains are determined as the set of review logic chains to be corrected.
[0063] After the above processing, the candidate entity cluster set and the directed relation unit set are further organized into a review logic chain set. Based on this, the logical completeness, directional consistency, chain support, and fusion judgment value are uniformly calculated, ultimately yielding a fusionable review logic chain set. This processing allows the subsequent intelligent knowledge graph construction process to continue at the chain-level objects, reducing the context breakage problem caused by discrete relation units directly participating in fact merging, and enhancing the continuity, consistency, and computability of knowledge organization results in multi-source heterogeneous data fusion scenarios.
[0064] In one embodiment of the present invention, after determining the set of fusionable review logic chains, the relational facts in the set of fusionable review logic chains are further merged and conflict-classified to generate a core fact set, a supporting fact set, and a conflict record set. The core of this process is to further compress the chain-level relational objects formed in the previous stage into fact-level objects, and to complete the merging of similar objects, version differentiation, and conflict layering within the fact-level objects. This ensures that the subsequent construction of the multi-layer knowledge graph no longer directly faces chain-like objects, but rather a unified set of facts that has already been classified and marked for conflict.
[0065] Specifically, the process first reads the starting entity cluster identifier, relationship type sequence, ending entity cluster identifier, result information sequence, basis information, time information, region information, fusion judgment value, and chain support degree corresponding to each fusionable review logic chain, and encapsulates them in a fixed order to obtain the relational facts. The relational facts refer to the factual expression objects extracted from a single fusionable review logic chain. These objects include both ends of the entity chain and the relational expression, as well as result information, basis information, time information, and region information. They also inherit the fusion judgment value and chain support degree corresponding to that fusionable review logic chain. In other words, the relational facts do not merely retain a simplified result of who has what kind of relationship with whom, but rather retain all the core contextual information required for subsequent factual judgment.
[0066] After obtaining the relational facts, relational facts with the same starting entity cluster identifier, relation type sequence, and ending entity cluster identifier are grouped into the same merging unit. The merging unit refers to an intermediate object for centralized management of multiple relational facts under the same relational skeleton, specifically defined by the starting entity cluster identifier, relation type sequence, and ending entity cluster identifier. That is, relational facts that are consistent in the three aspects of entity start point, relation expression, and entity end point are first considered as candidates of the same category, and then further differentiated into different versions within this category. Subsequently, the result information sequence, basis information, time information, and region information of the relational facts within the same merging unit are compared; those with all four being identical are determined to be the same fact version. The fact version refers to a group of relational facts under the same relational skeleton that have the same result information, the same basis, the same time, and the same region constraint. For example, two relational facts may have the same starting entity cluster identifier and ending entity cluster identifier, but if their time information is different, they are not considered the same fact version, but rather different fact versions within the same merging unit.
[0067] After the fact versions are determined, the support strength of the relational facts in each fact version is further calculated. Specifically, the fusion judgment value and chain support degree of each relational fact are multiplied and summed to obtain the version support value. The version support value refers to the cumulative support strength obtained by the same fact version, which reflects both the overall fusionability of the review logic chain on which the fact version depends and the degree of relational support within the chain. Unlike simply counting the number of relational facts, this embodiment incorporates both the fusion judgment value and the chain support degree into the version support value calculation process, so that the support degree of the fact version no longer depends on the quantity itself, but on the cumulative result after quality constraints.
[0068] Subsequently, the degree of conflict between fact versions is determined based on the proportion of the largest version's support value in the cumulative result of all version support values. The degree of conflict refers to the quantified result of the degree of disagreement between different fact versions within the same merging unit. A lower degree of conflict occurs when a fact version's support value significantly outweighs all version support values; a higher degree of conflict occurs when the version support values of multiple fact versions are close. In other words, the degree of conflict is not directly determined by the number of fact versions, but rather by the support distribution of each fact version.
[0069] After obtaining the version support value and conflict degree, further fact classification processing is performed. When the merging unit contains only one fact version, the corresponding relationship facts are assigned to the core fact set. The core fact set refers to the set of facts that are subsequently written directly as fact edges in the main graph. When the merging unit contains multiple fact versions and the conflict degree is not higher than a preset conflict threshold, the fact version with the highest version support value is assigned to the core fact set, and the rest are assigned to the supporting fact set. The supporting fact set refers to the set of facts that are not written as dominant facts into the core layer but are still retained as supporting reference facts. When the conflict degree is higher than the preset conflict threshold, all fact version relationship facts, along with their version support values and conflict degrees, are assigned to the conflict record set. The conflict record set is a set specifically recording highly divergent fact versions and their supporting distributions, used for subsequently constructing conflict record nodes and conflict relationships. For example, in the same merging unit, if the version support values of two fact versions are close, it indicates that there is not a sufficiently obvious dominant version within the merging unit. In this case, all fact versions enter the conflict record set instead of directly entering the core fact set.
[0070] After the above processing, the set of fusionable review logic chains is further organized into a core fact set, a supporting fact set, and a conflict record set. This processing enables the subsequent intelligent knowledge graph construction process to continue on fact-level objects, reducing the hierarchical confusion caused by directly adding chain-level objects to the graph. At the same time, by introducing version support values and conflict levels, relational facts from different sources, times, regions, and bases can be managed hierarchically under a unified framework, enhancing the consistency and computability of fact merging and conflict expression in multi-source heterogeneous data fusion scenarios.
[0071] In one embodiment of the present invention, after generating the core fact set, the supporting fact set, and the conflict record set, a multi-layered knowledge graph with quality status is further constructed based on the core fact set, the supporting fact set, the conflict record set, and the source constraint evidence package set. The core of this process lies in further organizing the hierarchical fact objects formed in the previous stage into a unified graph structure, and synchronously writing source backreference relationships, local quality status, and repair identifiers into this graph structure, so that subsequent incremental updates are no longer based on external detection results, but directly on the quality results within the graph. The fact objects include core fact edges, supporting fact edges, and conflict association edges.
[0072] Specifically, firstly, entity nodes, core fact edges, supporting fact edges, conflict record nodes, and conflict association edges are established based on the core fact set, supporting fact set, and conflict record set. The entity node refers to the graph node object formed by the starting entity cluster identifier and the ending entity cluster identifier involved in the core fact set and supporting fact set; the core fact edge refers to the main relation edge established between corresponding entity nodes by fact objects in the core fact set; the supporting fact edge refers to the auxiliary relation edge established between corresponding entity nodes by fact objects in the supporting fact set; the conflict record node refers to the node object that centrally carries the conflict facts in the conflict record set; and the conflict association edge refers to the association edge used to connect the conflict record node with the corresponding entity node or fact object.
[0073] After establishing entity nodes, core fact edges, corroborating fact edges, conflict record nodes, and conflict association edges, the corresponding source constraint evidence packages are determined from the set of source constraint evidence packages based on the basis information, time information, and regional information corresponding to each fact object. Source back-reference edges are then established to obtain a multi-layered knowledge graph. The source back-reference edge refers to the intra-graph reference relationship between a fact object and a source constraint evidence package. Its significance lies in enabling each fact edge or conflict association edge in the graph to reference back to the original evidence object supporting that fact, without relying on external tracing. For example, two core fact edges may have the same relationship type, but if their basis information and time information correspond to different source constraint evidence packages, their source back-reference edges will connect to different evidence objects respectively. Through this process, the multi-layered knowledge graph simultaneously carries information from the fact layer, conflict layer, and evidence layer.
[0074] After the multi-layered knowledge graph is formed, corresponding supporting fact edges, conflict record nodes, conflict association edges, and source referencing edges are extracted around each core fact edge to construct a local subgraph. A local subgraph refers to a local analysis object extracted around a single core fact edge, which includes at least the core fact edge itself and its directly related auxiliary relationships, conflict relationships, and evidence referencing relationships. Subsequently, completeness, consistency, traceability, and timeliness are calculated for each local subgraph. Completeness is determined by the proportion of fact objects that simultaneously possess a starting entity cluster identifier, a relationship type sequence, a terminating entity cluster identifier, a result information sequence, basis information, time information, regional information, and a source referencing edge to the total number of fact objects in the local subgraph. Consistency is determined by the proportion of fact objects that have not established associations with conflict record nodes to the total number of fact objects in the local subgraph. Traceability is determined by the proportion of fact objects with source referencing edges to the total number of fact objects in the local subgraph. Timeliness is determined by the proportion of fact objects whose recording time corresponding to the source constraint evidence package falls within a preset effective time window to the total number of fact objects in the local subgraph.
[0075] After obtaining the completeness, consistency, traceability, and timeliness, the completeness is multiplied by its corresponding preset weight, the consistency by its corresponding preset weight, the traceability by its corresponding preset weight, and the timeliness by its corresponding preset weight. The results are then summed to determine the quality status. This quality status is a unified result of a comprehensive evaluation of a single local subgraph in terms of structural completeness, conflict distribution, source referencing, and timeliness. Then, based on the comparison results of completeness, consistency, traceability, and timeliness with their respective preset thresholds, repair markers are determined. These repair markers are structured labels indicating the direction of subsequent corrections to the local subgraph. When completeness is below the corresponding preset threshold, the repair marker points to information supplementation; when consistency is below the corresponding preset threshold, the repair marker points to conflict reclassification; when traceability is below the corresponding preset threshold, the repair marker points to source link supplementation; and when timeliness is below the corresponding preset threshold, the repair marker points to local update. Finally, the quality status and repair identifier are written into the corresponding local subgraph to obtain a multi-layer knowledge graph with quality status.
[0076] After the above processing, the core fact set, the circumstantial fact set, the conflict record set, and the source-constrained evidence package set are further organized into a multi-layered knowledge graph with quality status. This processing enables the knowledge graph to not only express entity relationships and conflict relationships in multi-source heterogeneous review data, but also to store evidence sources, quality results, and repair directions within the graph, reducing the disconnect between graph construction and graph maintenance, and enhancing the integrated processing capabilities of this invention in graph structure organization, quality control, and subsequent incremental updates.
[0077] In one embodiment of the present invention, after constructing a multi-layered knowledge graph with quality states, a new set of source constraint evidence packages is further obtained. Based on the new set of source constraint evidence packages and the multi-layered knowledge graph with quality states, the affected subgraphs are determined. Steps 2 to 6 are then re-executed on the affected subgraphs to obtain an incrementally updated knowledge graph. The core of this process is that instead of performing a full reconstruction of the existing multi-layered knowledge graph, the scope of influence of the new evidence on the existing local subgraphs is first determined, and then reconstruction and write-back are performed only on the affected local regions. This ensures that the knowledge graph update process maintains locality, continuity, and traceability.
[0078] Specifically, firstly, a new set of source constraint evidence packages is obtained, and step 2 is re-executed on this new set to generate a new set of structured semantic units and a new set of source constraint features. Subject information, object information, basis information, and region information are then extracted from the new set of structured semantic units to generate a new set of entity references. The formation method of the new set of structured semantic units and the new set of source constraint features is consistent with the processing method of the original set of source constraint evidence packages in step 2, thereby ensuring that the new evidence and the existing knowledge graph are in the same expression system in terms of semantic granularity and constraint granularity. The new set of entity references refers to the set of reference objects extracted from the new structured semantic units that can be used for entity correspondence comparison with existing local subgraphs.
[0079] After obtaining the newly added set of structured semantic units, the newly added set of source constraint features, and the newly added set of entity references, the source overlap, entity overlap, time overlap, and repair trigger value are further determined for each local subgraph in the multi-layered knowledge graph with quality status. Specifically, the source overlap is determined by the proportion of newly added source constraint evidence packages that meet the source consistency condition; the entity overlap is determined by the proportion of newly added entity references that meet the entity consistency condition; the time overlap is determined by the proportion of newly added source constraint evidence packages whose recording time falls within the time range; and the repair trigger value is determined based on whether the repair identifier is empty. The source consistency condition refers to the consistency between the newly added source constraint evidence package and the existing source constraint evidence package associated with the local subgraph in terms of source category or source identifier; the entity consistency condition refers to the correspondence between the newly added entity reference and the existing entity node in the local subgraph in terms of subject information, object information, basis information, or region information; the time range refers to the time coverage interval formed by the recording time of the source constraint evidence package corresponding to the fact object in the local subgraph. The repair trigger value reflects whether the local subgraph has been marked as an object to be corrected in the previous stage of quality assessment.
[0080] Subsequently, based on the overlap of source, entity, time, repair trigger value, and quality status, a comprehensive evaluation is performed according to their respective preset weights to obtain the affected judgment value. Local subgraphs with affected judgment values not lower than a preset affected threshold are included in the affected subgraph set. The affected judgment value is a unified quantitative result of the degree of correlation between new evidence and existing local subgraphs. It considers not only the proximity of new evidence and local subgraphs in terms of source, entity, and time, but also the original quality status and repair needs of the local subgraph. In other words, the affected judgment in this embodiment does not directly trigger an update simply because new evidence exists; rather, it requires a sufficient degree of correlation between the new evidence and the local subgraph, or that the local subgraph itself is already in a state requiring priority correction. For example, although a local subgraph overlaps with new evidence in time, if there is no correlation in terms of source and entity, and the original quality status is high, its affected judgment value may not be sufficient to be included in the affected subgraph set.
[0081] After determining the set of affected subgraphs, for each affected subgraph, its associated set of source constraint evidence packages is combined with the newly added source constraint evidence package that triggered the affected subgraph to form a set of local update evidence packages. Steps 2 to 6 are then re-executed based solely on this set of local update evidence packages to obtain the reconstructed local subgraph. The set of local update evidence packages refers to the reorganized set of local evidence inputs for a single affected subgraph, including both the original source constraint evidence packages that the affected subgraph relied on and the newly added source constraint evidence packages that triggered the update. Through this process, the reconstruction of the local subgraph is not a regeneration detached from the original evidence foundation, but rather a local reconstruction based on the original local evidence with the introduction of new evidence. In this way, the original valid relational structure is preserved, and the new changes are absorbed into the reconstruction result.
[0082] After obtaining the reconstructed local subgraphs, the corresponding local subgraphs in the multi-layered knowledge graph with quality states are replaced with the reconstructed local subgraphs, and the version number corresponding to the reconstructed local subgraphs is incremented to obtain the incrementally updated knowledge graph. The version number incrementing means that, while keeping the versions of unaffected local subgraphs unchanged, only the reconstructed local subgraphs are assigned the updated version identifier. Thus, the incrementally updated knowledge graph preserves the stable structure of the unaffected parts of the existing knowledge graph while enabling the affected parts to undergo local replacement and version updates. For example, when a newly added source constraint evidence package only affects two local subgraphs, only these two local subgraphs are replaced with the reconstructed local subgraphs, while the remaining local subgraphs remain in their original state, thereby avoiding redundant calculations of the entire knowledge graph due to local changes.
[0083] After the above processing, the newly added set of source-constrained evidence packages is systematically incorporated into the multi-layered knowledge graph update process with quality status, ultimately resulting in an incrementally updated knowledge graph. This processing enables the knowledge graph to no longer rely on full reconstruction when faced with continuously arriving multi-source heterogeneous review data. Instead, it can perform local updates based on the real relationships between new evidence and existing local subgraphs, reducing redundant processing issues caused by excessively large update scopes, while maintaining the continuous correspondence between graph structure, quality status, and version information.
[0084] like Figure 3 This invention also provides a multi-source heterogeneous fair competition review knowledge graph construction system, including:
[0085] The evidence encapsulation module is used to acquire multi-source heterogeneous review data in the field of fair competition review, perform unified text processing and cleaning on the multi-source heterogeneous review data, construct a source-constrained evidence package set, and configure source credibility weights for the source-constrained evidence packages in the source-constrained evidence package set.
[0086] The semantic regularization module is used to perform semantic segmentation on the source constraint evidence package set, generate a set of structured semantic units, and generate a set of source constraint features based on the source credibility weight and the completeness and structural purity of each structured semantic unit.
[0087] The entity relationship module is used to generate a candidate entity cluster set and a directed relationship unit set based on the set of structured semantic units and the set of source constraint features;
[0088] The link construction module is used to construct review logic chains based on the candidate entity cluster set and the directed relation unit set, and to determine the set of review logic chains that can be merged.
[0089] The fact classification module is used to merge and classify the relational facts in the set of fusionable review logic chains, and generate a core fact set, a supporting fact set, and a conflict record set.
[0090] The graph quality control module is used to construct a multi-layered knowledge graph with quality status based on the core fact set, the circumstantial fact set, the conflict record set, and the source constraint evidence package set.
[0091] The incremental update module is used to obtain the newly added source constraint evidence package set, determine the affected subgraph based on the newly added source constraint evidence package set and the multi-layer knowledge graph with quality state, and re-execute steps 2 to 6 on the affected subgraph to obtain the incrementally updated knowledge graph.
[0092] It should be noted that the interval and threshold sizes are set for ease of comparison. The size of the threshold depends on the amount of sample data and the base number set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless calculations, and the formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0093] The embodiments of the present invention have been described above, but the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of the present embodiments, all of which are within the protection scope of the present embodiments.
Claims
1. A method for constructing a multi-source heterogeneous fair competition review knowledge graph, characterized in that, Includes the following steps: Step 1: Obtain multi-source heterogeneous review data in the field of fair competition review, perform unified text processing and cleaning on the multi-source heterogeneous review data, construct a source-constrained evidence package set, and configure source credibility weights for the source-constrained evidence packages in the source-constrained evidence package set; Step 2: Semantically segment the source constraint evidence package set to generate a set of structured semantic units, and generate a set of source constraint features based on the source credibility weight and the completeness and structural purity of each structured semantic unit; Step 3: Based on the set of structured semantic units and the set of source constraint features, generate a set of candidate entity clusters and a set of directed relation units; Step 4: Based on the candidate entity cluster set and the directed relation unit set, construct the review logic chain and determine the set of review logic chains that can be merged; Step 5: Merge and classify the relational facts in the set of fusionable review logic chains to generate a core fact set, a supporting fact set, and a conflict record set; Step 6: Based on the core fact set, the circumstantial fact set, the conflict record set, and the source constraint evidence package set, construct a multi-layer knowledge graph with quality status; Step 7: Obtain the newly added source constraint evidence package set. Based on the newly added source constraint evidence package set and the multi-layer knowledge graph with quality state, determine the affected subgraph. Re-execute steps 2 to 6 on the affected subgraph to obtain the incrementally updated knowledge graph.
2. The method for constructing a multi-source heterogeneous fair competition review knowledge graph according to claim 1, characterized in that, Acquire multi-source heterogeneous review data in the field of fair competition review, perform unified textual processing and cleaning on the multi-source heterogeneous review data, construct a source-constrained evidence package set, and configure source credibility weights for the source-constrained evidence packages in the source-constrained evidence package set, including: Step 11: Obtain multi-source heterogeneous review data, determine the source category tag and format tag corresponding to each original review data, and perform unified text processing on each original review data according to the format tag to generate unified text content and location mapping information; Step 12: Clean the uniform text content, and encapsulate the cleaned uniform text content along with the source category marker, record time, record number, format marker and location mapping information to obtain the source constraint evidence package set; Step 13: Determine the credibility of the source category based on the source category marker and the preset correspondence between source categories; determine the completeness of the metadata based on the number of valid items in the record time, record number, and format marker; determine the location citation degree based on the proportion of valid text fragments in the location mapping information that can be referenced back to the original review data to the total number of valid text fragments; and perform weighted processing on the source category credibility, metadata completeness, and location citation degree to configure the source credibility weight for the source constraint evidence package in the source constraint evidence package set.
3. The method for constructing a multi-source heterogeneous fair competition review knowledge graph according to claim 1, characterized in that, The source constraint evidence set is semantically segmented to generate a set of structured semantic units. Based on the source credibility weight and the completeness and structural purity of each structured semantic unit, a source constraint feature set is generated, including: Step 21: Semantically segment each source constraint evidence package in the source constraint evidence package set according to paragraph boundaries, title boundaries, table boundaries, and action phrase boundaries to obtain a semantic fragment set; Step 22: Extract subject information, action information, object information, result information, basis information, time information, and region information from each semantic segment in the semantic segment set, and encapsulate them in the order of subject information, action information, object information, result information, basis information, time information, and region information to generate a set of structured semantic units; Step 23: Count the number of valid information items in the subject information, action information, object information, result information, basis information, time information and region information of each structured semantic unit, and determine the completeness of each structured semantic unit by the ratio of the number of valid information items to the total number of the seven information items. Step 24: Content tags that can be uniquely classified into one of the following categories are identified as valid content tags: subject information, action information, object information, result information, basis information, time information, and region information. The structural purity is determined by the ratio of the number of valid content tags to the total number of content tags. The source credibility weight, completeness, and structural purity are combined according to their correspondence to generate a source constraint feature set that corresponds one-to-one with each structured semantic unit.
4. The method for constructing a multi-source heterogeneous fair competition review knowledge graph according to claim 1, characterized in that, Based on the set of structured semantic units and the set of source constraint features, a set of candidate entity clusters and a set of directed relation units are generated, including: Step 31: Perform weighted processing based on the source credibility weight, completeness and structural purity of each source constraint feature in the source constraint feature set to determine the confidence of each structured semantic unit, and extract the entity candidates corresponding to the subject information, object information, basis information and region information; Step 32: For entity candidates with the same slot category, determine the degree of name consistency, the degree of context consistency, the degree of slot category consistency, and the degree of coupling of structured semantic unit confidence, respectively. The degree of coupling of structured semantic unit confidence is determined based on the smaller value of the structured semantic unit confidence of the two entity candidates. Step 33: Perform a weighted summation of name consistency, context consistency, slot category consistency, and structured semantic unit confidence coupling to determine the merging score; group entity candidates with merging scores not lower than the preset merging threshold into the same candidate entity cluster, and place the remaining entity candidates into different candidate entity clusters to obtain a set of candidate entity clusters; Step 34: Determine the head entity cluster identifier and tail entity cluster identifier based on the candidate entity cluster to which the subject information and object information in each structured semantic unit belong, and generate a set of directed relation units by combining action information, result information, basis information, time information, and region information; determine the relation support of each directed relation unit as the structured semantic unit confidence corresponding to the structured semantic unit that generates each directed relation unit; perform a unique correspondence verification on the head entity cluster identifier and tail entity cluster identifier, and retain the directed relation units that pass the verification.
5. The method for constructing a multi-source heterogeneous fair competition review knowledge graph according to claim 4, characterized in that, Based on the candidate entity cluster set and the directed relation unit set, a review logic chain is constructed, and a set of fusionable review logic chains is determined, including: Step 41: Based on the candidate entity cluster set and the directed relation unit set, the tail entity cluster identifier of the previous directed relation unit is the same as the head entity cluster identifier of the next directed relation unit, and the directed relation units with consistent information, time information and regional information are assembled in order to obtain the review logic chain set. Step 42: Count the number of valid items in the starting entity cluster identifier, relation type sequence, ending entity cluster identifier, result information sequence, basis information, time information and region information in each review logic chain. Use the number of valid items as the numerator and the total number of the seven items as the denominator to obtain the logical completeness. Step 43: Count the number of directed relation units in each review logic chain that are consistent with the assembly direction. Use the number of directed relation units as the numerator and the total number of directed relation units as the denominator to obtain the direction consistency. Add up the relation support of each directed relation unit as the numerator and use the total number of directed relation units as the denominator to obtain the chain support. Step 44: Multiply the logical completeness by the preset weight corresponding to the logical completeness, multiply the directional consistency by the preset weight corresponding to the directional consistency, multiply the chain support by the preset weight corresponding to the chain support, and sum the three multiplications to obtain the fusion judgment value; determine the review logic chains whose fusion judgment value is not lower than the preset fusion threshold as the set of review logic chains that can be fused, and determine the remaining review logic chains as the set of review logic chains to be corrected.
6. The method for constructing a multi-source heterogeneous fair competition review knowledge graph according to claim 1, characterized in that, The relationship facts in the set of fusionable review logic chains are merged and conflict-classified to generate a core fact set, a supporting fact set, and a conflict record set, including: Step 51: Read the starting entity cluster identifier, relation type sequence, ending entity cluster identifier, result information sequence, basis information, time information, region information, fusion judgment value and chain support degree corresponding to each fusion review logic chain, and encapsulate them in a fixed order to obtain the relation facts; Step 52: Group relation facts with the same starting entity cluster identifier, relation type sequence, and ending entity cluster identifier into the same merging unit; compare the result information sequence, basis information, time information, and region information of relation facts within the same merging unit; those with all four items being the same are determined to be the same fact version. Step 53: For the relational facts in each fact version, multiply the fusion judgment value of each relational fact by the chain support degree respectively and then sum them up to obtain the version support value; then determine the degree of conflict between each fact version based on the proportion of the maximum version support value in the cumulative result of all version support values. Step 54: When the merging unit contains only one fact version, the corresponding relationship facts are included in the core fact set; when the conflict degree is not higher than the preset conflict threshold, the fact version corresponding relationship facts with the largest version support value are included in the core fact set, and the rest are included in the circumstantial fact set; when the conflict degree is higher than the preset conflict threshold, all fact version corresponding relationship facts, together with the version support value and conflict degree, are included in the conflict record set.
7. The method for constructing a multi-source heterogeneous fair competition review knowledge graph according to claim 1, characterized in that, Based on the core fact set, the circumstantial fact set, the conflict record set, and the source-constrained evidence package set, a multi-layered knowledge graph with quality status is constructed, including: Step 61: Based on the core fact set, the corroborating fact set, and the conflict record set, establish entity nodes, core fact edges, corroborating fact edges, conflict record nodes, and conflict association edges; then, based on the basis information, time information, and regional information corresponding to each fact object, determine the corresponding source constraint evidence package in the source constraint evidence package set, and establish source back-reference edges to obtain a multi-layer knowledge graph. Step 62: Using each core fact edge as the center, extract the corresponding supporting fact edges, conflict record nodes, conflict association edges, and source reference edges to construct a local subgraph. Completeness is determined by the proportion of fact objects that simultaneously possess a starting entity cluster identifier, relation type sequence, ending entity cluster identifier, result information sequence, basis information, time information, region information, and source reference edges to the total number of fact objects in the local subgraph. Consistency is determined by the proportion of fact objects that have not established a relationship with conflict record nodes to the total number of fact objects in the local subgraph. Traceability is determined by the proportion of fact objects with source reference edges to the total number of fact objects in the local subgraph. Timeliness is determined by the proportion of fact objects whose record time falls within a preset effective time window to the total number of fact objects in the local subgraph. The fact objects include core fact edges, supporting fact edges, and conflict association edges. Step 63: Multiply the completeness by the preset weight corresponding to the completeness, multiply the consistency by the preset weight corresponding to the consistency, multiply the traceability by the preset weight corresponding to the traceability, and multiply the timeliness by the preset weight corresponding to the timeliness. Then, sum the results to determine the quality status. Next, based on the comparison results of completeness, consistency, traceability, and timeliness with the corresponding preset thresholds, determine the repair identifier. Write the quality status and repair identifier into the corresponding local subgraph to obtain a multi-layer knowledge graph with quality status.
8. The method for constructing a multi-source heterogeneous fair competition review knowledge graph according to claim 1, characterized in that, Obtain the newly added source constraint evidence package set. Based on the newly added source constraint evidence package set and the multi-layer knowledge graph with quality states, determine the affected subgraphs. Re-execute steps 2 to 6 on the affected subgraphs to obtain the incrementally updated knowledge graph, including: Step 71: Obtain the newly added source constraint evidence package set, re-execute step 2 on the newly added source constraint evidence package set, generate the newly added structured semantic unit set and the newly added source constraint feature set, and extract subject information, object information, basis information and region information from the newly added structured semantic unit set to generate the newly added entity reference set; Step 72: For each local subgraph in the multi-layer knowledge graph with quality status, the source overlap is determined by the proportion of the number of newly added source constraint evidence packages that meet the source consistency condition, the entity overlap is determined by the proportion of the number of newly added entity references that meet the entity consistency condition, the time overlap is determined by the proportion of the number of newly added source constraint evidence packages whose recorded time falls within the time range, and the repair trigger value is determined based on whether the repair identifier is empty. Step 73: Based on the source overlap, entity overlap, time overlap, repair trigger value, and quality status, a comprehensive evaluation is performed according to their respective preset weights to obtain the affected judgment value; local subgraphs with affected judgment values not lower than the preset affected threshold are included in the affected subgraph set. Step 74: For each affected subgraph, combine its associated source constraint evidence package set with the newly added source constraint evidence package that triggered the affected subgraph to form a local update evidence package set. Then, re-execute steps 2 to 6 based only on the local update evidence package set to obtain the reconstructed local subgraph. Replace the corresponding local subgraph in the multi-layer knowledge graph with quality status with the reconstructed local subgraph and increment the version number corresponding to the reconstructed local subgraph to obtain the incremental update knowledge graph.
9. A multi-source heterogeneous fair competition review knowledge graph construction system, characterized in that, The method for constructing a multi-source heterogeneous fair competition review knowledge graph as described in any one of claims 1-8 includes: The evidence encapsulation module is used to acquire multi-source heterogeneous review data in the field of fair competition review, perform unified text processing and cleaning on the multi-source heterogeneous review data, construct a source-constrained evidence package set, and configure source credibility weights for the source-constrained evidence packages in the source-constrained evidence package set. The semantic regularization module is used to perform semantic segmentation on the source constraint evidence package set, generate a set of structured semantic units, and generate a set of source constraint features based on the source credibility weight and the completeness and structural purity of each structured semantic unit. The entity relationship module is used to generate a candidate entity cluster set and a directed relationship unit set based on the set of structured semantic units and the set of source constraint features; The link construction module is used to construct review logic chains based on the candidate entity cluster set and the directed relation unit set, and to determine the set of review logic chains that can be merged. The fact classification module is used to merge and classify the relational facts in the set of fusionable review logic chains, and generate a core fact set, a supporting fact set, and a conflict record set. The graph quality control module is used to construct a multi-layered knowledge graph with quality status based on the core fact set, the circumstantial fact set, the conflict record set, and the source constraint evidence package set. The incremental update module is used to obtain the newly added source constraint evidence package set, determine the affected subgraph based on the newly added source constraint evidence package set and the multi-layer knowledge graph with quality state, and re-execute steps 2 to 6 on the affected subgraph to obtain the incrementally updated knowledge graph.