Knowledge base collaborative updating method based on multi-source corpus
Patent Information
- Application Number
- CN202611200654.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-10
- Publication Date
- 2026-09-04
AI Technical Summary
[0003]针对现有技术中缺乏对候选三元组冲突风险的量化判别以及对错误关系谓语缺少多源自动修正手段的问题,提供一种基于多源语料的知识库协同更新方法,通过计算候选三元组尾实体类型与知识库中既有三元组尾实体类型分布之间的类型偏移量,实现冲突等级的自动评估与差异化处理,同时利用多源语料的共现频次对高冲突候选三元组的关系谓语进行修正,提升知识库更新的准确性和自动化程度
通过采集目标知识库中与候选更新三元组具有相同头实体的既有三元组,提取这些既有三元组的尾实体类型分布,再根据候选更新三元组的尾实体类型与该分布之间的类型偏移量确定冲突等级。当候选三元组的尾实体类型与头实体已有关联的主要尾实体类型分布高度吻合时,类型偏移量小,冲突等级低,该三元组可以直接进入待审核队列,实现低风险知识的快速流转;当尾实体类型明显偏离既有类型分布时,类型偏移量增大,冲突等级升高,表明该候选三元组可能与知识库已有知识结构存在潜在矛盾,此时不直接写入,而是触发后续的佐证验证流程。这种分级处理策略使得知识库更新能够根据冲突程度自适应地调控候选三元组的审核路径,避免了将所有候选三元组不加区分地送入审核队列或直接写入所导致的知识冲突积累和类型分布失真,也减少了人工对低风险三元组的重复判断工作。
Smart Images

Figure CN122693799A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of knowledge base update technology, specifically a collaborative update method for knowledge bases based on multi-source corpora. Background Technology
[0002] The continuous updating of multi-source heterogeneous corpora provides a rich data source for the dynamic expansion of knowledge bases. How to automatically extract and integrate new knowledge triples from the incremental text streams of multiple corpora is a major problem in knowledge base maintenance. Existing knowledge base update methods typically involve directly extracting relations from the corpus, deduplicating the obtained triples, and then writing them into the knowledge base in batches, or relying entirely on manual review queues for line-by-line verification. When dealing with multi-source corpora, these methods lack quantitative assessment tools for the compatibility between new triples and existing knowledge in the knowledge base. The tail entity type carried by candidate triples may deviate from the existing tail entity type system of the head entity in the knowledge base; direct writing can lead to disordered type distribution within the knowledge base, increasing the bias in queries and inferences. Furthermore, when the relational predicates extracted from the corpus are ambiguous or erroneous, existing methods lack an automatic correction mechanism based on cross-validation of multi-source corpora, often requiring significant manual intervention for relation verification, making it difficult to balance update efficiency and accuracy. Faced with massive amounts of incremental text from multiple sources, how to automatically identify the conflict risk of candidate triples before writing them into the knowledge base, and how to self-correct the relational representations of high-risk triples to reduce invalid writing and manual intervention are problems that multi-source corpus-driven collaborative update technology for knowledge bases needs to solve. Summary of the Invention
[0003] To address the shortcomings of existing technologies, such as the lack of quantitative assessment of conflict risk in candidate triples and the absence of multi-source automatic correction methods for erroneous relational predicates, this paper proposes a knowledge base collaborative update method based on multi-source corpora. By calculating the type offset between the tail entity type of candidate triples and the existing tail entity type distribution of triples in the knowledge base, the method achieves automatic assessment and differentiated processing of conflict levels. Simultaneously, it utilizes the co-occurrence frequency of multi-source corpora to correct relational predicates of high-conflict candidate triples, thereby improving the accuracy and automation of knowledge base updates.
[0004] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides a knowledge base collaborative update method based on multi-source corpora. This method achieves high-quality continuous update of the target knowledge base by collaboratively processing and resolving conflicts of incremental text streams from multiple heterogeneous corpora.
[0005] In one technical solution of the present invention, incremental text streams from multiple heterogeneous corpora are acquired, and the incremental text streams are segmented into text segments to be verified according to time windows. Preferably, an independent polling period is set for each heterogeneous corpus, and incremental text entries are pulled from the change logs of each heterogeneous corpus according to their respective polling periods. All incremental text entries pulled from all heterogeneous corpora within the same time window are concatenated into an original text string in the order of receipt. The original text string is segmented according to sentence boundaries, and short sentences with a length less than a preset length threshold are filtered out. The remaining sentences are used as the text segments to be verified. Through differentiated polling mechanisms and time-window-based text aggregation, data omissions or temporal disorder caused by different update frequencies of each corpus can be avoided, ensuring the integrity and timeliness of the text segments to be processed.
[0006] After obtaining the text segment to be verified, the term dependency paths of the text segment are extracted, and these term dependency paths are mapped to candidate update triples for the target knowledge base. Specifically, the text segment to be verified is subjected to part-of-speech tagging and dependency parsing to extract the subject, predicate, and object words from the subject-verb-object structure. The subject word is then matched with the entity index of the target knowledge base using fuzzy string matching to determine the head entity identifier. The predicate word is mapped to a relation type encoding, and the object word is mapped to a tail entity identifier, thereby forming candidate update triples. This triple extraction method based on deep syntactic analysis can accurately capture the semantic relationships between entities from unstructured text, providing structured candidate knowledge units for knowledge base updates.
[0007] To effectively assess the reliability of new knowledge, this invention collects existing triples in the target knowledge base that share the same head entity as the candidate update triples, and extracts the tail entity type distribution of these existing triples. Using the head entity identifier of the candidate update triple as the search key, the inverted index of the target knowledge base is queried to obtain all existing triples with that head entity identifier as the subject. The tail entity identifier of each existing triple is read, and the corresponding type tag is obtained from the type registry. The occurrence frequency of each type tag is counted to construct the tail entity type distribution. By analyzing the type distribution of existing knowledge, a benchmark for measuring the consistency between new knowledge and the existing knowledge system can be established.
[0008] In the conflict assessment stage, the conflict level of the candidate update triplet is determined based on the type offset between the tail entity type of the candidate update triplet and the distribution of the tail entity type. Preferably, the most frequently occurring primary type is extracted from the tail entity type distribution, and the semantic distance between the tail entity type of the candidate update triplet and this primary type is calculated. Different offset values are assigned when the semantic distance falls within different intervals. Then, the type offset is multiplied by the confidence coefficient of the candidate update triplet in the text segment to be verified, and the conflict level is finally obtained. This step combines the type semantic distance with the extraction confidence, enabling the conflict determination to comprehensively reflect the degree of deviation of the new knowledge and its own reliability.
[0009] When the conflict level is lower than the preset conflict threshold, the candidate update triplet is directly entered into the pending review queue of the target knowledge base. To ensure review efficiency, the current queue length of the target knowledge base can be read, and the preset conflict threshold can be dynamically adjusted based on the current queue length, making the preset conflict threshold inversely proportional to the current queue length. The conflict level is compared with the adjusted preset conflict threshold; if it is lower than the threshold, a pending review record containing the candidate update triplet and its timestamp is generated and appended to the tail of the pending review queue. This adaptive threshold adjustment mechanism can dynamically balance review efficiency and data entry quality when the system load changes.
[0010] When the conflict level is not lower than a preset conflict threshold, a set of supporting texts associated with the head entity of the candidate update triple is recalled from multiple heterogeneous corpora. The relational predicate of the candidate update triple is corrected based on the co-occurrence frequency of the supporting text set, and the corrected candidate update triple is sent to the review queue. For supporting text recall, a set of synonymous entities corresponding to the head entity identifier is extracted, and each synonym is used as an extended search term. Searches are performed in the title and summary fields of multiple heterogeneous corpora to obtain matching documents, and contextual paragraphs containing the extended search terms are extracted as the supporting text set. For relational predicate correction, the head entity identifier and tail entity identifier are used as joint search terms to search the full text of multiple heterogeneous corpora. The frequency of each relational predicate word in the recalled supporting statements is counted, and the relational predicate word with the highest frequency is selected as the candidate relational predicate. When the candidate relational predicate is inconsistent with the original relational predicate, it is replaced. By cross-verifying multiple corpora and correcting relationships based on co-occurrence statistics, we can effectively correct relationship errors caused by single-source bias or extraction errors, and improve the accuracy of knowledge updates under high-conflict conditions.
[0011] Finally, according to the timestamp order of the triples in the queue to be reviewed, each triple is written to the main storage area of the target knowledge base in sequence. During writing, the triple to be reviewed and its corresponding timestamp are read from the head of the queue, the current version number of the main storage area is obtained, and the operation log is queried using the corresponding timestamp and the current version number as a key to confirm whether the triple has been written. If it has not been written, the write operation is performed after appending the successor version number of the current version number, and the operation log is updated. By using idempotent write control based on timestamps and version numbers, duplicate updates can be prevented, ensuring data consistency of the knowledge base in the main storage area.
[0012] After the write operation is completed, the header entity identifiers of all triples involved in this write operation are read. The latest triples with that header entity identifier as the subject are then retrieved from the target knowledge base. A type-relation co-occurrence matrix for that header entity is constructed using the relation type encoding as the row index and the tail entity type as the column index. When the change in the row and column dimensions of the type-relation co-occurrence matrix exceeds a threshold, the index reconstruction process of the target knowledge base is triggered. This optimization process can perceive the evolution of the knowledge structure in real time and update the index structure promptly when the entity relationship graph undergoes significant reorganization, ensuring that subsequent query and reasoning services always operate based on the latest knowledge organization.
[0013] The technical effects and advantages provided by the present invention in the above technical solution are as follows: By collecting existing triples with the same head entity as the candidate update triples from the target knowledge base, the tail entity type distribution of these existing triples is extracted. The conflict level is then determined based on the type offset between the tail entity type of the candidate update triple and this distribution. When the tail entity type of a candidate triple highly matches the distribution of the main tail entity types already associated with the head entity, the type offset is small, the conflict level is low, and the triple can directly enter the review queue, enabling rapid flow of low-risk knowledge. When the tail entity type significantly deviates from the existing type distribution, the type offset increases, and the conflict level rises, indicating that the candidate triple may have a potential contradiction with the existing knowledge structure of the knowledge base. In this case, it is not written directly but triggers the subsequent verification process. This hierarchical processing strategy allows the knowledge base update to adaptively adjust the review path of candidate triples according to the degree of conflict, avoiding the accumulation of knowledge conflicts and distortion of type distribution caused by indiscriminately sending all candidate triples into the review queue or writing them directly. It also reduces the repetitive manual judgment work for low-risk triples.
[0014] When the conflict level is not lower than a preset conflict threshold, a set of supporting texts associated with the head entity of the candidate update triple is recalled from multiple heterogeneous corpora. Utilizing the context of co-occurrence of the head and tail entities in the supporting text set, the frequency of each relational predicate is statistically analyzed, and the relational predicate with the highest frequency is selected to correct the original candidate relational predicate. When relational predicates extracted from a single corpus or short text contain expression biases or errors, the co-occurrence frequency statistics of multi-source corpora can provide a more robust choice of relational representations. Cross-validation across corpora automatically identifies and replaces accidental or erroneous relational predicates, and the corrected triples are then sent to the review queue. This mechanism achieves self-correction of relational predicates by relying on the statistical characteristics of the multi-source corpora themselves without manual annotation, effectively reducing the proportion of erroneous entries caused by errors in relational predicate extraction and improving the construction quality of the knowledge base in multi-source corpus collaborative update scenarios. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0016] Figure 1 This is a flowchart of a knowledge base collaborative update method based on multi-source corpora; Figure 2 This is a flowchart of the term dependency path extraction and candidate update triple mapping process for the text segment to be verified. Figure 3 This is a flowchart of the conflict level determination method; Figure 4 This is a fuzzy matching score distribution diagram of candidate updated triples; Figure 5 This is a graph showing the relationship between the conflict level of candidate update triples and the dynamically adjusted preset conflict threshold. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] See Figure 1This invention provides a collaborative knowledge base update method based on multi-source corpora, comprising: acquiring incremental text streams from multiple heterogeneous corpora; segmenting the incremental text streams into text segments to be verified according to time windows; extracting term dependency paths from the text segments to be verified; mapping the term dependency paths to candidate update triples in the target knowledge base; collecting existing triples in the target knowledge base that have the same head entities as the candidate update triples; extracting the tail entity type distribution of the existing triples; and determining the candidate update triples based on the type offset between the tail entity type and the tail entity type distribution of the candidate update triples. The conflict level of the new triplet; when the conflict level is lower than the preset conflict threshold, the candidate updated triplet is directly entered into the pending review queue of the target knowledge base; when the conflict level is not lower than the preset conflict threshold, the set of supporting texts associated with the head entity of the candidate updated triplet is recalled from multiple heterogeneous corpora, the relational predicate of the candidate updated triplet is corrected based on the co-occurrence frequency of the supporting text set, and the corrected candidate updated triplet is sent into the pending review queue; according to the timestamp order of each triplet in the pending review queue, each triplet is written into the main storage area of the target knowledge base in sequence.
[0019] Example 1 In practice, the process of acquiring incremental text streams from multiple heterogeneous corpora and dividing the incremental text streams into text segments to be verified according to time windows is as follows.
[0020] An independent polling cycle is set for each heterogeneous corpus. It is understood that the data update frequency of each heterogeneous corpus differs, and using a uniform polling cycle may lead to a backlog of incremental text entries in high-frequency corpora or excessive empty requests in low-frequency corpora. In one optional implementation, the number of change log entries for each heterogeneous corpus within a predetermined statistical period is obtained, and the historical change frequency of each heterogeneous corpus is calculated. The historical change frequency of each heterogeneous corpus is denoted as Historical change frequency The unit is entries per second. The average historical change frequency is calculated based on the historical change frequencies of all heterogeneous corpora. Average historical change frequency Depend on Received, among which The total number of heterogeneous corpora. Set the baseline polling period. Baseline polling cycle The value is set to 60 seconds. The basis for setting the value to 60 seconds is: when the historical change frequency of a single heterogeneous corpus equals the average historical change frequency. At this time, retrieving incremental text entries at 60-second intervals can keep the retrieval delay of incremental text entries within a reasonable range without significantly increasing the query load of heterogeneous corpora. Subsequently, according to the formula... Determine the first The polling cycle of a heterogeneous corpus, among which Indicates the first The polling cycle of the heterogeneous corpus. As can be seen from the formula, the historical change frequency... Higher than the average historical change frequency The heterogeneous corpus is assigned a shorter polling cycle. Historical change frequency Lower than the average historical change frequency The heterogeneous corpus is assigned a longer polling cycle. .
[0021] Incremental text entries are retrieved from the change log according to the polling cycle of each heterogeneous corpus. For any heterogeneous corpus, when its polling cycle arrives, a request is sent to the change log query interface of that corpus, carrying the generation timestamp of the last incremental text entry recorded in the previous retrieval operation as the starting offset. The change log query interface returns all incremental text entries newly generated since the starting offset, each containing a text content field and a generation timestamp field. After the retrieval is complete, the generation timestamp of the last incremental text entry in the returned result is extracted and persisted as the starting offset for the next retrieval request.
[0022] All incremental text entries retrieved from all heterogeneous corpora within the same time window are concatenated into the original text string in the order of receipt. The start time of the time window is the end time of the previous time window, and the end time is the start time plus a preset time window length, which is set to 300 seconds. Within the time interval defined by the preset time window length, all incremental text entries asynchronously retrieved from various heterogeneous corpora are temporarily stored in a buffer. Each incremental text entry is appended with a receipt timestamp when it enters the buffer; the receipt timestamp is the precise moment the entry was received by the local system. When the time window ends, all incremental text entries in the buffer are arranged in ascending order of their receipt timestamps. The text content fields of each incremental text entry are then extracted sequentially and concatenated to form the original text string.
[0023] The original text string is segmented according to sentence boundaries, and short sentences are filtered out to obtain the text segments to be verified. The original text string is scanned using a set of sentence-ending symbols, including periods, question marks, exclamation marks, and newlines. When any sentence-ending symbol is encountered, the character sequence from the previous segmentation point to the current sentence-ending symbol is divided into an independent sentence. After segmentation, the length of each independent sentence is checked one by one. For Chinese text, sentence length is measured by the number of Chinese characters; for English text, sentence length is measured by the number of words. The length of each independent sentence is compared with a preset length threshold, which is set according to the primary language: 5 Chinese characters when the primary language of the text segment to be verified is Chinese; 3 words when the primary language is English. Independent sentences shorter than the preset length threshold are discarded, while those longer than the preset length threshold are retained. The retained independent sentences are the text segments to be verified, which serve as input for subsequent extraction of candidate update triples.
[0024] Example 2 In specific implementation, please refer to Figure 2 The process of extracting the term dependency paths of the text segment to be verified and mapping the term dependency paths to candidate update triples of the target knowledge base is as follows.
[0025] The text segment to be validated is subjected to part-of-speech tagging and dependency parsing. Part-of-speech tagging provides word class information for dependency parsing, while dependency parsing provides grammatical relation support for extracting subject-verb-object structures. In one optional implementation, the text segment to be validated is input into a sequence labeling model based on a bidirectional long short-term memory network combined with a conditional random field. The sequence labeling model outputs the part-of-speech tag for each word in the text segment. The tagged text segment is then input into a graph dependency analyzer based on a dual affine attention mechanism. The graph dependency analyzer outputs the dependency arc tags and dependency directions of each word with other words in the sentence, forming a dependency tree. Starting from the word marked as the root node in the dependency tree, the algorithm traverses towards the leaf nodes, locating words whose part-of-speech tag is a verb and whose dependency arc tag indicates a core predicate. These located words are recorded as predicate words. Subsequently, the dependency tree is searched for dependent nodes whose parent node is the predicate and whose dependency arc labels conform to the subject relation label set. The subject relation label set includes grammatical role labels such as "nsubj" and "csubj". The word corresponding to the dependent node is recorded as the subject. Similarly, the dependency tree is searched for dependent nodes whose parent node is the predicate and whose dependency arc labels conform to the object relation label set. The object relation label set includes grammatical role labels such as "dobj" and "iobj". The word corresponding to the dependent node is recorded as the object. If multiple subjects or objects exist in the dependency tree, the subject and object with the closest syntactic distance to the predicate are selected as the final subject and object.
[0026] The subject is matched against the entity index of the target knowledge base using fuzzy string matching to determine the head entity identifier. The entity index of the target knowledge base stores a mapping between the standard name and entity identifier of each entity; the entity identifier is a unique code that identifies an entity within the target knowledge base. The fuzzy string matching operation is calculated based on the similarity of character sets. Specifically, the subject is denoted as... The standard name of each entity to be matched in the entity index of the target knowledge base is denoted as... The similarity between the subject and the standard name of each entity to be matched is calculated using the following fuzzy matching score calculation formula:
[0027] in, Subject Standard name of the entity to be matched The fuzzy matching score, Subject The character set, Indicates the standard name of the entity to be matched. The character set, This indicates the number of elements in the character set of the subject word. This indicates the number of elements in the character set of the standard name of the entity to be matched. This represents the number of elements in the intersection of the subject word character set and the character set of the entity standard name to be matched. After traversing all entity standard names in the entity index of the target knowledge base, the entity standard name with the highest fuzzy matching score that exceeds the preset similarity threshold is selected as the matching entity. The entity identifier corresponding to the matching entity is read and used as the head entity identifier. The preset similarity threshold is set to 0.8. The basis for this value is that when the Dice coefficient of the character sets of two Chinese words is not lower than 0.8, the probability that they point to the same entity in most contexts exceeds 95%, which can balance matching accuracy and matching recall.
[0028] The predicate is mapped to a relation type code in the target knowledge base. The target knowledge base has a pre-defined relation type mapping table. Each record in the table contains the standard predicate text, the relation type code, and a list of allowed variants. During the mapping operation, the predicate is compared with the standard predicate text and the list of allowed variants for all records in the relation type mapping table using exact string matching. If the predicate is exactly the same as the standard predicate text or any variant word in the list of allowed variants for a record, the relation type code corresponding to that record is read and used as the relation type code obtained from the predicate mapping. If the predicate cannot be exactly matched with any record, the stem of the predicate is extracted, and the extracted stem is again used for exact matching with the standard predicate text and the list of allowed variants for all records in the relation type mapping table. If a match still cannot be found, the fuzzy matching score calculation method described above is used to calculate the fuzzy matching score between the predicate and each standard predicate text. The relation type code corresponding to the standard predicate text with the highest fuzzy matching score is selected as the relation type code obtained from the predicate mapping. The step of mapping object words to tail entity identifiers adopts the same string fuzzy matching process as mapping subject words to head entity identifiers. That is, using the entity index of the target knowledge base, the similarity between the object word and the standard name of each entity is calculated using the same fuzzy matching score calculation formula, and the entity identifier with the highest matching score and exceeding the preset similarity threshold of 0.8 is selected as the tail entity identifier.
[0029] After obtaining the head entity identifier, relation type code, and tail entity identifier, they are combined in the form of triples to form a candidate update triple. The format of the candidate update triple is (head entity identifier, relation type code, tail entity identifier).
[0030] See Figure 4In the figure, the horizontal axis represents the fuzzy string matching score between the subject of the candidate update triple and the standard name of the entity in the target knowledge base entity index, ranging from 0 to 1. The vertical axis represents the frequency of the corresponding fuzzy matching score. The shaded bars in the legend show the frequency distribution of the fuzzy matching scores, the solid black line represents the kernel density estimation curve calculated based on this frequency distribution, and the dashed line is the reference line with a preset similarity threshold of 0.8.
[0031] The figure shows a bimodal distribution of fuzzy matching scores. The first peak is between approximately 0.2 and 0.4, with a low frequency and relatively dispersed distribution. The second peak is significantly concentrated between 0.9 and 1.0, with a significantly higher frequency. Furthermore, the kernel density curve in this range shows a sharp rise followed by a rapid decline, indicating that most subject words have matching scores higher than 0.8 with the standard entity name. This bimodal distribution reflects that during the process of extracting subject words and mapping them to head entity identifiers, the matching results are concentrated in two different intervals: low similarity and high similarity.
[0032] A preset similarity threshold of 0.8 is indicated by a dashed line, located at a clear numerical interval between the two peak distributions. This threshold divides the fuzzy matching score into a high similarity interval and a low similarity interval. In the high similarity interval (score ≥ 0.8), a large number of matching results are concentrated, indicating that these subjects are highly similar to the target knowledge base entities, effectively ensuring the accuracy and recall of the matching. In the low similarity interval (score < 0.8), the matching frequency is low and the distribution is relatively scattered, indicating that some subjects have a low degree of matching with the standard entity names, possibly corresponding to incorrect or ambiguous matches, and are therefore filtered out in this implementation.
[0033] Example 3 In specific implementation, please refer to Figure 3 The process of collecting existing triples with the same head entity as the candidate update triples from the target knowledge base and extracting the tail entity type distribution of the existing triples is as follows.
[0034] Using the head entity identifier of the candidate update triplet as the search key, the inverted index of the target knowledge base is queried. The inverted index of the target knowledge base uses the head entity identifier as the primary key and stores the storage location information of all existing triplets with the corresponding head entity identifier as the subject. Each index entry in the inverted index contains a head entity identifier field and a triplet identifier list field. The triplet identifier list field records the unique identifiers of all related existing triplets in the main storage area. During the query, the head entity identifier of the candidate update triplet is precisely matched against the head entity identifier field in the inverted index. If a match is successful, the triplet identifier list is returned. Based on each triplet identifier in the triplet identifier list, complete existing triplet records are sequentially read from the main storage area to obtain all existing triplets with the head entity identifier of the candidate update triplet as the subject. If no index entry is matched, the set of all existing triplets is empty.
[0035] After obtaining all existing triples, read the tail entity identifier of each existing triple. For each tail entity identifier, access the type registry of the target knowledge base. The type registry is a system table in the target knowledge base that stores the mapping relationship between entity identifiers and type tags. Each record in the type registry contains an entity identifier field and a type tag field. The type tag field records the ontology type to which the entity belongs, and the type tag uses the type code corresponding to the leaf node in the hierarchical classification system. Search for matching records in the type registry using the tail entity identifier, and extract the type tags from the matching records. After traversing all existing triples, obtain a sequence of type tags that corresponds one-to-one with the tail entity identifiers of all existing triples.
[0036] Count the occurrences of each type label in all existing triples. Use the type label as the statistical key and accumulate the frequency of the type label sequence. Each time a type label is read, check if its current count exists. If not, initialize the count to 1; otherwise, increment the count by 1. After traversing all existing triples, treat each type label and its corresponding occurrence count as a key-value pair. All key-value pairs together constitute the tail entity type distribution. The tail entity type distribution can be represented in set form as follows: ,in Indicates type label, This indicates the number of times the corresponding type of tag appears. This indicates the number of distinct type labels in the tail entity type distribution. The tail entity type distribution is an empty set when the set of all existing triples is empty.
[0037] After obtaining the tail entity type distribution, the conflict level of the candidate update triple is determined based on the type offset between the tail entity type of the candidate update triple and the tail entity type distribution. It can be understood that the tail entity type distribution represents the existing fact type clustering of the same head entity in the target knowledge base, and the degree of deviation between the tail entity type of the candidate update triple and the existing clustering trend reflects the potential conflict level.
[0038] Extract the most frequent primary type from the tail entity type distribution. Iterate through all key-value pairs in the tail entity type distribution and compare their occurrence counts. The value of the most frequently occurring type tag is used as the primary type. If there are multiple most frequently occurring type tags, the type tag with the shallowest level in the type hierarchy is selected as the primary type; if the level depths are the same, the type tag with the smallest lexicographical order of its type code is selected as the primary type.
[0039] The semantic distance between the tail entity type and the primary type of the candidate update triple is calculated. The tail entity type of the candidate update triple is obtained by querying the type registry using the tail entity identifier of the candidate update triple, in the same way as the method described above for extracting the tail entity type label of the existing triple. The semantic distance is calculated based on the type hierarchy tree, which is constructed from the parent-child relationships of all type labels in the type registry of the target knowledge base. Each node in the type hierarchy tree corresponds to a type label, and the edges represent the direct inclusion relationship from the parent type to the child type. Starting from the node corresponding to the tail entity type of the candidate update triple, the algorithm moves along the edges to the upper or lower level of the type hierarchy tree until it reaches the node corresponding to the primary type. The number of edges traversed is recorded as the semantic distance. If there is no path between the two types in the type hierarchy tree, the semantic distance is set to a maximum fixed value. The maximum fixed value is the maximum possible value of the shortest path length between any two nodes in the type hierarchy tree plus 1. In this embodiment, the maximum fixed value is 100. When the tail entity type distribution is an empty set, there is no primary type, and the semantic distance is directly assigned the maximum fixed value of 100.
[0040] The type offset is determined based on semantic distance. A first distance threshold and a second distance threshold are set, with the first threshold set to 2 and the second to 5. The first distance threshold of 2 is chosen because, in the type hierarchy of a general knowledge graph, when the semantic distance is 2, the two types are related as grandparents / grandchildren or uncles / nephews, indicating a perceptible difference in conceptual connotation, but still maintaining a strong semantic connection. The second distance threshold of 5 is chosen because, when the semantic distance reaches 5, the nearest common ancestor of the two types in the type hierarchy tree has risen to a higher level of abstraction, and semantic separation begins; direct replacement could easily lead to factual conflicts. When the semantic distance is less than the first distance threshold (0 or 1), the type offset is assigned the first offset value of 0.2; when the semantic distance is not less than the first distance threshold and less than the second distance threshold (2, 3, or 4), the type offset is assigned the second offset value of 0.5; when the semantic distance is not less than the second distance threshold (greater than or equal to 5), the type offset is assigned the third offset value of 1.0. The first offset value of 0.2, the second offset value of 0.5, and the third offset value of 1.0 are set based on the following: standardizing the range of the conflict level to an interval that is multipliable by the confidence coefficient, while maintaining the linear risk increment relationship between the offset levels, so that the type offset directly reflects the basic probability of factual conflict.
[0041] In one optional implementation, the confidence coefficient of the candidate update triple in the text segment to be verified is obtained through the confidence output during the dependency parsing process. During the aforementioned dependency parsing process that extracts the subject, predicate, and object, the graph dependency parser synchronously outputs the confidence score of each dependency arc, with the confidence score ranging from (0,1). The subject dependency arc score pointing from the subject to the predicate is extracted and denoted as confidence component one; the object dependency arc score pointing from the object to the predicate is extracted and denoted as confidence component two; the core arc score, where the predicate is the root node of the syntax tree, is extracted and denoted as confidence component three. The confidence coefficient of the candidate update triple in the text segment to be verified is the product of confidence component one, confidence component two, and confidence component three, expressed by the following formula: ,in This represents the confidence coefficient of the candidate updated triple in the text segment to be verified. Indicates a subject-dependent arc fraction. Indicates an object-dependent arc fraction. This represents the core arc fraction. When the same predicate in the text segment to be validated is associated with multiple sets of subject and object words, the confidence coefficient is taken as the maximum value among the products of each set.
[0042] After obtaining the type offset and confidence coefficient, the conflict level of the candidate update triple is calculated. The formula for calculating the conflict level is:
[0043] in, Indicates the conflict level of candidate update triples. Indicates the type offset. This represents the confidence coefficient of the candidate updated triple in the text segment to be validated. As shown in the formula, the conflict level... The value range is (0,1], when the type offset Take the third offset value of 1.0 and the confidence coefficient When it approaches 1, the conflict level A value close to 1 indicates an extremely high probability of conflict; when the type offset... Take the first offset value of 0.2 and the confidence coefficient At a lower level, the conflict level A value close to 0 indicates an extremely low probability of conflict.
[0044] Example 4 In practice, the current queue length of the target knowledge base is read. The pending review queue of the target knowledge base is a first-in, first-out (FIFO) data structure stored in memory. Each element in the queue is a pending review record, containing a candidate update triplet and its corresponding timestamp. The current queue length is obtained by calling the length query interface provided by the pending review queue, and the obtained current queue length is denoted as... Current queue length The unit is the number of records to be reviewed.
[0045] The preset conflict threshold is dynamically adjusted based on the current queue length, making it inversely proportional to the current queue length. The formula for dynamically adjusting the preset conflict threshold is:
[0046] in, This indicates the adjusted preset conflict threshold. Indicates the threshold scaling factor. This represents the queue length smoothing constant. Indicates the current queue length. Threshold scaling factor. The value is set to 50, based on the following: the current queue length when the queue to be reviewed is empty. The value is 0, at which point the adjusted preset conflict threshold is... Pick To avoid sending all candidate update triples into the evidence recall process due to excessively high preset conflict thresholds in an empty queue state, the threshold scaling factor is adjusted. Smoothing constant with queue length The ratio was set below the upper limit of the common conflict level range. Analysis of the conflict level distribution showed that the conflict level range was (0,1], with common conflict levels concentrated between 0.3 and 0.7. A value of 0.9 allows the threshold to be slightly lower than the theoretical maximum of 1.0 when the queue is empty, but it can still intercept high-collision triples. (Queue length smoothing constant) The value is set to 55 so that when the queue is empty... This is close to the target value of 0.9, and the threshold decrease rate is appropriate as the queue length increases. According to the formula, as the current queue length... Increase and adjust the preset conflict threshold The threshold for high-conflict candidate update triples to be directly entered into the pending review queue is correspondingly reduced and increased.
[0047] The conflict level of the candidate updated triplet is compared with the adjusted preset conflict threshold. The conflict level of the candidate updated triplet is obtained by the conflict level calculation process described in Example 3, and is denoted as... Comparison operation judgment Is it less than the adjusted preset conflict threshold? .like If the conflict level is determined to be lower than the adjusted preset conflict threshold, a record to be reviewed, containing candidate update triples and their timestamps, is generated. The timestamp of the candidate update triple refers to the generation timestamp carried by the incremental text entry from which the candidate update triple originates. When a candidate update triple is derived from multiple incremental text entries, the earliest timestamp among the generation timestamps of each entry is used. The format of the record to be reviewed is a quadruple (header entity identifier, relation type code, tail entity identifier, timestamp). The generated quadruple is converted into a binary serialization format and used as the record to be reviewed. The serialized record to be reviewed is appended to the tail of the queue to be reviewed. The append operation is completed under the protection of a mutex lock to ensure the consistency of the queue state in a concurrent environment.
[0048] When the conflict level is not lower than the preset conflict threshold, that is In this case, the original candidate updated triplet will no longer be directly written into the pending review queue, but instead the supporting text recall and relational predicate correction process will be executed.
[0049] Extract the set of synonym entities corresponding to the head entity identifier of the candidate update triple. The target knowledge base has a pre-defined synonym entity mapping table, which uses the entity identifier as the primary key and records a list of synonym entity identifiers for each entity identifier. Each synonym entity identifier points to an entity in the target knowledge base that represents the same real-world object as the primary key entity identifier but has a different surface name. When querying the synonym entity mapping table, the head entity identifier of the candidate update triple is used as the query key. If a record is found, all entries in the corresponding synonym entity identifier list are read, and the head entity identifier of the candidate update triple is also added to the set, forming the synonym entity set. If no record is found, the synonym entity set only contains the head entity identifier of the candidate update triple.
[0050] Each synonym in the synonym entity set is used as an extended search term. Extended search terms are retrieved from the title and abstract fields of multiple heterogeneous corpora. For each heterogeneous corpus, a search query is constructed, with the target fields limited to the title and abstract fields. The search criteria are any extended search term that exactly matches the content of the target field or contains the stem of the extended search term. Each heterogeneous corpus returns a list of document identifiers for matching documents that meet the search criteria. A document identifier is a unique identifier assigned to each document within the heterogeneous corpus. All document identifier lists from all heterogeneous corpora are merged and deduplicated to obtain a single list of document identifiers without duplicates.
[0051] Extract context paragraphs containing the extended search terms from the documents corresponding to the document identifier list. For each document identifier in the document identifier list, retrieve the full text of the document from the corresponding heterogeneous corpus based on the document identifier. Locate the position where the extended search term appears in the full text of the document. Using the position where the extended search term appears as the center, extract a first preset number of sentences forward and a second preset number of sentences backward, which together with the sentence containing the extended search term to form a context paragraph. The first preset number is set to 2, and the second preset number is set to 2, meaning that the context paragraph includes the sentence containing the extended search term, plus the two sentences before and after it, for a maximum of 5 sentences in total. If the extended search term appears multiple times in the document, a context paragraph is extracted independently for each occurrence position. Combine all context paragraphs from all documents into a supporting text set.
[0052] Using the head and tail entity identifiers of candidate update triples as joint search terms, a search is conducted in the full text of multiple heterogeneous corpora to obtain several supporting statements containing the joint search terms. A joint search query is constructed, requiring that the standard entity name corresponding to both the head and tail entity identifiers appear simultaneously in the full text of the document. These two standard entity names are obtained by querying the entity index of the target knowledge base. The search scope is set to the full-text index of the heterogeneous corpora. Each heterogeneous corpus returns a list of matching documents. From these matching documents, all complete sentences containing both standard entity names are extracted, with each complete sentence serving as a supporting statement. Supporting statements must satisfy the condition that the character spacing between the head and tail entity standard entity names does not exceed a preset character limit (set to 50 characters) to avoid the two entity names being too far apart in the text and lacking actual semantic connection. The supporting statements returned by all heterogeneous corpora are then merged into a supporting statement set.
[0053] The frequency of each relational predicate in several supporting statements is counted. For each supporting statement in the set, dependency parsing is performed, using the same model as the one used to analyze the text segment to be verified in Example 2. The subject, predicate, and object of each supporting statement are extracted. It is determined whether the extracted subject matches the standard entity name corresponding to the head entity identifier of the candidate update triple, and whether the object matches the standard entity name corresponding to the tail entity identifier of the candidate update triple. If both match, the predicate of the supporting statement is included in the count as a relational predicate. Frequency statistics use a hash mapping structure, with the relational predicate text string as the key, and the number of occurrences is accumulated. If the set of supporting statements is empty, the frequency statistics result for the relational predicate is empty.
[0054] The relational predicate with the highest frequency is selected as the candidate relational predicate. When multiple relational predicates have the same highest frequency, the relational predicate that appears earliest in the supporting statement set is selected as the candidate relational predicate. If the frequency statistics result is empty, the original relational predicate of the candidate updated triple is used as the candidate relational predicate.
[0055] The candidate relation predicate is compared with the relation predicate of the candidate update triple. The relation predicate of the candidate update triple refers to the standard predicate text corresponding to the relation type code in the candidate update triple. The standard predicate text is obtained by looking up the relation type mapping table in the target knowledge base. If the text string of the candidate relation predicate is completely identical to the text string of the relation predicate of the candidate update triple, then the candidate relation predicate is considered to be identical to the relation predicate of the candidate update triple. In this case, no replacement operation is performed, and the original candidate update triple and its timestamp are generated as a record to be reviewed and appended to the tail of the queue to be reviewed. If the text string of the candidate relation predicate is inconsistent with the text string of the relation predicate of the candidate update triple, the relation predicate of the candidate update triple is replaced with the relation type code corresponding to the candidate relation predicate. The replacement method is as follows: the candidate relation predicate is mapped to the relation type code again through the relation type mapping table. If the candidate relation predicate cannot be mapped to any existing relation type code, a new record is added to the relation type mapping table, and a new relation type code is assigned. The newly assigned relation type code adopts the form of an auto-incrementing integer. After mapping or assigning the relation type code, a corrected candidate update triplet is generated. The corrected candidate update triplet has the same head entity identifier and tail entity identifier as the original candidate update triplet, but the relation type code is updated to the relation type code obtained from the mapping or assignment. The original timestamp is appended to the corrected candidate update triplet, generating a record to be reviewed. The record to be reviewed is appended to the tail of the queue to be reviewed, completing the operation of sending the corrected candidate update triplet into the queue to be reviewed.
[0056] See Figure 5 In the graph, the horizontal axis represents the current queue length, i.e., the number of records awaiting review in the target knowledge base's queue, expressed in records. The vertical axis represents the conflict level of the candidate update triplet, ranging from 0 to 1. In the legend, gray solid dots indicate conflict levels below the adjusted preset conflict threshold. Candidate update triples are indicated by their inclusion in the pending review queue; gray hollow diamonds represent conflict levels not lower than the adjusted preset conflict threshold. The candidate updated triples indicate that such triples will enter the corroboration recall process. The black curve represents the adjusted preset conflict threshold. With the current queue length The changing function curve conforms to the formula given in Example 4. ,in , .
[0057] As shown in the graph, as the length of the pending queue gradually increases from 0 to over 600, the adjusted preset conflict threshold... The threshold value exhibits a monotonically decreasing trend, starting at approximately 0.9 and decreasing to approximately 0.07 when the queue length reaches 600 queues. This trend reflects the system's strategy of dynamically adjusting the conflict threshold; that is, the longer the queue, the lower the threshold, increasing the proportion of high-conflict triples that directly enter the corroboration recall process and alleviating queue backlog.
[0058] On the vertical axis, the gray dots are mainly concentrated in the lower conflict level range (below the threshold curve), indicating that most low-conflict triples are directly entered into the pending review queue. The gray hollow rhombuses are relatively evenly distributed and all are located above the threshold curve, which conforms to the rule that when the conflict level is not lower than the threshold, they enter the evidence recall process. As the queue length increases, the overall triple conflict level distribution does not show a significant trend towards high or low values, indicating that the conflict level distribution of candidate update triples is well distinguished under dynamic adjustment.
[0059] Example 5 In practice, the process of writing each triplet into the main storage area of the target knowledge base in the order of its timestamp in the queue to be reviewed is as follows.
[0060] Read a triplet and its corresponding timestamp from the head of the pending review queue. The pending review queue is implemented as a thread-safe queue with a blocking mechanism. The head read operation removes the first pending record enqueued from the head of the queue and returns it, while retrieving the triplet data and timestamp value carried in that record. The triplet data contains three fields: head entity identifier, relation type code, and tail entity identifier. The timestamp value is a timestamp appended when the pending record was generated, and its format is consistent with the timestamp format uniformly used within the target knowledge base.
[0061] Retrieve the current version number of the target knowledge base's main storage area. The target knowledge base's main storage area maintains a monotonically increasing sequence of version numbers. Each time data changes occur in the main storage area, the version number increments according to a predetermined rule. The current version number is the version number assigned by the most recent data change operation, obtained by calling the version number query interface provided by the main storage area.
[0062] Using the retrieved timestamp and the current version number as the key, query the operation log of the target knowledge base. The operation log records historical information about each write operation performed on the main storage area. Each record in the operation log includes the head entity identifier, relation type code, tail entity identifier, timestamp of the write operation, the main storage area version number at the time of the write operation, and the operation type identifier for the write triple. When querying the operation log, an exact match is performed between the timestamp field and the timestamp value of the write operation, and simultaneously, an exact match is performed between the main storage area version number at the time of the write operation and the current version number. If there is a record in the operation log where the timestamp field of the write operation equals the timestamp value and the main storage area version number at the time of the write operation equals the current version number, then it is confirmed that the triple to be reviewed has been written under the current version number. If there is no record in the operation log that meets the above key matching conditions, then it is confirmed that the triple to be reviewed has not been written.
[0063] When it is confirmed that the triplet to be reviewed has been written under the current version number, skip the current write operation, directly mark the triplet to be reviewed as processed, and return to the step of reading the next triplet to be reviewed from the queue.
[0064] When it is confirmed that the triplet to be reviewed has not been written, the successor version number of the current version number is appended to the triplet to be reviewed, and then the write operation to the main storage area is performed. The successor version number is generated by incrementing the current version number by a predetermined step size, which is set to 1. That is, the successor version number is obtained by adding 1 to the current version number. The write operation in the main storage area is an atomic operation. The main storage area organizes data in the form of a triplet storage table. Each record in the triplet storage table contains a primary key identifier field, a header entity identifier field, a relation type code field, a tail entity identifier field, a creation timestamp field, and a version number field. During the write operation, a unique primary key identifier is assigned to the triplet to be reviewed, the header entity identifier, relation type code, and tail entity identifier are filled into the corresponding fields, the timestamp value is filled into the creation timestamp field, the successor version number is filled into the version number field, and the entire record is inserted into the triplet storage table. After the write operation is completed, a new record is appended to the operation log. The timestamp field of the write operation in the new record is set to the timestamp value, the main storage version number field when the write operation occurred is set to the successor version number, and the operation type identifier field is set to the identifier code corresponding to the insert operation.
[0065] After writing all the triples to be reviewed in the queue into the main storage area of the target knowledge base in sequence, the index reconstruction judgment and triggering process is executed.
[0066] Read the header entity identifiers of all triples involved in this write operation. Iterate through all triple records successfully written to the main storage area in this write operation, extracting the value of the header entity identifier field for each triple record. Deduplicate all extracted header entity identifiers to obtain a deduplicated header entity identifier set. If the deduplicated header entity identifier set is empty, this write operation will not trigger the subsequent index rebuild process.
[0067] For each header entity identifier in the deduplicated header entity identifier set, query the target knowledge base for all the latest triples with that header entity identifier as the subject. The query operation uses the header entity identifier as the query condition, traversing the triple storage table in the main storage area of the target knowledge base to obtain all target triple records that match the header entity identifier field. For multiple triple records with the same relation type code and the same tail entity identifier but different version numbers, only the record with the largest version number field value is retained, discarding the remaining older version records, thus obtaining the set of all the latest triples with that header entity identifier as the subject.
[0068] Using the relation type codes in all the latest triples as row indices and the tail entity types as column indices, a type-relation co-occurrence matrix is constructed for the head entity. The tail entity type is obtained as follows: for each tail entity identifier in the latest triples, the corresponding tail entity type label is obtained by querying the type registry of the target knowledge base. The query method is consistent with the method of obtaining type labels from the type registry in Example 3. When constructing the type-relation co-occurrence matrix, the frequency of each relation type code and each tail entity type label co-occurring in all the latest triples is counted. The rows of the matrix correspond to unique relation type codes, and the columns correspond to unique tail entity type labels. Line number The matrix element values of the column represent the first column. The first relation type encoding and the first The co-occurrence frequency of tail entity type tags. The dimension of the type-relation co-occurrence matrix is jointly determined by the number of duplicate relation type codes and the number of duplicate tail entity type tags. It can be understood that the row and column dimensions of the type-relation co-occurrence matrix can characterize the semantic richness and structural complexity associated with the head entity.
[0069] Calculate the change in row and column dimensions of the co-occurrence matrix of the relation type. The formula for calculating the change in row and column dimensions is:
[0070] in, This represents the change in row and column dimensions of the type-relationship co-occurrence matrix. This indicates the number of row dimensions of the type-relation co-occurrence matrix reconstructed after this write operation. This indicates the number of row dimensions in the type-relation co-occurrence matrix of the same head entity stored in the target knowledge base before this write operation. This indicates the number of column dimensions in the type-relationship co-occurrence matrix reconstructed after this write operation. This indicates the number of column dimensions in the type-relationship co-occurrence matrix of the same head entity stored in the target knowledge base before this write operation. The type-relationship co-occurrence matrix stored in the target knowledge base before this write operation refers to the head entity dimension snapshot recorded during the last index rebuild or after the last write. Dimension snapshots are stored in the entity metadata table of the target knowledge base. When a head entity does not have a historical dimension snapshot record in the entity metadata table of the target knowledge base, [the following will occur]. and All are considered as 0.
[0071] Changes in the row and column dimensions of the type-relationship co-occurrence matrix Compare with the change threshold. The change threshold is set to 5, based on the following: an increase in the type-relationship co-occurrence matrix dimension indicates a structural expansion of the semantic relationship network of the head entity. When the sum of the newly added row and column dimensions reaches 5, the scope covered by the index entries for that head entity in the original inverted index is insufficient to efficiently respond to queries involving new relationships and new tail entity types. Triggering index rebuilding at this point can avoid performance degradation due to incomplete index coverage. If the change is less than the threshold, the index rebuild process will not be triggered. If the change is greater than or equal to the threshold, the index reconstruction process of the target knowledge base is triggered. The specific operations of the index reconstruction process are as follows: all index entries in the inverted index of the target knowledge base involving the header entity identifier are marked as invalid; all the latest triples with the header entity identifier as the subject are read from the main storage area of the target knowledge base; the inverted index entries of these latest triples are reconstructed; and the newly constructed inverted index entries replace the old index entries marked as invalid, completing the partial reconstruction of the header entity's index. The index reconstruction process is executed asynchronously in the background and does not affect the normal read and write services of the main storage area of the target knowledge base. After the index reconstruction is completed, the dimension snapshot record corresponding to the header entity identifier in the entity metadata table of the target knowledge base is updated, and the row dimension number is updated to [value missing]. Update the column dimension number to .
[0072] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A knowledge base collaborative update method based on multi-source corpora, characterized in that, include: Obtain incremental text streams from multiple heterogeneous corpora, and segment the incremental text streams into text segments to be verified according to time windows; Extract the term dependency path of the text segment to be verified, and map the term dependency path to the candidate update triple of the target knowledge base; Collect existing triples in the target knowledge base that have the same head entity as the candidate update triples, and extract the tail entity type distribution of the existing triples. The conflict level of the candidate update triple is determined based on the type offset between the tail entity type of the candidate update triple and the distribution of the tail entity type. When the conflict level is lower than the preset conflict threshold, the candidate update triplet is directly entered into the pending review queue of the target knowledge base; When the conflict level is not lower than the preset conflict threshold, the set of supporting texts associated with the head entity of the candidate update triplet is recalled from the multiple heterogeneous corpora. The relational predicate of the candidate update triplet is corrected based on the co-occurrence frequency of the supporting text set, and the corrected candidate update triplet is sent to the queue to be reviewed. According to the timestamp order of each triple in the queue to be reviewed, each triple is written into the main storage area of the target knowledge base in sequence.
2. The knowledge base collaborative update method based on multi-source corpus according to claim 1, characterized in that, The process of acquiring incremental text streams from multiple heterogeneous corpora and segmenting these incremental text streams into text segments to be verified according to time windows includes: Set an independent polling cycle for each heterogeneous corpus, and pull incremental text entries from the change log of each heterogeneous corpus according to their respective polling cycles; All incremental text entries retrieved from all heterogeneous corpora within the same time window are concatenated into the original text string in the order of receipt. The original text string is segmented according to sentence boundaries, short sentences with a length less than a preset length threshold are filtered out, and the remaining sentences are used as the text segments to be verified.
3. The knowledge base collaborative update method based on multi-source corpus according to claim 1, characterized in that, The step of extracting the term dependency paths of the text segment to be verified and mapping the term dependency paths to candidate update triples of the target knowledge base includes: The text segment to be verified is subjected to part-of-speech tagging and dependency parsing to extract the subject, predicate, and object words in the subject-verb-object structure; The subject term is matched with the entity index of the target knowledge base to determine the head entity identifier. The predicate is mapped to the relation type encoding of the target knowledge base, and the object is mapped to the tail entity identifier; The candidate update triple is formed by the head entity identifier, the relation type code, and the tail entity identifier.
4. The knowledge base collaborative update method based on multi-source corpus according to claim 1, characterized in that, The step of collecting existing triples with the same head entity as the candidate update triples from the target knowledge base and extracting the tail entity type distribution of the existing triples includes: Using the head entity identifier of the candidate updated triple as the retrieval key, query the inverted index of the target knowledge base to obtain all existing triples with the head entity identifier as the subject; Read the tail entity identifier of each existing triple in all existing triples, and obtain the type tag corresponding to each tail entity identifier from the type registry of the target knowledge base; The occurrence count of each type of label in all existing triples is counted, and the tail entity type distribution is constructed by combining each type of label and its occurrence count.
5. The knowledge base collaborative update method based on multi-source corpus according to claim 1, characterized in that, The step of determining the conflict level of the candidate update triple based on the type offset between the tail entity type of the candidate update triple and the tail entity type distribution includes: Extract the most frequently occurring primary type from the tail entity type distribution, and calculate the semantic distance between the tail entity type of the candidate update triple and the primary type; When the semantic distance is less than the first distance threshold, the type offset is assigned a first offset value; when the semantic distance is not less than the first distance threshold and less than the second distance threshold, the type offset is assigned a second offset value; when the semantic distance is not less than the second distance threshold, the type offset is assigned a third offset value. The type offset is multiplied by the confidence coefficient of the candidate update triple in the text segment to be verified to obtain the conflict level.
6. The knowledge base collaborative update method based on multi-source corpus according to claim 1, characterized in that, When the conflict level is lower than a preset conflict threshold, the candidate update triplet is directly entered into the pending review queue of the target knowledge base, including: Read the current queue length of the target knowledge base, and dynamically adjust the preset conflict threshold according to the current queue length, so that the preset conflict threshold is inversely proportional to the current queue length; The conflict level is compared with the adjusted preset conflict threshold. If the conflict level is lower than the adjusted preset conflict threshold, a record to be reviewed containing the candidate update triplet and its timestamp is generated. The record to be reviewed is appended to the tail of the queue to be reviewed.
7. The knowledge base collaborative update method based on multi-source corpus according to claim 1, characterized in that, When the conflict level is not lower than the preset conflict threshold, the set of supporting texts associated with the head entity of the candidate update triple is recalled from the plurality of heterogeneous corpora, including: Extract the set of synonymous entities corresponding to the head entity identifier of the candidate update triplet, and use each synonymous entity in the set of synonymous entities as an expanded search term; The extended search terms are retrieved from the title and summary fields of the multiple heterogeneous corpora to obtain a list of document identifiers for matching documents; Extract the context paragraphs containing the extended search terms from the documents corresponding to the document identifier list, and combine all the context paragraphs into the supporting text set.
8. The knowledge base collaborative update method based on multi-source corpus according to claim 1, characterized in that, The step of modifying the relational predicate of the candidate update triple based on the co-occurrence frequency of the supporting text set includes: Using the head entity identifier and tail entity identifier of the candidate update triple as joint search terms, a search is conducted in the full text of the multiple heterogeneous corpora to obtain several supporting statements containing the joint search terms. The frequency of each relational predicate word in the aforementioned supporting statements is counted, and the relational predicate word with the highest frequency of occurrence is selected as the candidate relational predicate; When the candidate relation predicate is inconsistent with the relation predicate of the candidate updated triple, the relation predicate of the candidate updated triple is replaced with the candidate relation predicate.
9. The knowledge base collaborative update method based on multi-source corpus according to claim 1, characterized in that, The step of writing each triplet in the pending review queue into the main storage area of the target knowledge base in sequence according to its timestamp order includes: Read a triplet to be reviewed and its corresponding timestamp from the head of the queue to be reviewed, and obtain the current version number of the main storage area of the target knowledge base; Using the corresponding timestamp and the current version number as the key, query the operation log of the target knowledge base to confirm whether the triplet to be reviewed has been written under the current version number; If it is not written, then append the successor version number of the current version number to the triplet to be reviewed, perform the write operation to the main storage area, and update the operation log.
10. The knowledge base collaborative update method based on multi-source corpus according to claim 1, characterized in that, After writing each triplet in the pending review queue into the main storage area of the target knowledge base in the order of its timestamp, the method further includes: Read the header entity identifiers of all triples involved in this write operation, and query the target knowledge base for all the latest triples with the header entity identifier as the subject; Construct a type-relation co-occurrence matrix for the head entity using the relation type encoding in all the latest triples as the row index and the tail entity type as the column index; When the change in the row and column dimensions of the type-relationship co-occurrence matrix exceeds the change threshold, the index reconstruction process of the target knowledge base is triggered.