Multi-index knowledge base and its updating method and query self-recovery method

CN122432172BActive Publication Date: 2026-09-18ZHEJIANG CHUANGLIN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610908263.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-18
Estimated Expiration
2046-06-23

AI Technical Summary

Technical Problem

[0004]本发明针对现有技术中多索引知识库系统在数据同步构建与查询过程中存在的索引一致性难以保证、重建粒度过粗以及中断恢复效率低的缺点,提供了一种多索引知识库及其更新与查询自愈方法

Benefits of technology

所提出的多索引数据库以内容分块为统一单元构建或更新各类索引对象并通过分块结构快照统一版本基准、索引依赖映射明确关联关系,有效提高索引一致性,还能令后续数据维护过程中可针对内容分块确定待重建的索引对象以实现细粒度重建。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432172B_ABST
    Figure CN122432172B_ABST
Patent Text Reader

Abstract

The application discloses a multi-index knowledge base, an updating method thereof and a query self-recovery method. The multi-index knowledge base comprises a plurality of content blocks, each content block having a corresponding block structure snapshot and an index dependency mapping relationship. The block structure snapshot is used for indicating a document to which the content block belongs and a chapter path, a block order and a content fingerprint corresponding to the content block. The index dependency mapping relationship is used for indicating a plurality of index objects associated with the corresponding content block. The application constructs or updates various index objects by taking the content block as a unified unit, and unifies a version benchmark by the block structure snapshot and explicitly associates the relationship by the index dependency mapping, thereby effectively improving the index consistency. In addition, the application can determine the index object to be reconstructed for the content block in a subsequent data maintenance process, so that fine-grained reconstruction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of knowledge base construction technology, and in particular to a multi-index knowledge base, as well as an update method and a query self-healing method based on the multi-index knowledge base. Background Technology

[0002] Existing knowledge base systems generally maintain chapter structure indexes, keyword indexes, entity relationship indexes, and vector indexes simultaneously. During the user query phase, candidate data is retrieved by combining multiple types of indexes, and the final question and answer results are then output by the result generation module.

[0003] At the index building, updating, and task recovery levels, existing technologies mostly adopt file-level or batch-level overall reconstruction methods, or only perform independent incremental updates on a single index; after an abnormal task interruption, the general approach is to use task-level overall status recording for restart and recovery; during the query application phase, the system assumes that various index data always remain consistent, lacking proactive cross-index consistency verification and anomaly repair mechanisms. Summary of the Invention

[0004] This invention addresses the shortcomings of existing multi-index knowledge base systems in data synchronization construction and querying processes, such as difficulty in ensuring index consistency, coarse reconstruction granularity, and low interruption recovery efficiency. It provides a multi-index knowledge base and its self-healing update and query method.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A multi-index knowledge base, the multi-index knowledge base comprising several content blocks, each content block having a corresponding block structure snapshot and index dependency mapping relationship; The segmented structure snapshot is used to indicate the document to which the content segment belongs, as well as the chapter path, segmentation order, and content fingerprint corresponding to the content segment; The index dependency mapping relationship is used to indicate several index objects associated with the corresponding content chunks.

[0006] As one possible implementation method: The block structure snapshot includes the following fields: The content includes: chunk identifier, the document identifier to which the chunk belongs, the chapter path, the chunk order, the hash value of the chunk text content as a content fingerprint, the snapshot version number and the last update timestamp, as well as the index attributes of various index objects. Index types include chapter structure indexes, keyword indexes, entity relationship indexes, and / or vector indexes.

[0007] As one possible implementation method: The index dependency mapping relationship uses the corresponding block identifier as the key to record the set of index object identifiers associated with the content block, and the index types corresponding to the set of index object identifiers are different for each other.

[0008] As one possible implementation, the content segmentation is an independent content unit obtained by segmenting the document using chapter boundaries, paragraph boundaries, or sentence boundaries as segmentation boundaries.

[0009] Secondly, the update method for multi-index knowledge bases includes the following steps; The document to be processed is split into several corresponding content blocks, and a snapshot of the block structure corresponding to each content block is generated. Based on the historical block structure snapshots and the current block structure snapshots corresponding to the document to be processed, determine the addition, deletion, and modification of blocks; Each newly added block, each deleted block, and each modified block is taken as the target block, the index type affected by the target block is taken as the target index type, and the modification task is based on the index dependency mapping relationship of the target block and / or the corresponding target index type. Construct corresponding task sets based on each change task; Update the index based on the task set.

[0010] As one possible implementation, a local construction task is performed on the task set, dividing the local index construction process into a structure synchronization stage and an index writing stage, and recording stage checkpoints containing input hash values ​​in each stage; When the task resumes, the input hash value is recalculated, and a determination is made as to whether to reuse historical results based on the input hash value and the input hash value recorded at the previous stage checkpoint.

[0011] Thirdly, the query self-healing method of multi-indexed knowledge bases includes the following steps; Based on the query request, perform multi-index retrieval to obtain several corresponding candidate blocks and construct a candidate result set; Based on the block structure snapshot and index dependency mapping relationship, perform cross-index consistency verification on each candidate block, and identify abnormal candidate blocks based on the verification results; Perform local repairs on abnormal candidate blocks, and update the candidate result set based on the repair results; Once all abnormal candidate blocks have undergone partial recovery processing, the corresponding query results are generated based on the updated candidate result set.

[0012] As one possible implementation method; The block structure snapshot includes the index attributes of various index objects; The steps for performing cross-index consistency verification on the current candidate block are as follows: Read the index attributes of various index objects of the current candidate block from the block structure snapshot; Based on the index dependency mapping relationship, obtain various index objects associated with the current candidate block, and obtain several objects to be verified; Determine the existence of various index objects in their corresponding index storage to obtain the first verification result; Based on the index type classification, calculate the attribute values ​​corresponding to each type of verification object, obtain the corresponding verification attributes, perform attribute verification based on the corresponding index attributes and verification attributes, and obtain the second verification result; If the first or second verification result fails, the candidate block is marked as an abnormal candidate block; if all verifications pass, it is marked as a candidate block that has passed the consistency verification.

[0013] As one possible implementation method: The chapter path serves as the index attribute for the chapter structure index under the corresponding content block; The index attributes also include the hash values ​​of the keyword set, the entity set, the vector model, and the configuration parameters; The hash value of the keyword set is used to represent the index attribute corresponding to the keyword index under the corresponding content block; The hash value of the entity set is used to characterize the index attribute corresponding to the entity index under the corresponding content block; The hash value of the vector model and configuration parameters represents the index attribute corresponding to the vector index under the corresponding content block.

[0014] As one possible implementation, local repair is performed on the current abnormal candidate block. The specific steps are as follows: The index objects to be repaired associated with the current candidate anomaly block are determined based on the index dependency mapping relationship; The reconstruction operations of the index objects are performed sequentially according to a preset reconstruction order, which includes: chapter structure objects, keyword objects, entity relationship objects, and vector objects. After the index object is rebuilt, update the version number of the block structure snapshot corresponding to the abnormal candidate block.

[0015] This invention, by adopting the above technical solutions, has significant technical effects: The proposed multi-index database constructs or updates various index objects using content blocks as a unified unit, and unifies version benchmarks through block structure snapshots and clarifies the relationship through index dependency mapping, effectively improving index consistency. It also enables fine-grained reconstruction by determining the index objects to be rebuilt for content blocks during subsequent data maintenance. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a method for updating a multi-index knowledge base according to the present invention. Figure 2 This is a flowchart illustrating a self-healing query method for a multi-index knowledge base according to the present invention. Detailed Implementation

[0018] The present invention will be further described in detail below with reference to the embodiments. The following embodiments are explanations of the present invention, but the present invention is not limited to the following embodiments.

[0019] Existing multi-index knowledge base systems suffer from problems such as difficulty in guaranteeing index consistency and overly coarse reconstruction granularity, specifically including: The reconstruction granularity is too coarse. Existing technologies only perform index reconstruction at the document level or batch level, which cannot locate the affected fine-grained content units, resulting in high consumption of computing resources and low construction efficiency. The lack of unified consistency control in indexes means that chapter structure indexes, keyword indexes, entity relationship indexes, and vector indexes are constructed independently in existing technologies, and there is a lack of unified data association and version control mechanisms between them, which can easily lead to inconsistencies in cross-index data. To address the aforementioned shortcomings, this application proposes a multi-index knowledge base that uses content blocks as the smallest management unit. It generates content blocks through document parsing and segmentation, constructs block structure snapshots, and establishes multi-index dependency mappings. In subsequent updates, it accurately locates affected index objects by identifying differences between old and new snapshots and using index dependency mapping relationships. It only performs local construction and recovery on affected index objects, resulting in lower resource consumption and higher construction efficiency compared to existing technologies.

[0020] The multi-index knowledge base includes several content blocks, each of which has a corresponding block structure snapshot and index dependency mapping relationship; In this embodiment, the index types include chapter structure index, keyword index, entity relationship index, and / or vector index.

[0021] The content segmentation is an independent content unit obtained by segmenting the document using chapter boundaries, paragraph boundaries, or sentence boundaries as segmentation boundaries.

[0022] The segmented structure snapshot is used to indicate the document to which the content segment belongs, as well as the chapter path, segmentation order, and content fingerprint corresponding to the content segment. The content fingerprint is the hash value of the corresponding segment text content, which can be calculated based on SHA-256 for example. Chapter paths and chunk order are used to indicate the location of content chunks in the source document, for example, to identify changes in chapter structure; Content fingerprints are used to identify the text content of content blocks and to recognize changes in the text within a block. For example, when the content corresponding to the block's main text, table transcribing, or summary content changes, the hash value of the corresponding content block will also change.

[0023] As one possible implementation, the block structure snapshot includes the following fields: The table includes the chunk identifier, the document identifier to which the chunk belongs (e.g., " / Chapter 1 / Section 2"), the chapter path, the chunk order, the hash value of the chunk text content, the snapshot version number (monotonically increasing), the last update timestamp, and the index attributes of various index objects.

[0024] The segment identifier is a unique identity code for each content segment in the multi-index knowledge base. It is used to uniquely identify a single content segment and serves as a unified primary key for linking and binding segment structure snapshots, index dependency mappings, stage checkpoints, and various index objects.

[0025] In this embodiment, the chapter path serves as the index attribute of the chapter structure index under the corresponding content block; The index attributes also include the hash values ​​of the keyword set, the entity set, the vector model, and the configuration parameters; The hash value of the keyword set is the hash value of all keyword objects in the corresponding content block, which is used to represent the index attribute corresponding to the keyword index under the corresponding content block; The hash value of the entity set is the hash value of all entity objects in the corresponding content block, which is used to represent the index attribute corresponding to the entity index under the corresponding content block; The vector model is used to extract vectors from content blocks to generate corresponding vector objects. The vector model and configuration parameters (model parameters) can affect the vector extraction of content blocks and the generation of vector objects. Therefore, in this embodiment, the hash value based on the vector model and configuration parameters represents the index attribute corresponding to the vector index under the corresponding content block.

[0026] The index dependency mapping relationship is used to indicate several index objects associated with the corresponding content chunks.

[0027] The index dependency mapping relationship uses the corresponding block identifier as the key to record the set of index object identifiers associated with the content block. When a content block changes, the index dependency mapping relationship can be used to identify and locate the index object that needs to be reconstructed. For example, the index dependency mapping relationship is used to record the association between the content block and the keyword, and is used to represent the keyword index object corresponding to the content block.

[0028] The multi-index database proposed in this embodiment constructs or updates various index objects by using content blocks as a unified unit, and unifies the version benchmark through block structure snapshots and clarifies the relationship through index dependency mapping, which effectively improves index consistency. It also enables the determination of the index objects to be rebuilt for content blocks in the subsequent data maintenance process to achieve fine-grained reconstruction.

[0029] In addition to the coarse-grained reconstruction, existing technologies also suffer from coarse-grained recovery mechanisms. Specifically, existing technologies only record task-level states and cannot perform precise recovery based on blocks or stages. After an interruption, the completed parts need to be re-executed. Furthermore, the lack of a fine-grained recovery mechanism based on data states means that intermediate results cannot be reused after a task is interrupted, leading to redundant calculations and wasted system resources. To address the aforementioned shortcomings, this application proposes a method for updating a multi-index knowledge base. It uses content chunks as the smallest management unit, generating content chunks through document parsing and segmentation, constructing chunk structure snapshots, and establishing multi-index dependency mappings. It identifies changed chunks by utilizing the differences between old and new snapshots, accurately deriving the affected index subsets. The construction process is divided into two stages: structure synchronization and index writing, with breakpoint resume functionality achieved through stage checkpoints and input hashes. Figure 1 As shown, it includes the following steps: S100: Segment the document to be processed to obtain several corresponding content blocks, and generate a snapshot of the block structure corresponding to each content block; S110. Obtain the document to be processed; The document to be processed includes a target document and a document identity identifier of the target document. In this embodiment, when the target document is a newly added document, the corresponding document identity identifier of the target document is used to form a document to be processed. When the target document is an updated document, the document identity identifier corresponding to the target document is obtained and marked to form a corresponding document to be processed. This is prior art, so it is not limited in detail.

[0030] In this embodiment, the target documents include documents in formats such as PDF, DOCX, and Markdown.

[0031] S120. Segment the document to be processed; In this embodiment, the document to be processed is parsed, chapters are identified, and content is segmented to obtain several corresponding content blocks; Specifically: S121, Document parsing and chapter recognition; In this embodiment, after receiving the document to be processed, document parsing and chapter recognition are performed to generate chapter structure information and construct a complete chapter structure tree. The chapter structure tree is used to indicate the chapter level. For example, the chapter structure tree includes several branch nodes, each branch node represents a chapter, and each branch node includes at least one leaf node, each leaf node represents a section.

[0032] Specifically, the document to be processed is parsed to identify its main text structure, heading structure, hierarchical directory, chapter titles, page number ranges, and other structural information, and a corresponding chapter structure tree is generated.

[0033] S122. Based on the chapter recognition results, the document to be processed is segmented into multiple independent content blocks; All subsequent index building, version comparison, and difference identification will use content blocks as the smallest processing unit; In this embodiment, chapters are divided into blocks according to preset division rules such as boundaries and upper limits of segment length, resulting in multiple content blocks. Those skilled in the art can design their own division rules based on actual needs, such as using chapter boundaries, paragraph boundaries, and sentence boundaries as boundaries. This specification does not need to limit them in detail.

[0034] S130. Determine the block identifier corresponding to each content block and construct a block structure snapshot; This embodiment determines the block identifier corresponding to each content block and constructs / or updates the block structure snapshot, using the block structure snapshot as a unified benchmark for multi-index version comparison, difference identification and consistency verification; The segment identifier is a unique identity code for each content segment in the multi-index knowledge base. It is used to uniquely identify a single content segment and serves as a unified primary key for the segment structure snapshot, index dependency mapping relationship, stage checkpoint and association with various index objects. The segmented structure snapshot is used to indicate the document to which the content segment belongs (i.e., the corresponding document to be processed), as well as the chapter path, segmentation order, and content fingerprint corresponding to the content segment.

[0035] In this embodiment, the segmented structure snapshot includes the following fields: segment identifier, document identifier (document identity identifier) ​​to which the segment belongs, chapter path (e.g., " / Chapter 1 / Section 2"), segment order, hash value (SHA-256) of the segment text content, i.e., content fingerprint, hash value of keyword set, hash value of entity set, hash value of vector model and configuration parameters (used to identify specific vector model and corresponding model parameters), snapshot version number (monotonically increasing) and last update timestamp information.

[0036] When the document to be processed is a newly added document, a new block identifier is assigned to each content block, and a new block structure snapshot is built.

[0037] When the document to be processed is an updated document, the corresponding block identifier is determined based on the document identifier, chapter path, block order, and / or hash value of the block text content, and the corresponding block structure snapshot is updated based on the block identifier. As an example: Search for historical block structure snapshots based on the hash value of the block text content. When a corresponding historical block structure snapshot is found, update the block structure snapshot using the block identifier of the historical block structure snapshot. When no corresponding historical block structure snapshot is found based on the hash value of the block text content, the corresponding historical block structure snapshot is determined based on the chapter path and block order. When a matching historical block structure snapshot is available, the block structure snapshot is updated using the block identifier of the historical block structure snapshot. If no corresponding historical block structure snapshot is found based on the chapter path and block order, a new block identifier is generated for the corresponding content block, and a corresponding block structure snapshot is added based on the newly generated block identifier.

[0038] S200. Based on the historical block structure snapshot corresponding to the document to be processed and the current block structure snapshot, determine the addition, deletion and modification of blocks; This embodiment uses content blocks as the basic object, extracts chapter structure, keywords, entity relationships and generates vector features, and constructs a multi-index relationship of chapter structure index, keyword index, entity relationship index and vector index respectively. This is the prior art and will not be described in detail in this specification. The specific steps are as follows: S210. Obtain a snapshot of the historical block structure corresponding to the document to be processed; That is, based on the document identity of the document to be processed, search for the corresponding historical chunk structure snapshot, where the historical chunk structure snapshot is the same as the chunk structure snapshot corresponding to the previous version. S220. Compare the historical block structure snapshot corresponding to the document to be processed with the current block structure snapshot, and identify newly added blocks, deleted blocks and modified blocks based on the comparison results. S221, Add a new block; The newly added chunk is a content chunk in the document to be processed that lacks a corresponding historical chunk structure snapshot; When the document to be processed is a newly added document, it lacks a snapshot of the historical chunk structure. In this case, all its content chunks are judged as newly added chunks. The method for constructing the multi-index relationship corresponding to the newly added block is as follows: Synchronously create various index objects corresponding to the newly added blocks, and establish the index dependency mapping relationship between the content blocks and various index objects; Index objects may include chapter structure objects, keyword objects, document keyword association objects, entity objects, relation objects, and vector objects. Among them, chapter structure objects include article objects, chapter objects, paragraph objects, etc., which record the hierarchical information of article and chapter structure. Keyword objects are used to indicate the keyword nodes in the corresponding knowledge base. Entity objects are used to indicate the entity nodes in the corresponding knowledge base. Relation objects are used to indicate the set of relation edges in the knowledge base. Vector objects are used to indicate the vector representation corresponding to the content block. Those skilled in the art can determine the constructed index objects based on the index types involved in their knowledge base. The construction methods for various types of index objects are all existing technologies, and this specification will not elaborate on them.

[0039] In this embodiment, the index dependency mapping relationship uses the block identifier as the key to record the set of multiple index object identifiers associated with the block. When a content block changes, relying on the index dependency mapping relationship, only the index object associated with the block is processed locally, without having to rebuild the entire document or the entire knowledge base.

[0040] S222, Delete block; The deleted chunk is a historical content chunk that only has a snapshot of its historical chunk structure and lacks a corresponding snapshot of its current chunk structure. This indicates that the content chunk in the document to be processed has been deleted.

[0041] S223, Modify block division; The change blocks can be divided into content change blocks, structure path change blocks, and only configuration change blocks; This embodiment identifies the type of changed chunk by comparing the content features (hash value of chunk text content), structural path features (chapter path), and construction configuration features (hash value of vector model and configuration parameters) of the corresponding historical chunk structure snapshot with the current chunk structure snapshot; During identification, a cascading comparison is performed based on the priority of content features, structural path features, and construction configuration features. The specific steps are as follows: When the hash value of the text content of a block changes, it is determined to be a block with changed content. This indicates that the main text content, table transcription content, or summary content of the block has changed. When the content hash value remains unchanged but the chapter path changes, it is determined to be a block with a change in structural path. This indicates that the chapter link, title path, hierarchical position, page number range, or structural affiliation of the corresponding content block has changed. When the content hash value and structural path characteristics remain unchanged, but the hash values ​​of the vector model and configuration parameters change, it is determined that only the configuration change block is constructed, which indicates that the construction parameters for generating the index have changed. As one possible implementation, the construction configuration features may also include construction rules such as chunk granularity, overlap length, extraction model, prompt word template, entity pattern configuration, vector model, or enabled index type. That is, the original document content and its structural position in the document remain unchanged, but any change in any of the above construction configuration features is considered as a construction configuration change only. Such changes result in the need to rebuild the index results corresponding to the chunks without re-parsing the document content itself.

[0042] S300. Based on each newly added block, each deleted block, and each modified block, construct the corresponding task set; The task set is used to indicate the need to add / delete / change corresponding index objects under the corresponding index type. Changing the index object can be achieved by deleting the previous index object and adding the changed index object. S310. Take each newly added block, each deleted block, and each changed block as the target block, determine the target index type of the target block according to the preset propagation rules, and make a change task based on the index dependency mapping relationship of the target block and / or the change task corresponding to the target index type. Index dependency mappings are used to indicate several index objects associated with a corresponding content chunk; Change tasks are used to indicate the addition / deletion of corresponding index objects.

[0043] According to the preset propagation rules, based on the change difference type (addition, deletion, and modification) of the target block, and combined with the corresponding index dependency mapping relationship, the impact of changes to a single content block is propagated to the corresponding index object. The propagation rules are used to indicate the index types affected by the change difference type. Those skilled in the art can set propagation rules according to actual needs to indicate the index types affected by various change difference types, and take the index type affected by the target block as the corresponding target index type. In this embodiment, the index types include chapter structure index, keyword index, entity relationship index, and vector index, specifically: For newly added blocks, the propagation scope is all index types. That is, based on the newly added blocks, corresponding change tasks are created according to various index types to create corresponding index objects. For deletion blocks, the propagation scope is also all index types. That is, based on the index dependency mapping relationship of the deletion block, corresponding change tasks are established for all index objects of the deletion block to delete the corresponding index objects. For content change blocks, the propagation scope is also all index types. This is because the main content block is the common input source for chapter structure index, keyword index, entity relationship index and vector index. After the content changes, all four types of indexes need to be rebuilt. Therefore, based on the index dependency mapping relationship of the content change block, it is necessary to establish corresponding change tasks for all index objects of the content change block in order to update the corresponding index objects. For blocks with structural path changes, the propagation scope is limited to chapter structure objects because the content of the blocks remains unchanged. The results of the three types of indexes—keywords, entity relationships, and vectors—are not affected by the structural path changes. Therefore, only the chapter structure objects associated with the blocks are included in the propagation scope. That is, based on the index dependency mapping relationship of the blocks with structural path changes, a corresponding change task is established for the chapter structure objects of the blocks with structural path changes to update the chapter structure objects. For blocks that only involve configuration changes, the propagation scope is limited to vector indices; that is, when the vector model or configuration parameters change, the propagation extends to the vector objects. As one possible implementation, the art can also propagate configuration items that change according to specific needs along the dependency mapping based on actual requirements, specifically: When the granularity of the chunk or the overlap length changes, the propagation to the re-segmented chunk and all its associated index objects is possible (in practical applications, the content features are compared first, and such content chunks will be determined as content-changed chunks or newly added chunks). When the extraction model, prompt word template, or entity schema configuration changes, the changes are propagated to the keyword object and entity relationship object. When an index type changes, the changes are propagated to all index objects corresponding to the newly added or deactivated index type.

[0044] S320. Construct corresponding task sets based on each changed task; The change tasks for each target block are summarized according to the index type, and after deduplication, corresponding task sets are formed. In this embodiment, the affected index object sets are selected according to chapter structure, keywords, entities, relations, and vectors, and then aggregated to generate the affected index subset to be partially updated and rebuilt, i.e., the task set.

[0045] Perform local build tasks only on the affected subset of indexes, skipping unaffected chunks and index objects, and avoiding invalid operations on irrelevant index objects.

[0046] S400, Update the index based on the task set; Based on the task set, corresponding index objects are deleted or constructed. When deleting the original index object, the dependency mapping between the content chunk and the original index object is removed; after constructing the new index object, the dependency mapping between the content chunk and the new index object is established or updated.

[0047] In summary, this embodiment uses content blocks as the basic synchronous management unit for a multi-index knowledge base. It accurately identifies changed content units based on block structure snapshots, performing partial construction and updates only on affected index subsets. This abandons the traditional document-level and batch-level full-scale reconstruction mode, effectively reducing the scope of knowledge base reconstruction, significantly decreasing the overhead of text parsing and semantic vectorization calculations, and database read / write I / O losses, thus significantly shortening index construction time. Simultaneously, it establishes dependency mapping relationships between chapter structure indexes, keyword indexes, entity relationship indexes, and vector indexes around unified content blocks, enabling collaborative updates of multiple indexes around the same block object. This establishes a unified data association and version control mechanism, effectively avoiding problems such as cross-index version inconsistencies, missing index structures, and mapping mismatches, thereby improving the overall accuracy of knowledge retrieval and question-answering results.

[0048] As one possible implementation, this embodiment performs a local construction task on a task set, divides the local index construction process into a structure synchronization stage and an index writing stage, records stage checkpoints containing input hash values ​​in each stage, and determines whether to reuse historical results based on the input hash values ​​and stage status when the task is resumed. Specifically: After the partial build task is triggered, based on the chapter structure information, content chunks, and corresponding chunk structure snapshots obtained from document parsing, the structure synchronization phase begins: Complete the writing of document structure objects, chapter structure objects, chunk objects, and the mapping relationship between chunks and chapters, and record the checkpoints and input hash values ​​of the structure synchronization phase. That is, based on the changes of the corresponding content chunks in the local build task, perform add, update, or delete operations on the corresponding document structure objects, chapter structure objects, chunk objects, and the mapping relationship between chunks and chapters to complete the structure synchronization.

[0049] After the structure synchronization phase is completed, the index writing phase begins: Keyword nodes, entity nodes, relation edges, and vector records are written, and the index writing stage checkpoints and input hash values ​​(the comprehensive hash value input in the stage) are recorded. Specifically, based on the block text, keyword information, entity information, relation information, and vector configuration parameters corresponding to the affected content blocks, the corresponding keyword nodes, entity nodes, relation edges, and vector records are generated or located, and combined with the historical index dependency mapping relationship to determine whether they belong to newly added, updated, or deleted objects, and then the corresponding write processing is performed.

[0050] The phase checkpoint data parameters include: a unique identifier for the construction task, phase type, phase status, hash value of the phase input, set of blocks involved in the phase, and timestamp of the checkpoint record.

[0051] As an example, when an exception occurs during the build task, determine whether to reuse historical results when the task resumes: When a build task is abnormally interrupted, fails to execute, is manually terminated, or needs to be resumed after a system restart, the input hash value for the current stage is regenerated and compared with the checkpoint record of the previous stage: If the input hash value of the current stage is consistent with the input hash value recorded at the previous stage checkpoint, and the stage status recorded at the previous stage checkpoint is completed, then the stage is skipped and the historical results are reused directly, such as skipping the structure synchronization stage and directly entering the index writing stage. If a previous checkpoint does not exist, the status is not completed, or the input hash value is inconsistent, the current stage is re-executed.

[0052] Each stage is not executed on the entire batch of documents, but rather on the set of affected content chunks corresponding to a local build task. The stage checkpoint records the set of chunk identifiers involved in that stage, and the input hash value is also generated based on the set of affected content chunks and their chunk structure snapshots. Therefore, during task recovery, the current stage is not re-executed on all documents, but only on the set of affected content chunks where checkpoints are incomplete, the status is abnormal, or the input hash values ​​are inconsistent. This allows the recovery to reflect both stage-level control and chunk-level scope limitation.

[0053] This embodiment divides the index building process into two independent stages: structure synchronization and index writing. It also introduces a stage checkpoint management mechanism with input hash values. When a task is interrupted, fails, or restarted, the completed intermediate execution results can be accurately reused based on the stage status and the input hash comparison results. Only the incomplete links are continued, realizing fine-grained breakpoint continuation at the block and stage levels. This avoids the overall task running repeatedly and improves the efficiency of task recovery and system resource utilization.

[0054] Existing query solutions based on multi-index databases lack a mechanism to verify the consistency of candidate data across multiple indexes during the query process, and cannot identify abnormal data caused by missing indexes, inconsistent versions, or incomplete index construction. Existing systems typically rely on offline reconstruction to complete index repair, and lack technical means to repair abnormal data in real time during the query process.

[0055] To address the aforementioned issues, this application also proposes a query self-healing method for multi-index knowledge bases. During the query phase, multi-dimensional cross-index consistency checks are performed on recalled candidate blocks to identify anomalous candidate blocks. Local online repair is then performed on the anomalous blocks according to a preset priority, updating snapshots and checkpoints and replacing candidate results. Only trusted and repaired blocks are used to generate the final query results; unrepaired anomalous candidate blocks do not participate in the result generation process. Figure 2 As shown, the specific steps include: S500: Based on the query request, perform multi-index recall to obtain several corresponding candidate blocks and construct a candidate result set; That is, after receiving a query request, the corresponding content blocks are recalled as candidate blocks based on the chapter structure index, keyword index, entity relationship index and vector index, and a candidate result set is constructed; this is existing technology and will not be described in detail in this specification.

[0056] S600: Perform cross-index consistency verification on each candidate block and identify abnormal candidate blocks based on the verification results; Based on the block structure snapshot and index dependency mapping relationship, cross-index consistency verification is performed on each candidate block. Specifically: Query the block structure snapshot and index dependency mapping relationship corresponding to the current candidate block; The following checks are performed sequentially: chapter structure object verification, keyword object verification, entity relationship object verification, vector object verification, and structure snapshot version consistency verification.

[0057] During the cross-index consistency verification process, for each candidate block, the following processing is performed: The chapter path corresponding to the current candidate block is obtained based on the block structure snapshot, namely, keyword hash, entity hash and vector configuration hash information; Based on the index dependency mapping relationship, obtain the various index objects associated with the current candidate block, and obtain several objects to be verified, that is, take the associated chapter structure objects, keyword objects, entity objects and vector objects as objects to be verified; Determine the existence of each object to be verified in its corresponding index storage to obtain the first verification result; Based on the attribute values ​​corresponding to various verification objects, a verification attribute corresponding to the index type is obtained. The verification attribute is then verified based on the index attribute to obtain a second verification result. If the first or second verification result fails, the candidate block is marked as an abnormal candidate block; if all verifications pass, it is marked as a candidate block that has passed the consistency verification.

[0058] As one possible implementation method, the existence determination is based on the index dependency mapping relationship; For each candidate block, first obtain the set of various index object identifiers associated with the block from the index dependency mapping, and then verify them one by one in the corresponding index storage using these identifiers as query conditions: Chapter structure object: Based on the chapter identifier or chapter path, query whether the corresponding node exists in the graph structure storage; Keyword object: Based on the keyword identifier, check if the corresponding record exists in the keyword index; Entity-relationship object: Based on the entity identifier and the relationship triplet identifier, query whether the corresponding record exists in the entity-relationship index; Vector object: Based on the block identifier or entity identifier, query whether the corresponding vector record exists in the vector storage.

[0059] If a record matching the target identifier exists in the index storage, the index object is determined to exist; otherwise, it is determined to be missing. When it is determined to be missing, the first verification result is failure.

[0060] After the existence check is completed, an attribute value consistency check is performed separately to distinguish cases where the object exists but the attribute values ​​are inconsistent with the snapshot records of the block structure.

[0061] As one possible implementation, attribute value verification refers to further verifying whether the key attributes of the object are consistent with those recorded in the block structure snapshot after finding the corresponding index object in the index storage, rather than simply verifying a single hash value. Specifically, the key fields of the current index object are recalculated for hash, and the calculation result is compared one by one with the hash values ​​of the corresponding fields in the block structure snapshot. Specifically, the verification field for chapter structure objects corresponds to the chapter path `structure_path`, the verification field for keyword objects corresponds to the hash value of the keyword set `keyword_hash`, the verification field for entity relationship objects corresponds to the hash value of the entity set `entity_hash`, the verification field for vector objects corresponds to the hash value of the vector model and configuration parameters `vector_config_hash`, and the version information corresponds to the snapshot version number `version` field. If the hash comparison results of all fields are consistent, the attribute value is determined to be consistent; if any field is inconsistent, the attribute value of the index object is determined to be outdated or abnormal.

[0062] For example, the keyword_hash in the chunk structure snapshot corresponds to the overall hash value of the keyword set associated with that content chunk. Therefore, during verification, all keyword objects corresponding to the candidate chunk should be obtained first, the keywords should be normalized and organized to form a keyword set, the hash of the keyword set should be recalculated, and then compared with the keyword_hash in the chunk structure snapshot.

[0063] Furthermore, cross-index consistency verification also includes stage checkpoint verification, that is, obtaining the stage checkpoint status corresponding to the candidate block, determining whether it is in the completed state, and marking it as an abnormal candidate block when it is not in the completed state. Furthermore, cross-index consistency verification also includes snapshot version verification, which verifies whether the structure snapshot version number of the candidate block is consistent with the version identifier of each index object based on the version identifier associated with the corresponding content block recorded in each index object. When they are inconsistent, they are also marked as abnormal candidate blocks.

[0064] In this embodiment, when the version number in the block structure snapshot is inconsistent with the version identifier of any index object, the version number of the block structure snapshot shall be used as the standard to trigger the rewrite operation of the index object, thereby ensuring the eventual consistency of the data.

[0065] S700: Perform local repair on abnormal candidate blocks, and update the candidate result set based on the repair results; After a block is identified as an abnormal candidate block through consistency verification, the subset of index objects to be repaired and the repair order are determined based on the index dependency mapping relationship.

[0066] For content blocks marked as anomalous candidate blocks, the following local repair process is performed: 1. Determine the index object to be repaired associated with the current candidate anomaly block based on the index dependency mapping relationship; 2. Perform the reconstruction operation of the index objects in the preset reconstruction order, which includes: chapter structure object, keyword object, entity relationship object and vector object; 3. After the index object is rebuilt, update the version number of the block structure snapshot corresponding to the current candidate block for anomalies; This embodiment adds a cross-index multi-dimensional consistency and integrity verification mechanism for candidate blocks during the query phase. It can proactively identify abnormal candidate blocks with missing indexes, outdated versions, and incomplete construction, effectively isolating retrievable index data from reliable and usable data, preventing abnormal and expired data from participating in the generation of question and answer results, and improving the reliability of online query services from the source. Furthermore, it builds a query-driven online self-healing closed loop based on offline index maintenance, without triggering offline reconstruction of the entire database. It can perform local index repair in real time for abnormal blocks hit by the query, and immediately write back and update the repair results to the current query candidate process, realizing dynamic perception and online self-healing of index anomalies.

[0067] As one possible implementation, after the deletion, construction or repair of the index object in the corresponding stage is completed, the corresponding stage checkpoint is updated synchronously to reflect the current execution result and input consistency status of the stage. The updated content includes at least the stage status and input hash value, and may further include information involving the block identifier set, checkpoint record time and version information.

[0068] When the same content chunk is updated simultaneously by both the offline build task (the update method described above) and the online query repair operation (the current query cure method), an optimistic locking mechanism is employed: 4. Add the repaired abnormal candidate blocks back to the candidate result set and replace the results corresponding to the original abnormal candidate blocks; As an example, consider the concurrent scenario where offline build tasks and online repair operations simultaneously update the same content chunk: Before each update to the content chunks and associated index objects, the current chunk structure snapshot version number is read. If the version number has been modified and updated by other transactions, a concurrency conflict is determined, and the current transaction chooses to automatically retry or directly abandon the execution. After the online partial repair operation is completed, the version number of the chunk structure snapshot is automatically incremented and the last update timestamp is refreshed.

[0069] During the partial repair process, if an index write failure is detected or a preset processing time threshold is exceeded, the current block repair operation is terminated and the content block is removed from the candidate result set. If the repair fails due to model service unavailability, database write timeout, or network anomaly, the block will be marked as an unrepaired anomaly candidate block and will not participate in the generation of this result. The system generates answers based only on candidate blocks that have passed consistency checks and repaired candidate blocks; unrepaired abnormal candidate blocks are not included in the result generation process.

[0070] In this embodiment, a timeout threshold is set for a single local repair, which is 500 milliseconds by default. If the timeout is exceeded, the process will be forcibly interrupted to avoid blocking the online query link.

[0071] S800: After all abnormal candidate blocks have undergone local recovery processing, the corresponding query results are generated based on the updated candidate result set. That is, sorting or fusion processing is performed based on the updated candidate result set to generate the final query result. This is existing technology and will not be described in detail in this application.

[0072] In summary, this embodiment adopts an isolation and control strategy for abnormal candidate blocks that fail to be repaired in scenarios such as local repair timeout, model service anomaly, database write failure, and network failure. It restricts unrepaired abnormal blocks from participating in result generation and only relies on the fusion of qualified and successfully repaired candidate blocks to generate the final answer. This effectively reduces the probability of the system outputting incorrect results based on inconsistent abnormal data and further enhances the stability and operational reliability of the multi-index knowledge base query service.

[0073] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0074] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0075] This invention is described with reference to flowchart illustrations and / or block diagrams of the method, terminal device (system), and computer program product according to the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0076] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, the instruction means being implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0078] It should be noted that: The phrase "an embodiment" or "an embodiment" used in this specification means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. Therefore, the phrase "an embodiment" or "an embodiment" appearing in various places throughout the specification does not necessarily refer to the same embodiment.

[0079] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments.

[0080] Furthermore, it should be noted that the specific embodiments described in this specification may differ in the shape and name of their components. All equivalent or simple variations made to the structure, features, and principles described in this patent concept are included within the scope of protection of this patent.

Claims

1. A query self-healing method for multi-indexed knowledge bases, used in the query phase. Its characteristics are: The multi-index knowledge base includes several content blocks, each of which has a corresponding block structure snapshot and index dependency mapping relationship; The segmented structure snapshot is used to indicate the document to which the content segment belongs, as well as the chapter path, segmentation order, and content fingerprint corresponding to the content segment, and also includes the index attributes of various index objects; The index dependency mapping relationship is used to indicate several index objects associated with the corresponding content chunk; The steps to find self-healing methods include: Based on the query request, perform multi-index retrieval to obtain several corresponding candidate blocks and construct a candidate result set; Based on the block structure snapshot and index dependency mapping relationship, perform cross-index consistency verification on each candidate block, and identify abnormal candidate blocks based on the verification results; Perform local repairs on abnormal candidate blocks, and update the candidate result set based on the repair results; When all abnormal candidate blocks have undergone partial recovery processing, the corresponding query results are generated based on the updated candidate result set. The steps for performing cross-index consistency verification on the current candidate block are as follows: Read the index attributes of various index objects of the current candidate block from the block structure snapshot; Based on the index dependency mapping relationship, obtain various index objects associated with the current candidate block, and obtain several objects to be verified; Determine the existence of various index objects in their corresponding index storage to obtain the first verification result; Based on the index type classification, calculate the attribute values ​​corresponding to each type of verification object, obtain the corresponding verification attributes, perform attribute verification based on the corresponding index attributes and verification attributes, and obtain the second verification result; If the first verification result or the second verification result is unsuccessful, the candidate block will be marked as an abnormal candidate block. If all checks pass, the block is marked as a candidate block that has passed the consistency check.

2. The query self-healing method for a multi-indexed knowledge base according to claim 1, characterized in that: The snapshot of the block structure also Includes the following fields: The information includes the chunk identifier, the document identifier to which the chunk belongs, the chapter path, the chunk order, the hash value of the chunk text content as a content fingerprint, the snapshot version number, and the last update timestamp.

3. The query self-healing method for a multi-indexed knowledge base according to claim 2, characterized in that: The index dependency mapping relationship uses the corresponding block identifier as the key to record the set of various index object identifiers associated with the content block, and the index types corresponding to the set of index object identifiers are different from each other; The index dependency mapping relationship is used to indicate several index objects associated with the corresponding content blocks. The index types include chapter structure index, keyword index, entity relationship index and / or vector index.

4. The query self-healing method for a multi-indexed knowledge base according to claim 3, characterized in that: The chapter path serves as the index attribute for the chapter structure index under the corresponding content block; The index attributes also include the hash values ​​of the keyword set, the entity set, the vector model, and the configuration parameters; The hash value of the keyword set is used to represent the index attribute corresponding to the keyword index under the corresponding content block; The hash value of the entity set is used to characterize the index attribute corresponding to the entity index under the corresponding content block; The hash value of the vector model and configuration parameters represents the index attribute corresponding to the vector index under the corresponding content block.

5. The query self-healing method for a multi-indexed knowledge base according to claim 4, characterized in that, The specific steps for performing local repair on the current abnormal candidate blocks are as follows: The index objects to be repaired associated with the current candidate anomaly block are determined based on the index dependency mapping relationship; The reconstruction operations of the index objects are performed sequentially according to a preset reconstruction order, which includes: chapter structure objects, keyword objects, entity relationship objects, and vector objects. After the index object is rebuilt, update the version number of the block structure snapshot corresponding to the abnormal candidate block.

6. The query self-healing method for a multi-indexed knowledge base according to claim 1, characterized in that, The content segmentation is an independent content unit obtained by dividing the document using chapter boundaries, paragraph boundaries, or sentence boundaries as segmentation boundaries.

Citation Information

Patent Citations

  • Office knowledge base-oriented cloud document duplicate removal archiving and tracing management method and system

    CN122173448A