Context reconstruction method and system based on parent-child document matching

By analyzing and integrating the hierarchical structure and content characteristics of parent-child documents, dynamic semantic alignment processing is carried out, and reconstructed context data blocks are generated, which solves the problem of fragmentation and inconsistency of document context information in the existing technology, and realizes efficient and accurate document content reconstruction and version update.

CN120106083AActive Publication Date: 2025-06-06JIEHELIX (SHANGHAI) MEDICAL TECH CO LTD

Patent Information

Application Number
CN202510573612.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-06
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

Existing document processing technologies are difficult to effectively integrate and process related documents, and cannot fully utilize the context information in the child document to enrich and improve the content of the parent document, resulting in fragmentation and incompleteness of the parent document.

Method used

By obtaining the target parent document and its associated sub-document collection, parsing the context paragraphs and paragraph identifiers in the sub-document unit, determining the hierarchical affiliation parameters between the sub-document and the parent document, performing dynamic semantic alignment processing, generating reconstructed context data blocks, and outputting an optimized version to trigger version updates.

Benefits of technology

It significantly improves the accuracy and efficiency of document context reconstruction, ensures the logical and readability of the reconstructed document content, and solves the problems of fragmentation of context information and semantic inconsistency in traditional technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106083A_ABST
    Figure CN120106083A_ABST
Patent Text Reader

Abstract

The invention provides a context reconstruction method and system based on father-child document matching, and the method comprises the steps: firstly obtaining a target father document and an associated child document set thereof, the child document set comprises a plurality of child document units, and each unit comprises a context paragraph and a paragraph identifier; extracting global structure features including paragraph hierarchical distribution and title nesting depth, and local content features including paragraph core word sequences and semantic coherence indexes; determining hierarchical membership relation parameters of the sub-document unit paragraph identifiers and the parent document according to the sub-document unit paragraph identifiers, matching context paragraphs with corresponding paragraph hierarchies, performing dynamic semantic alignment processing including paragraph boundary calibration and semantic redundancy elimination on the context paragraphs on the basis of the characteristics, generating reconstructed context data blocks, and storing the reconstructed context data blocks in the parent document. And generating a target parent document optimization version and outputting the target parent document optimization version to a document storage system to trigger version updating operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and in particular to a context reconstruction method and system based on parent-child document matching. Background Art

[0002] In the current document processing and information management field, with the explosive growth of information volume and the increasing complexity of document structure, how to efficiently and accurately reconstruct the context information of documents to improve the readability, logic and maintainability of documents has become a key issue that needs to be solved. However, existing document processing technologies have significant limitations in this regard.

[0003] On the one hand, existing document processing methods often focus on parsing and processing a single document, while ignoring the possible hierarchical affiliation and contextual associations between documents. In practical applications, many documents do not exist in isolation, but are related to each other in the form of parent documents and child documents, forming a complete information system. Existing technologies lack the ability to effectively integrate and process these related documents, and cannot fully utilize the contextual information in child documents to enrich and improve the content of parent documents, resulting in fragmented and incomplete contextual information of parent documents.

[0004] On the other hand, even if some technologies attempt to process related documents, most of them remain at the level of simple text splicing or content duplication, lacking in-depth exploration of the document's hierarchical structure and semantic information. These technologies cannot accurately identify the hierarchical affiliation between child documents and parent documents, and cannot accurately match and dynamically align the context paragraphs of child documents based on this relationship. Therefore, when reconstructing the context information of the parent document, problems such as paragraph dislocation and semantic incoherence often occur, which seriously affects the quality and readability of the document.

[0005] Furthermore, existing document processing technologies usually use static semantic analysis methods when semantically aligning documents, and cannot dynamically adjust according to the hierarchical structure and content characteristics of the document. This static semantic alignment method cannot effectively deal with the semantic redundancy and boundary ambiguity in the document, resulting in a large amount of redundancy and ambiguity in the reconstructed context information, which reduces the information density and usability of the document.

[0006] Finally, in terms of document version updates, existing technologies mostly rely on manual operations or simple automated scripts and lack intelligent version update mechanisms. When the context information of the parent document changes, it is impossible to generate an optimized version and trigger a version update operation in a timely and automatic manner, resulting in inefficient document version management and prone to version confusion and information inconsistency. Summary of the invention

[0007] In view of this, the purpose of the present application is to provide a context reconstruction method and system based on parent-child document matching.

[0008] According to a first aspect of the present application, a context reconstruction method based on parent-child document matching is provided, the method comprising: Obtaining a target parent document and its associated sub-document set, wherein the sub-document set includes a plurality of sub-document units, each of which includes at least one context paragraph and a corresponding paragraph identifier; Performing hierarchical structure analysis on the target parent document to extract global structural features and local content features of the target parent document, wherein the global structural features include paragraph hierarchical distribution information and title nesting depth information, and the local content features include paragraph core word sequences and semantic coherence indicators; Determine, according to the paragraph identifier of the sub-document unit, a hierarchical affiliation parameter between the sub-document unit and the target parent document, and match the context paragraph of the sub-document unit with a corresponding paragraph level of the target parent document based on the hierarchical affiliation parameter; Based on the global structural features and the local content features, dynamic semantic alignment processing is performed on the context paragraphs of the sub-document unit to generate a reconstructed context data block, wherein the dynamic semantic alignment processing includes paragraph boundary calibration and semantic redundancy elimination; An optimized version of the target parent document is generated according to the reconstructed context data block, and the optimized version is output to a document storage system to trigger a version update operation.

[0009] In a possible implementation of the first aspect, performing hierarchical structure parsing on the target parent document to extract global structural features and local content features of the target parent document includes: Identify a title tag sequence in the target parent document, and generate paragraph level distribution information according to the nesting level of the title tag sequence, wherein the paragraph level distribution information includes the starting position, the ending position and the corresponding title level number of each paragraph; Performing semantic segmentation processing on each paragraph of the target parent document to extract a core word sequence of the paragraph, wherein the core word sequence of the paragraph is composed of words in the paragraph that meet a word frequency threshold and an inverse document frequency threshold; Calculating the semantic coherence index between adjacent paragraphs, wherein the semantic coherence index is obtained by weighting the overlap of core word sequences of adjacent paragraphs and the cosine similarity of semantic vectors; The title nesting depth information is determined according to the maximum value of the title level number, and the paragraph level distribution information, title nesting depth information, paragraph core word sequence and semantic coherence index are written into the structure feature cache and the content feature cache respectively.

[0010] In a possible implementation of the first aspect, determining the hierarchical affiliation parameter between the sub-document unit and the target parent document according to the paragraph identifier of the sub-document unit includes: Parsing the paragraph identifier of the sub-document unit, extracting the parent document reference number, the target paragraph level number and the paragraph position offset in the paragraph identifier; Matching the paragraph level distribution information corresponding to the parent document reference number from the structural feature cache to determine the maximum position offset range allowed by the target paragraph level number; If the paragraph position offset is within the maximum position offset range, then the structural similarity and content overlap between the context paragraph of the sub-document unit and the corresponding paragraph level of the target parent document are calculated, wherein the structural similarity is determined by the consistency of the title level number and the paragraph length ratio, and the content overlap is determined by the number of overlapping words in the core word sequence of the paragraph and the change in the semantic coherence index; A hierarchical affiliation parameter is generated according to the weighted sum of the structural similarity and the content overlap, and if the hierarchical affiliation parameter exceeds a preset affiliation threshold, the sub-document unit is marked as a valid matching unit.

[0011] In a possible implementation of the first aspect, the matching of the context paragraph of the sub-document unit with the corresponding paragraph level of the target parent document based on the level affiliation parameter includes: For each of the sub-document units marked as valid matching units, extracting the paragraph core word sequence and semantic coherence index of the paragraph level corresponding to the target parent document from the content feature cache; Performing synonym replacement detection on the context paragraph of the sub-document unit, generating a replacement word list, and updating the paragraph core word sequence according to the replacement word list; Calculating the dynamic overlap between the updated paragraph core word sequence and the paragraph core word sequence at the paragraph level corresponding to the target parent document, wherein the dynamic overlap is obtained by adjusting the word order consistency and the word meaning similarity; If the dynamic overlap degree is greater than or equal to a preset reconstruction threshold, the context paragraph of the sub-document unit is positionally bound to the corresponding paragraph level of the target parent document to generate a paragraph mapping relationship table.

[0012] In a possible implementation of the first aspect, the performing dynamic semantic alignment processing on the context paragraphs of the sub-document unit based on the global structural features and the local content features to generate a reconstructed context data block includes: According to the paragraph mapping relationship table, determining the paragraph level position at which the context paragraph of the sub-document unit needs to be inserted into the target parent document; Obtaining the semantic coherence index of the adjacent paragraphs at the paragraph level position from the structural feature cache, and calculating the insertion compatibility between the context paragraph and the adjacent paragraph, wherein the insertion compatibility is determined by the core word overlap rate and semantic vector direction consistency of the previous and next paragraphs; If the insertion compatibility exceeds a preset compatibility threshold, the paragraph boundary calibration is performed on the context paragraph, wherein the paragraph boundary calibration includes deleting the core word sequence repeated with the adjacent paragraph and adjusting the transition conjunctions of the starting sentence of the paragraph; Eliminate semantic redundancy on the calibrated context paragraphs to obtain processed context paragraphs; The processed context paragraph is encapsulated into a reconstructed context data block, and a version identifier and a timestamp are added to the reconstructed context data block.

[0013] In a possible implementation manner of the first aspect, performing semantic redundancy elimination on the calibrated context paragraph to obtain a processed context paragraph includes: Extract all sentences in the calibrated context paragraph to generate a sentence set; Performing semantic role labeling on each sentence in the sentence set, and extracting the subject-predicate-object structure and modifying components of the sentence; Calculate the semantic overlap between any two sentences, where the semantic overlap is determined by subject-verb-object structural consistency and modifier similarity; If the semantic overlap between two sentences exceeds a preset overlap threshold, the sentence with a higher density of core words is retained and the other sentence is deleted; The remaining sentences are checked for logical order, and the sentence arrangement order is adjusted according to the global structural features of the target parent document so that the semantic coherence index of adjacent sentences reaches a preset coherence threshold, thereby obtaining a processed context paragraph.

[0014] In a possible implementation of the first aspect, outputting the optimized version to a document storage system to trigger a version update operation includes: Parsing all reconstruction context data blocks in the optimized version, and extracting a version identifier and a timestamp of each reconstruction context data block; Acquire the current version metadata of the target parent document from the document storage system, wherein the current version metadata includes a historical version identifier set and a last updated timestamp; If the timestamp of the optimized version is later than the last updated timestamp, the optimized version is compared with the current version to generate a difference content report, which includes the number of newly added paragraphs, the positions of modified paragraphs, and the identifiers of deleted paragraphs; Generate a version update instruction according to the difference content report, wherein the version update instruction includes incremental update data and a rollback verification code; The version update instruction is sent to the document storage system, so that the document storage system replaces the target paragraph according to the incremental update data, and verifies the integrity of the replaced paragraph according to the rollback verification code.

[0015] In a possible implementation of the first aspect, sending the version update instruction to a document storage system so that the document storage system replaces a target paragraph according to the incremental update data and verifies the integrity of the replaced paragraph according to a rollback check code includes: Creating a temporary version branch in the document storage system, and writing the incremental update data into the temporary version branch; Performing an integrity check on the temporary version branch, the integrity check including a section identifier continuity check and a semantic logic conflict detection; If the integrity check passes, the temporary version branch is merged into the main version branch, and the last updated timestamp is updated; If a semantic logic conflict is detected, the write operation of the temporary version branch is undone according to the rollback verification code, and a manual intervention process is triggered.

[0016] In a possible implementation of the first aspect, before acquiring the target parent document and its associated child document set, the method further includes: Configuring a document association rule library, wherein the document association rule library defines matching conditions of parent and child documents, wherein the matching conditions include file format consistency, subject keyword overlap rate, and author authority identifier; Monitor the document input channel, and when a new document is detected to be uploaded, extract metadata features of the new document, wherein the metadata features include a creator identifier, a document subject tag, and a format type; Matching the metadata features of the new document with the metadata features of the existing parent document according to the document association rule library, and if the match is successful, adding the new document to the child document set of the corresponding parent document; If the match fails, create an independent parent document identifier for the new document and initialize its child document collection to be empty; The configuration document association rule base includes: Define a file format consistency rule, wherein the file format consistency rule requires that the format type of the parent and child documents is the same and the version compatibility identifier is consistent; Setting a topic keyword overlap rate threshold, wherein the topic keyword overlap rate threshold is dynamically adjusted according to the topic distribution statistics of the historical document collection; Bind author permission identifiers so that the creator identifier of a child document must be included in the list of authorized editors of the parent document; Assign a weight coefficient to each matching condition, and determine the matching priority of the parent and child documents based on the weighted sum of the weight coefficients; When multiple parent documents meet the matching conditions, the parent document with the highest matching priority is selected for association.

[0017] According to the second aspect of the present application, a context reconstruction system is provided, which includes a machine-readable storage medium and a processor, wherein the machine-readable storage medium stores machine-executable instructions, and when the processor executes the machine-executable instructions, the context reconstruction system implements the aforementioned context reconstruction method based on parent-child document matching.

[0018] According to a third aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed, the aforementioned context reconstruction method based on parent-child document matching is implemented.

[0019] According to any one of the above aspects, the technical effect of the present application is: The embodiment of the present application realizes the intelligent processing of the whole process from document structure analysis to content semantic alignment by constructing a context reconstruction method based on parent-child document matching, significantly improves the accuracy and efficiency of document context reconstruction, and provides an innovative technical solution for document content optimization and version update. Specifically, by obtaining the target parent document and its associated child document set, and parsing the context paragraphs and paragraph identifiers in the child document unit, the hierarchical affiliation between documents can be accurately grasped. On this basis, the hierarchical structure of the target parent document is parsed and processed to extract global structural features and local content features, which not only reveals the macro-architecture of the document, but also deeply analyzes the micro-semantics of the paragraph, providing multi-dimensional feature support for the dynamic semantic alignment of the context paragraph. By matching the context paragraph of the child document unit with the corresponding paragraph level of the target parent document through the hierarchical affiliation parameter, the precise positioning of cross-document paragraphs is achieved, effectively avoiding the dislocation and omission of context information. Furthermore, the context paragraphs of the sub-document units are dynamically semantically aligned based on global structural features and local content features. By calibrating paragraph boundaries and eliminating semantic redundancy, the semantic coherence and consistency of the reconstructed context data blocks are ensured, significantly improving the readability and logic of the document content. Finally, an optimized version of the target parent document is generated based on the reconstructed context data blocks, and a version update operation is triggered, effectively solving the problems of context information fragmentation and semantic inconsistency in traditional document processing methods, and improving the automation and intelligence level of document processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 A schematic diagram of a process flow of a context reconstruction method based on parent-child document matching provided in an embodiment of the present application is shown; Figure 2 A schematic diagram of the component structure of a context reconstruction system for implementing the above-mentioned context reconstruction method based on parent-child document matching provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0022] The embodiments of the present application are described below in conjunction with the drawings in the present application. It should be understood that the implementation methods described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.

[0023] It will be understood by those skilled in the art that, unless specifically stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application refer to that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the technical field. It should be understood that when an element is "connected" or "coupled" to another element, the element may be directly connected or coupled to the other element, or may refer to the element and the other element establishing a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling, and the term "and / or" used herein indicates at least one of the items defined by the term, such as "A and / or B" may be implemented as "A", or as "B", or as "A and B".

[0024] In order to make the purpose, technical solution and advantages of the present application clearer, the implementation mode of the present application will be further described in detail with reference to the accompanying drawings. The technical solution of the embodiment of the present application and the technical effect produced by the technical solution of the present application are explained below by describing several exemplary implementation modes. It should be pointed out that the following implementation modes can refer to, draw lessons from or combine with each other, and the same terms, similar features and similar implementation steps in different implementation modes are not described repeatedly.

[0025] Figure 1 The flowchart of the context reconstruction method and system based on parent-child document matching provided by the embodiment of the present application is shown. It should be understood that in other embodiments, the order of some steps in the context reconstruction method based on parent-child document matching of this embodiment can be shared with each other according to actual needs, or some steps can be omitted or maintained. The detailed steps of the context reconstruction method based on parent-child document matching include: Step S110: Acquire a target parent document and its associated sub-document set, wherein the sub-document set includes a plurality of sub-document units, and each sub-document unit includes at least one context paragraph and a corresponding paragraph identifier.

[0026] In this embodiment, taking the field of medical data retrieval as an example, the target parent document can be a review document on the diagnosis and treatment specifications of a certain disease. For example, the target parent document revolves around the diagnosis and treatment of heart disease, covering many aspects such as the diagnosis methods and treatment methods of heart disease. The associated sub-document set is composed of many research reports, case analysis and other documents related to the diagnosis and treatment of heart disease. Each sub-document unit, taking a research report on the therapeutic effect of a new drug for heart disease as an example, contains a context paragraph describing the drug experimental process, such as the dosage of the drug, the screening criteria for experimental subjects, etc. At the same time, each paragraph has a corresponding paragraph identifier, such as "P1-001" represents the first paragraph in the report describing the dosage of the drug, "P1-002" represents the second paragraph describing the screening criteria for experimental subjects, etc. These sub-document units together constitute a sub-document set.

[0027] Step S120: performing hierarchical structure analysis on the target parent document to extract global structural features and local content features of the target parent document, wherein the global structural features include paragraph hierarchical distribution information and title nesting depth information, and the local content features include paragraph core word sequences and semantic coherence indicators.

[0028] Next, the hierarchical structure of this summary document on the diagnosis and treatment of heart disease is parsed.

[0029] Step S121: Identify the title tag sequence in the target parent document, and generate paragraph level distribution information according to the nesting level of the title tag sequence, wherein the paragraph level distribution information includes the start position, end position and corresponding title level number of each paragraph.

[0030] For example, in this document summarizing the diagnosis and treatment standards for heart disease, the title tag sequence may be like "1. Diagnostic methods for heart disease", "1.1 Clinical symptom diagnosis", "1.2 Instrument detection diagnosis", etc. By identifying these title tag sequences, paragraph level distribution information is generated according to their nesting levels. For example, the starting position of the paragraph corresponding to "1. Diagnostic methods for heart disease" is assumed to be the 100th character of the document, and the ending position is the 500th character, and its title level number is 1; the starting position of the paragraph corresponding to "1.1 Clinical symptom diagnosis" is at the 501st character, and the ending position is at the 800th character, and the title level number is 2. In this way, the starting and ending positions of each paragraph and the corresponding title level number are determined in detail to form paragraph level distribution information.

[0031] Step S122: performing semantic segmentation processing on each paragraph of the target parent document to extract a paragraph core word sequence, wherein the paragraph core word sequence is composed of words in the paragraph that meet a word frequency threshold and an inverse document frequency threshold.

[0032] Take the paragraph "1.1 Clinical Symptom Diagnosis" as an example and perform semantic segmentation on it. Assume that the word frequency threshold is set to 5 times and the inverse document frequency threshold is set to 0.2. The content of this paragraph is "Common clinical symptoms of heart disease patients include chest pain, dyspnea, palpitations, etc., and chest pain is a more prominent symptom." After processing, words that meet the word frequency threshold of more than 5 times and the inverse document frequency threshold greater than 0.2, such as "chest pain", "dyspnea", "palpitations", etc., constitute the core word sequence of the paragraph.

[0033] Step S123: calculating the semantic coherence index between adjacent paragraphs, wherein the semantic coherence index is obtained by weighting the overlap of the core word sequences of the adjacent paragraphs and the cosine similarity of the semantic vectors.

[0034] Continuing with the review document of the diagnostic and treatment standards for heart disease as an example, assume that the adjacent paragraphs are "1.1 Clinical Symptom Diagnosis" and "1.2 Instrumental Detection Diagnosis". First, calculate the overlap of the core word sequences of the two paragraphs. For example, the core word sequence of "1.1 Clinical Symptom Diagnosis" is ["chest pain", "dyspnea", "palpitations"], and the core word sequence of "1.2 Instrumental Detection Diagnosis" is ["electrocardiogram", "cardiac ultrasound", "chest pain"], and the overlapping word is "chest pain". The overlap is assumed to be calculated as 1 / 3 (the number of overlapping words divided by the number of core words in "1.1 Clinical Symptom Diagnosis"). Then calculate the cosine similarity of the semantic vectors. Here, it is assumed that the cosine similarity of the semantic vectors of the two paragraphs is 0.6 through calculation. Then, through weighted calculation, assuming that the weights are 0.4 (coincidence weight) and 0.6 (semantic vector cosine similarity weight), the semantic coherence index is 0.4*1 / 3+0.6*0.6=0.4*0.333+0.6*0.6=0.133+0.36=0.493.

[0035] Step S124: determining the title nesting depth information according to the maximum value of the title level number, and writing the paragraph level distribution information, title nesting depth information, paragraph core word sequence and semantic coherence index into the structure feature cache and content feature cache respectively.

[0036] In the document reviewing the diagnosis and treatment of heart disease, assuming that the maximum value of the title level number is 3, such as "1. Diagnosis of heart disease" (level 1) - "1.1 Clinical symptom diagnosis" (level 2) - "1.1.1 Characteristics of chest pain" (level 3), the title nesting depth information is 3. Then the previously obtained paragraph level distribution information, title nesting depth information, core word sequence of each paragraph, and semantic coherence index of adjacent paragraphs are written into the structural feature cache and content feature cache respectively for use in subsequent steps.

[0037] Step S130: Determine the hierarchical affiliation parameters between the sub-document unit and the target parent document according to the paragraph identifier of the sub-document unit, and match the context paragraph of the sub-document unit with the corresponding paragraph level of the target parent document based on the hierarchical affiliation parameters.

[0038] Then, we will take the example that the sub-document unit is a detailed case analysis report on a specific treatment for heart disease.

[0039] Step S131: parsing the paragraph identifier of the sub-document unit, and extracting the parent document reference number, target paragraph level number and paragraph position offset in the paragraph identifier.

[0040] Assume that the paragraph identifier of the case analysis report is "P2-003-10-20", where "P2" represents the parent document reference number, corresponding to the review document of heart disease diagnosis and treatment standards; "003" represents the target paragraph level number, which is assumed to correspond to a certain level under "2. Treatment of heart disease" in the review document; "10" is the paragraph position offset, indicating the offset relative to the starting position of the target paragraph level.

[0041] Step S132: Match the paragraph level distribution information corresponding to the parent document reference number from the structural feature cache to determine the maximum position offset range allowed by the target paragraph level number.

[0042] From the information previously written into the structural feature cache, find the paragraph level distribution information of the heart disease diagnosis and treatment specification review document corresponding to the parent document reference number "P2". Assuming that the paragraph of the level "2. Treatment of heart disease" starts at the 1500th character of the document and ends at the 2500th character, according to the predefined rules (for example, setting the maximum offset range to 10% of the length of the paragraph at this level), the maximum position offset range allowed for the target paragraph level number is from 1500-(2500-1500)*0.1=1400 characters to 1500+(2500-1500)*0.1=1600 characters.

[0043] Step S133: If the paragraph position offset is within the maximum position offset range, then calculate the structural similarity and content overlap between the context paragraph of the sub-document unit and the corresponding paragraph level of the target parent document, wherein the structural similarity is determined by the consistency of the title level number and the paragraph length ratio, and the content overlap is determined by the number of overlapping words in the paragraph core word sequence and the change in the semantic coherence index.

[0044] If the paragraph position offset 10 is within the above maximum position offset range. Calculate the structural similarity, assuming that the target paragraph level number is 3, the corresponding level number of the sub-document unit is also 3, the title level number consistency is 1 (completely consistent), the sub-document unit context paragraph length is 200 characters, the target parent document corresponding paragraph level paragraph length is 500 characters, and the paragraph length ratio is 200 / 500=0.4, then the structural similarity is (1+0.4) / 2=0.7. Calculate the content overlap, the sub-document unit context paragraph core word sequence is ["specific drugs", "treatment effect", "recovery status"], the target parent document corresponding paragraph level core word sequence is ["specific drugs", "treatment process", "complications"], the number of overlapping words is 1, assuming that the change in the semantic coherence index of adjacent paragraphs is 0.1 (calculated by comparing with adjacent paragraphs), then the content overlap is 1 / 3+0.1=0.333+0.1=0.433.

[0045] Step S134: generating a hierarchical affiliation parameter according to the weighted sum of the structural similarity and the content overlap, and if the hierarchical affiliation parameter exceeds a preset affiliation threshold, marking the sub-document unit as a valid matching unit.

[0046] Assuming that the weight of structural similarity is 0.6 and the weight of content overlap is 0.4, the hierarchical affiliation parameter is 0.6*0.7+0.4*0.433=0.42+0.173=0.593. The preset affiliation threshold is assumed to be 0.5. Since 0.593 is greater than 0.5, the sub-document unit is marked as a valid matching unit.

[0047] Step S135: for each of the sub-document units marked as valid matching units, extract the paragraph core word sequence and semantic coherence index of the paragraph level corresponding to the target parent document from the content feature cache.

[0048] Taking the sub-document unit marked as a valid matching unit as an example, the core word sequence of the paragraph corresponding to the paragraph level (assuming it is the "2.1 Drug Treatment" level) of the heart disease diagnosis and treatment standard review document is extracted from the content feature cache, such as ["names of various drugs", "dosage", "side effects of drugs"] and semantic coherence index (for example, the semantic coherence index with the adjacent level paragraph is 0.5).

[0049] Step S136: Perform synonym replacement detection on the context paragraph of the sub-document unit, generate a replacement word list, and update the paragraph core word sequence according to the replacement word list.

[0050] In the context paragraph of the sub-document unit about specific drug treatment, for example, the paragraph content is "This special drug has a significant effect on the treatment of heart disease". After synonym replacement detection, it is found that "special drug" is a synonym for "drug" in the "names of various drugs" in the core word sequence of the target parent document, and a replacement word list ["special drug"-"drug"] is generated. Then, the core word sequence of the paragraph is updated according to the replacement word list, and "special drug" is replaced with "drug". The updated core word sequence of the paragraph becomes ["drug", "treatment effect", "recovery status"].

[0051] Step S137: Calculate the dynamic overlap between the updated paragraph core word sequence and the paragraph core word sequence at the paragraph level corresponding to the target parent document, wherein the dynamic overlap is obtained by adjusting the word order consistency and the word meaning similarity.

[0052] The updated paragraph core word sequence is ["drug", "treatment effect", "recovery status"], and the corresponding paragraph-level core word sequence of the target parent document is ["names of various drugs", "drug dosage", "drug side effects"]. Calculate the word order consistency, assuming that according to the predefined rules, the word order consistency calculated by the position difference of the same word "drug" in the two sequences is 0.6 (full score 1). Calculate the word meaning similarity, and assume that the word meaning similarity between "treatment effect" and "drug side effects" is 0.3, and the word meaning similarity between "recovery status" and "drug dosage" is 0.2. Through weighted calculation (assuming that the word order consistency weight is 0.6 and the word meaning similarity weight is 0.4), the dynamic overlap is 0.6*0.6+0.4*(0.3+0.2) / 2=0.36+0.4*0.25=0.36+0.1=0.46.

[0053] Step S138: If the dynamic overlap degree is greater than or equal to a preset reconstruction threshold, the context paragraph of the sub-document unit is positionally bound to the corresponding paragraph level of the target parent document to generate a paragraph mapping relationship table.

[0054] The preset reconstruction threshold is assumed to be 0.4. Since 0.46 is greater than or equal to 0.4, the context paragraph of the sub-document unit is positionally bound to the level "2.1 Drug Therapy" in the heart disease diagnosis and treatment standard review document, and a paragraph mapping relationship table is generated to record the specific position of the sub-document unit paragraph "P2-003" corresponding to the target parent document "2.1 Drug Therapy" level.

[0055] Step S140: Based on the global structural features and the local content features, dynamic semantic alignment processing is performed on the context paragraphs of the sub-document unit to generate a reconstructed context data block, wherein the dynamic semantic alignment processing includes paragraph boundary calibration and semantic redundancy elimination.

[0056] Next, take the sub-document unit context paragraph for which a position binding has been established as an example.

[0057] Step S141: According to the paragraph mapping relationship table, determine the paragraph level position of the target parent document into which the context paragraph of the sub-document unit needs to be inserted.

[0058] According to the generated paragraph mapping relationship table, it is determined that the context paragraph of the sub-document unit about specific drug treatment needs to be inserted into a specific position under the "2.1 Drug Treatment" level of the heart disease diagnosis and treatment specification review document, assuming it is in the middle of the paragraph at this level.

[0059] Step S142: obtaining the semantic coherence index of the adjacent paragraphs at the paragraph level position from the structural feature cache, and calculating the insertion compatibility between the context paragraph and the adjacent paragraphs, wherein the insertion compatibility is determined by the core word overlap rate and semantic vector direction consistency of the previous and next paragraphs.

[0060] Get the semantic coherence index of the adjacent paragraphs at the insertion position at the "2.1 Drug Therapy" level from the structural feature cache, assuming that the core word sequence of the previous paragraph is ["drug development background", "R&D team"], the core word sequence of the next paragraph is ["clinical trial process", "test results"], and the core word sequence of the sub-document unit context paragraph is ["drug ingredients", "treatment effect"]. Calculate the core word overlap rate of the previous paragraph and the sub-document unit context paragraph, the overlapping words are 0, and the overlapping rate is 0; the core word overlap rate of the next paragraph and the sub-document unit context paragraph, the overlapping words are 0, and the overlapping rate is 0. Assume that the semantic vector directional consistency obtained by a certain semantic vector calculation method is 0.7 (full score 1). Through weighted calculation (assuming the core word overlap rate weight is 0.4, and the semantic vector directional consistency weight is 0.6), the insertion compatibility is 0.4*(0+0) / 2+0.6*0.7=0.42.

[0061] Step S143: If the insertion compatibility exceeds a preset compatibility threshold, the paragraph boundary of the context paragraph is calibrated, and the paragraph boundary calibration includes deleting the core word sequence repeated with the adjacent paragraph and adjusting the transition conjunctions of the starting sentence of the paragraph.

[0062] The preset compatibility threshold is assumed to be 0.4. Since 0.42 is greater than 0.4, the paragraph boundary of the context paragraph is calibrated. For example, if the starting sentence of the sub-document unit context paragraph is found to be "This special medicine has a unique ingredient", if the adjacent paragraph has mentioned "drug" related content, delete the repeated core word "drug" and adjust the starting sentence to "This special-effect item with a unique ingredient has", and add appropriate transitional conjunctions according to semantic logic to make the connection with the adjacent paragraph more natural.

[0063] Step S144: eliminating semantic redundancy of the calibrated context paragraph to obtain a processed context paragraph.

[0064] Step S1441: extract all sentences in the calibrated context paragraph to generate a sentence set.

[0065] The content of the calibrated context paragraph is "This unique ingredient has a good therapeutic effect. It can effectively relieve the symptoms of heart disease. This effect has been verified in clinical trials." Extract all sentences and generate a sentence set ["This unique ingredient has a good therapeutic effect.", "It can effectively relieve the symptoms of heart disease.", "This effect has been verified in clinical trials."].

[0066] Step S1442: perform semantic role labeling on each sentence in the sentence set, and extract the subject-predicate-object structure and modifying components of the sentence.

[0067] For the first sentence, "The special-effect item of this unique ingredient has a good therapeutic effect.", the semantic role annotation is performed, and the subject-predicate-object structure is "item (subject)-has (predicate)-effect (object)", and the modifying component is "the unique ingredient's" and "good". For the second sentence, "It can effectively relieve the symptoms of heart disease.", the subject-predicate-object structure is "it (subject)-relieve (predicate)-symptoms (object)", and the modifying component is "can effectively" and "heart disease". For the third sentence, "This effect has been verified in clinical trials.", the subject-predicate-object structure is "effect (subject)-verified (predicate)", and the modifying component is "this" and "in clinical trials".

[0068] Step S1443: Calculate the semantic overlap between any two sentences, where the semantic overlap is determined by subject-verb-object structure consistency and modifying component similarity.

[0069] Calculate the semantic overlap between the first sentence and the second sentence, subject-predicate-object structure consistency, subject "item" and "it" (referring to the item) have a certain correlation, consistency is assumed to be 0.6; predicate "have" and "relieve" are different, consistency is 0; object "effect" and "symptom" are different, consistency is 0. Similarity of modifying components, "this unique component" and "can be effective" are assumed to be 0.2. Through weighted calculation (assuming the weight of subject-predicate-object structure consistency is 0.6, and the weight of modifying component similarity is 0.4), the semantic overlap is 0.6*(0.6+0+0) / 3+0.4*0.2=0.12+0.08=0.2. Calculate the semantic overlap between other sentences in a similar way.

[0070] Step S1444: If the semantic overlap between two sentences exceeds a preset overlap threshold, the sentence with a higher core word density is retained and the other sentence is deleted.

[0071] The preset overlap threshold is assumed to be 0.3. Since the semantic overlap between the first sentence and the second sentence is 0.2, which is less than 0.3, no deletion is performed. Assuming that the semantic overlap between other sentences exceeds 0.3, such as two sentences, the core word density is calculated. Assuming that the number of core words in the first sentence is 3, the sentence length is 20 characters, and the core word density is 3 / 20=0.15; the number of core words in the second sentence is 4, the sentence length is 25 characters, and the core word density is 4 / 25=0.16, the second sentence with a higher core word density is retained, and the first sentence is deleted.

[0072] Step S1445: perform a logical sequence check on the remaining sentences, and adjust the sentence arrangement order according to the global structural features of the target parent document so that the semantic coherence index of adjacent sentences reaches a preset coherence threshold, thereby obtaining a processed context paragraph.

[0073] The remaining sentences are checked in logical order according to the global structural features of the target parent document "2.1 Drug Therapy". Assuming that a certain semantic coherence calculation method is used, the sentence order is adjusted so that the semantic coherence index of adjacent sentences reaches the preset coherence threshold (assuming it is 0.5), and finally the processed context paragraph is obtained.

[0074] Step S145: Encapsulate the processed context paragraph into a reconstructed context data block, and add a version identifier and a timestamp to the reconstructed context data block.

[0075] The processed context paragraphs about the specific drug treatment are encapsulated into a reconstructed context data block, and a version identifier and a timestamp are added to the reconstructed context data block.

[0076] Continuing with the heart disease diagnosis and treatment specification review document and its related sub-documents as an example, the processed context paragraphs about specific drug treatments are encapsulated into reconstructed context data blocks. In this process, the integrity and standardization of the encapsulation must be ensured so that the data block can be accurately identified and processed by subsequent operations.

[0077] Next, add a version identifier. The version identifier can be a string combination with predefined rules, which is used to uniquely identify the version information of the reconstruction context data block. For example, the rule for setting the version identifier is "V" plus the last two digits of the year, the two digits of the month, the two digits of the day, and a self-increasing three-digit serial number. Assuming that the current processing time is October 15, 2024, and this is the fifth reconstruction context data block generated that day, the version identifier can be set to "V241015005". Such a version identifier contains both time information and a self-increasing serial number, which is convenient for distinguishing and managing different versions of data blocks in subsequent processes.

[0078] Then add a timestamp, which records the exact time when the reconstruction context data block is generated. The timestamp can be in a common time format, such as a time count in milliseconds. For example, the number of milliseconds from a fixed starting time point (such as 00:00:00 UTC on January 1, 1970) to the time when the data block is generated is used as the timestamp. Assuming that the reconstruction context data block is generated at 14:30:15 and 200 milliseconds on October 15, 2024, through the corresponding time calculation method (for example, converting the year, month, day, hour, minute, second, and millisecond into milliseconds and adding them together), the number of milliseconds from 00:00:00 UTC on January 1, 1970 to this moment is 1718405415200, which is used as the timestamp of the reconstruction context data block.

[0079] Step S150: Generate an optimized version of the target parent document according to the reconstructed context data block, and output the optimized version to a document storage system to trigger a version update operation.

[0080] In the scenario of a document reviewing the diagnostic and treatment standards for heart disease, an optimized version of the target parent document is generated based on the previously generated reconstructed context data block. The specific operation is to accurately insert the reconstructed context data block into the corresponding paragraph level position of the target parent document according to the position determined by the paragraph mapping relationship table. For example, if it was previously determined that the reconstructed context data block about a specific drug treatment should be inserted into a specific position under the "2.1 Drug Treatment" level, the data block is inserted here to form an optimized version of the target parent document.

[0081] Step S151: parsing all reconstruction context data blocks in the optimized version, and extracting the version identifier and timestamp of each reconstruction context data block.

[0082] For the generated optimized version of the target parent document, all the reconstructed context data blocks are parsed. Taking the previously encapsulated reconstructed context data block about a specific drug treatment as an example, according to the encapsulated format and rules, the version identifier "V241015005" and the timestamp "1718405415200" are extracted from it. This extraction operation is performed for each reconstructed context data block to ensure that the accurate version and time information of each data block is obtained.

[0083] Step S152: Acquire the current version metadata of the target parent document from the document storage system, wherein the current version metadata includes a historical version identifier set and a last updated timestamp.

[0084] Assume that the document storage system uses a common storage structure to manage the version information of the document. In the scenario of the review document of the diagnosis and treatment of heart disease, the storage location corresponding to the target parent document is found from the document storage system, and the current version metadata is obtained from it. The historical version identifier set may be a list containing multiple version identifiers, such as ["V241010001", "V241012003", "V241014004"], which represent versions generated at different times. The last updated timestamp records the time when the target parent document was last updated, assuming it is "1718321000000", which is also expressed in milliseconds from 00:00:00 UTC on January 1, 1970.

[0085] Step S153: If the timestamp of the optimized version is later than the last updated timestamp, the optimized version is compared with the current version to generate a difference content report, which includes the number of newly added paragraphs, modified paragraph positions and deleted paragraph identifiers.

[0086] Compare the timestamp "1718405415200" of the reconstructed context data block in the optimized version with the last updated timestamp "1718321000000" obtained from the document storage system. Since "1718405415200" is greater than "1718321000000", that is, the timestamp of the optimized version is later than the last updated timestamp, a difference comparison is required.

[0087] In the process of difference comparison, the number of newly added paragraphs is first determined. For example, through structural analysis of the optimized version and the current version, it is found that a new paragraph is added in the optimized version because a reconstructed context data block about a specific drug treatment is inserted, so the number of newly added paragraphs is 1.

[0088] Next, determine the position of the modified paragraph. Check the impact of inserting the reconstruction context data block on the positions of surrounding paragraphs in the optimized version. Assuming that there were originally paragraphs A and B under the "2.1 Drug Therapy" level, after inserting the reconstruction context data block, the position of paragraph B moved backwards. The specific position of the movement can be determined by calculating the document character position. For example, the original starting position of paragraph B was at the 3000th character in the document. After inserting the reconstruction context data block, the starting position of paragraph B became the 3500th character. Then the modified paragraph position information is recorded as paragraph B moving from the 3000th character to the 3500th character.

[0089] Finally, the paragraph identifier is deleted. After comparison, if no paragraph is deleted in the optimized version, the deleted paragraph identifier is an empty list. If there is a deletion, for example, if a paragraph in the current version is deleted in the optimized version due to content duplication, etc., the identifier of the paragraph is recorded. Assuming that the paragraph identifier is "P3-007", the deleted paragraph identifier is recorded as ["P3-007"]. This information is collated to generate a difference content report.

[0090] Step S154: Generate a version update instruction according to the difference content report, wherein the version update instruction includes incremental update data and a rollback verification code.

[0091] Generate version update instructions based on the difference content report. The incremental update data mainly comes from the reconstruction context data block and the modification information determined in the difference comparison. For example, the incremental update data includes the reconstruction context data block content about a specific drug treatment, as well as the modified paragraph position information (such as moving paragraph B from character 3000 to character 3500).

[0092] The rollback check code is generated to ensure that the previous version can be rolled back when a problem occurs during the update process. The generation of the rollback check code can be based on a predefined algorithm, such as by hashing the key contents of the optimized version and the current version. First, the key parts of the optimized version and the current version (such as the core paragraph content of the document, the version identifier, etc.) are combined together. It is assumed that the combined content is "the main content of the document review of the diagnosis and treatment of heart disease + V241015005 + V241014004", and then it is calculated by a common hash algorithm (such as SHA-256) to obtain a fixed-length hash value, which is assumed to be "0x123abcdef45678901234567890abcdef12345678901234567890abcdef4567890", and this hash value is used as the rollback check code. The incremental update data and the rollback check code are combined together to form a version update instruction.

[0093] Step S155: Send the version update instruction to the document storage system, so that the document storage system replaces the target paragraph according to the incremental update data, and verifies the integrity of the replaced paragraph according to the rollback verification code.

[0094] The generated version update instruction is sent to the document storage system. After receiving the instruction, the document storage system starts to execute the update operation.

[0095] Step S1551: Create a temporary version branch in the document storage system, and write the incremental update data into the temporary version branch.

[0096] The document storage system first creates a temporary version branch in its internal structure. The temporary version branch is similar to an independent storage space for temporarily storing data during the update process. Then the incremental update data is written to the temporary version branch. For example, the incremental update data such as the reconstruction context data block content and paragraph position modification information about a specific drug treatment is accurately written to the storage location corresponding to the temporary version branch to ensure the integrity and accuracy of the data.

[0097] Step S1552: Perform an integrity check on the temporary version branch, wherein the integrity check includes a paragraph identifier continuity check and a semantic logic conflict detection.

[0098] Perform a paragraph identifier continuity check to see if the identifiers of all paragraphs in the temporary version branch are continuous in the expected order. For example, if there are paragraph identifiers "P1-001", "P1-002", and "P1-003" in the temporary version branch, check whether they are continuous in sequence without skipping numbers. If there are skipping numbers, such as missing "P1-002", the paragraph identifier continuity check fails.

[0099] Perform semantic logic conflict detection to analyze whether the semantic logic relationship between each paragraph in the temporary version branch is reasonable. Taking the review document of the diagnosis and treatment specifications for heart disease as an example, check whether the semantic logic of the surrounding paragraphs is coherent after the reconstructed context data block about specific drug treatment is inserted. For example, whether the inserted content is logically consistent with the description of drug treatment in the previous and next paragraphs, and whether there are contradictory statements. For example, if the previous paragraph mentions that the applicable symptom of a certain drug is A, and the inserted reconstructed context data block says that the applicable symptom of the drug is B, and A and B contradict each other, then there is a semantic logic conflict, and the semantic logic conflict detection fails.

[0100] Step S1553: If the integrity check passes, the temporary version branch is merged into the main version branch, and the last updated timestamp is updated.

[0101] If both the paragraph identifier continuity check and the semantic logic conflict detection pass, it means that the data in the temporary version branch is complete and logically reasonable. At this time, the document storage system merges the temporary version branch into the main version branch to update the main version branch to the latest state. At the same time, update the last updated timestamp to the time of the current operation. For example, assuming that the current operation time is 14:35:00 on October 15, 2024, the number of milliseconds converted from 00:00:00 UTC on January 1, 1970 is 1718405700000, and this value is updated as the last updated timestamp of the target parent document in the document storage system.

[0102] Step S1554: If a semantic logic conflict is detected, the write operation of the temporary version branch is canceled according to the rollback verification code, and the manual intervention process is triggered.

[0103] If a semantic logic conflict is detected during the integrity check, the document storage system will cancel the write operation of the temporary version branch based on the previously generated rollback check code. The specific operation is to compare and restore the rollback check code with the relevant data previously saved in the storage system. For example, the rollback check code "0x123abcdef45678901234567890abcdef12345678901234567890abcdef4567890" is used to find the version data before the update operation, and restore the temporary version branch to the state before the update. At the same time, the manual intervention process is triggered to notify relevant personnel (such as document administrators or domain experts) to handle the semantic logic conflict problem. Relevant personnel can analyze the cause of the conflict and make manual adjustments by viewing the difference content report and the data of the temporary version branch to ensure the accuracy and logic of the document content.

[0104] Before obtaining the target parent document and its associated child document set, the method further includes: step S210: configuring a document association rule library, wherein the document association rule library defines matching conditions of parent and child documents, wherein the matching conditions include file format consistency, subject keyword overlap rate, and author authority identifier.

[0105] In the overall process of medical data retrieval, before obtaining the target parent document and its associated child document set, you need to configure the document association rule base. Taking documents related to heart disease diagnosis and treatment as an example, the file format consistency rule requires that the format type of the parent and child documents is the same and the version compatibility identifier is consistent. For example, if the parent document is in PDF format and version 1.5, then the child document must also be in PDF format and the version must be between 1.4 and 1.6 (the set version compatibility range).

[0106] The threshold of the topic keyword overlap rate is dynamically adjusted through the topic distribution statistics of the historical document collection. Suppose that a topic analysis is performed on the past 100 historical documents related to heart disease diagnosis and treatment, and the topic keywords of each document are extracted. For example, the frequently appearing topic keywords in these documents are "heart disease", "treatment", "diagnosis", etc. Through statistical analysis, it is found that the documents related to "heart disease treatment" account for 60%, and the documents related to "heart disease diagnosis" account for 40%. According to these statistical results, the threshold of the topic keyword overlap rate is dynamically adjusted. If the current focus is on document associations in the treatment aspect, the threshold of the topic keyword overlap rate is set to 0.6, that is, the topic keyword overlap rate between the child document and the parent document reaches 0.6 or above to meet the matching conditions.

[0107] Bind the author permission identifier so that the creator identifier of the child document must be included in the list of authorized editors of the parent document. For example, if the list of authorized editors of the parent document is ["DoctorSmith", "DoctorJohnson", "ResearcherLi"], then only the child document with one of these creator identifiers will meet the matching condition.

[0108] Assign a weight coefficient to each matching condition, and determine the matching priority of the parent and child documents based on the weighted sum of the weight coefficients. Assume that the file format consistency weight coefficient is 0.3, the subject keyword overlap rate weight coefficient is 0.5, and the author permission identifier weight coefficient is 0.2. When judging the matching priority of a child document with a parent document, assuming that the file format consistency score of the child document and the parent document is 1 (completely consistent), the subject keyword overlap rate score is 0.7, and the author permission identifier score is 1 (the creator of the child document is in the authorized editor list), the matching priority is 0.3×1+0.5×0.7+0.2×1=0.3+0.35+0.2=0.85. When multiple parent documents meet the matching conditions, select the parent document with the highest matching priority for association.

[0109] Step S220: monitoring the document input channel, and when a new document upload is detected, extracting metadata features of the new document, the metadata features including a creator identifier, a document subject tag, and a format type.

[0110] In the medical data retrieval system, the document input channel is continuously monitored. For example, when a new document related to heart disease is uploaded, the event is captured immediately. Then the metadata features of the new document are extracted. Taking a newly uploaded research report on a new treatment for heart disease as an example, the creator identifier is assumed to be "ResearcherWang", and the document subject tag is extracted by analyzing the document content, which may be "new treatment for heart disease, efficacy evaluation", etc. The format type is detected as PDF format.

[0111] Step S230: matching the metadata features of the new document with the metadata features of the existing parent document according to the document association rule library, and if the match is successful, adding the new document to the child document set of the corresponding parent document.

[0112] Match the metadata features of the newly uploaded document with the metadata features of the parent document in the document association rule library. For example, there is a parent document on heart disease treatment, whose file format is PDF format, version 1.5, and the subject keywords include "heart disease treatment" and "drug treatment", etc. The authorized editor list is ["DoctorSmith", "DoctorJohnson", "ResearcherLi", "ResearcherWang"]. The newly uploaded document format is PDF format, which meets the file format consistency; the subject keyword overlap rate is calculated, and the overlap rate of the new document subject label "new heart disease treatment, efficacy evaluation" and the parent document subject keyword "heart disease treatment" is assumed to be 0.7, which exceeds the set 0.6 threshold; the creator identifier "ResearcherWang" is in the authorized editor list of the parent document. Therefore, the match is successful, and the newly uploaded research report on new heart disease treatment is added to the child document collection of the parent document.

[0113] Step S240: If the match fails, create an independent parent document identifier for the new document and initialize its child document set to empty.

[0114] Suppose a new document about a study of genetic factors of a rare heart disease is uploaded, and there is no matching parent document. Its file format is DOCX, which is inconsistent with the PDF format of most parent documents related to heart disease diagnosis and treatment. Even if the subject keywords partially overlap, the overall matching priority is lower than other parent documents due to the low file format consistency score, and the match fails. At this time, create an independent parent document identifier for this new document, such as "ParentDoc-001", and initialize its child document collection to empty, waiting for subsequent matching child documents to be uploaded.

[0115] Figure 2 A context reconstruction system 100 provided in an embodiment of the present application is shown, including a processor 1001 and a memory 1003 and program code stored in the memory 1003, and the processor 1001 executes the above program code to implement the steps of the context reconstruction method based on parent-child document matching.

[0116] Figure 2The context reconstruction system 100 shown includes: a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, such as through a bus 1002. Optionally, the context reconstruction system 100 may also include a transceiver 1004, which may be used for data interaction between the context reconstruction system and other context reconstruction systems, such as data transmission and / or data reception. It should be noted that in actual scheduling, the transceiver 1004 is not limited to one, and the structure of the context reconstruction system 100 does not constitute a limitation on the embodiments of the present application.

[0117] Processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 1001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0118] The bus 1002 may include a path to transmit information between the above components. The bus 1002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 1002 may be divided into an address bus, a data bus, a control bus, and the like.

[0119] The memory 1003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage medium, other magnetic storage devices, or any other medium that can be used to have or store program code and can be read by a computer, without limitation herein.

[0120] The memory 1003 is used to store program codes for executing the embodiments of the present application, and the execution is controlled by the processor 1001. The processor 1001 is used to execute the program codes stored in the memory 1003 to implement the steps shown in the above method embodiments.

[0121] An embodiment of the present application provides a computer-readable storage medium having program code stored thereon. When the program code is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.

[0122] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the implementation order of these steps is not limited to the order indicated by the arrows. Unless there is clear explanation in this article, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders based on demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages according to the actual implementation scenario, and some or all of these sub-steps or stages may be executed at the same time, and each sub-step or stage in these sub-steps or stages may also be executed at different times respectively. In different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured based on demand, and the embodiment of the present application does not limit this.

[0123] The above is only an optional implementation method for some implementation scenarios of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the scheme of the present application, other similar implementation methods based on the technical ideas of the present application are also within the protection scope of the embodiments of the present application.

Claims

1. A context reconstruction method based on parent-child document matching, characterized in that: The method comprises: Obtaining a target parent document and its associated sub-document set, wherein the sub-document set includes a plurality of sub-document units, each of which includes at least one context paragraph and a corresponding paragraph identifier; Performing hierarchical structure analysis on the target parent document to extract global structural features and local content features of the target parent document, wherein the global structural features include paragraph hierarchical distribution information and title nesting depth information, and the local content features include paragraph core word sequences and semantic coherence indicators; Determine, according to the paragraph identifier of the sub-document unit, a hierarchical affiliation parameter between the sub-document unit and the target parent document, and match the context paragraph of the sub-document unit with a corresponding paragraph level of the target parent document based on the hierarchical affiliation parameter; Based on the global structural features and the local content features, dynamic semantic alignment processing is performed on the context paragraphs of the sub-document unit to generate a reconstructed context data block, wherein the dynamic semantic alignment processing includes paragraph boundary calibration and semantic redundancy elimination; An optimized version of the target parent document is generated according to the reconstructed context data block, and the optimized version is output to a document storage system to trigger a version update operation.

2. The context reconstruction method based on parent-child document matching according to claim 1 is characterized in that: The step of performing hierarchical structure analysis on the target parent document to extract global structural features and local content features of the target parent document includes: Identify a title tag sequence in the target parent document, and generate paragraph level distribution information according to the nesting level of the title tag sequence, wherein the paragraph level distribution information includes the starting position, the ending position and the corresponding title level number of each paragraph; Performing semantic segmentation processing on each paragraph of the target parent document to extract a core word sequence of the paragraph, wherein the core word sequence of the paragraph is composed of words in the paragraph that meet a word frequency threshold and an inverse document frequency threshold; Calculating the semantic coherence index between adjacent paragraphs, wherein the semantic coherence index is obtained by weighting the overlap of core word sequences of adjacent paragraphs and the cosine similarity of semantic vectors; The title nesting depth information is determined according to the maximum value of the title level number, and the paragraph level distribution information, title nesting depth information, paragraph core word sequence and semantic coherence index are written into the structure feature cache and the content feature cache respectively.

3. The context reconstruction method based on parent-child document matching according to claim 1 is characterized in that: The step of determining the hierarchical affiliation parameter between the sub-document unit and the target parent document according to the paragraph identifier of the sub-document unit includes: Parsing the paragraph identifier of the sub-document unit, extracting the parent document reference number, the target paragraph level number and the paragraph position offset in the paragraph identifier; Matching the paragraph level distribution information corresponding to the parent document reference number from the structural feature cache to determine the maximum position offset range allowed by the target paragraph level number; If the paragraph position offset is within the maximum position offset range, then the structural similarity and content overlap between the context paragraph of the sub-document unit and the corresponding paragraph level of the target parent document are calculated, wherein the structural similarity is determined by the consistency of the title level number and the paragraph length ratio, and the content overlap is determined by the number of overlapping words in the core word sequence of the paragraph and the change in the semantic coherence index; A hierarchical affiliation parameter is generated according to the weighted sum of the structural similarity and the content overlap, and if the hierarchical affiliation parameter exceeds a preset affiliation threshold, the sub-document unit is marked as a valid matching unit.

4. The context reconstruction method based on parent-child document matching according to claim 3 is characterized in that: The matching of the context paragraph of the sub-document unit with the corresponding paragraph level of the target parent document based on the level affiliation parameter includes: For each of the sub-document units marked as valid matching units, extracting the paragraph core word sequence and semantic coherence index of the paragraph level corresponding to the target parent document from the content feature cache; Performing synonym replacement detection on the context paragraph of the sub-document unit, generating a replacement word list, and updating the paragraph core word sequence according to the replacement word list; Calculating the dynamic overlap between the updated paragraph core word sequence and the paragraph core word sequence at the paragraph level corresponding to the target parent document, wherein the dynamic overlap is obtained by adjusting the word order consistency and the word meaning similarity; If the dynamic overlap degree is greater than or equal to a preset reconstruction threshold, the context paragraph of the sub-document unit is positionally bound to the corresponding paragraph level of the target parent document to generate a paragraph mapping relationship table.

5. The context reconstruction method based on parent-child document matching according to claim 4 is characterized in that: The step of performing dynamic semantic alignment processing on the context paragraphs of the sub-document unit based on the global structural features and the local content features to generate a reconstructed context data block includes: According to the paragraph mapping relationship table, determining the paragraph level position at which the context paragraph of the sub-document unit needs to be inserted into the target parent document; Obtaining the semantic coherence index of the adjacent paragraphs at the paragraph level position from the structural feature cache, and calculating the insertion compatibility between the context paragraph and the adjacent paragraph, wherein the insertion compatibility is determined by the core word overlap rate and semantic vector direction consistency of the previous and next paragraphs; If the insertion compatibility exceeds a preset compatibility threshold, the paragraph boundary calibration is performed on the context paragraph, wherein the paragraph boundary calibration includes deleting the core word sequence repeated with the adjacent paragraph and adjusting the transition conjunctions of the starting sentence of the paragraph; Eliminate semantic redundancy on the calibrated context paragraphs to obtain processed context paragraphs; The processed context paragraph is encapsulated into a reconstructed context data block, and a version identifier and a timestamp are added to the reconstructed context data block.

6. The context reconstruction method based on parent-child document matching according to claim 5 is characterized in that: The semantic redundancy elimination is performed on the calibrated context paragraph to obtain a processed context paragraph, including: Extract all sentences in the calibrated context paragraph to generate a sentence set; Performing semantic role labeling on each sentence in the sentence set, and extracting the subject-predicate-object structure and modifying components of the sentence; Calculate the semantic overlap between any two sentences, where the semantic overlap is determined by subject-verb-object structural consistency and modifier similarity; If the semantic overlap between two sentences exceeds a preset overlap threshold, the sentence with a higher density of core words is retained and the other sentence is deleted; The remaining sentences are checked for logical order, and the sentence arrangement order is adjusted according to the global structural features of the target parent document so that the semantic coherence index of adjacent sentences reaches a preset coherence threshold, thereby obtaining a processed context paragraph.

7. The context reconstruction method based on parent-child document matching according to claim 1 is characterized in that: The step of outputting the optimized version to a document storage system to trigger a version update operation includes: Parsing all reconstruction context data blocks in the optimized version, and extracting a version identifier and a timestamp of each reconstruction context data block; Acquire the current version metadata of the target parent document from the document storage system, wherein the current version metadata includes a historical version identifier set and a last updated timestamp; If the timestamp of the optimized version is later than the last updated timestamp, the optimized version is compared with the current version to generate a difference content report, which includes the number of newly added paragraphs, the positions of modified paragraphs, and the identifiers of deleted paragraphs; Generate a version update instruction according to the difference content report, wherein the version update instruction includes incremental update data and a rollback verification code; The version update instruction is sent to the document storage system, so that the document storage system replaces the target paragraph according to the incremental update data, and verifies the integrity of the replaced paragraph according to the rollback verification code.

8. The context reconstruction method based on parent-child document matching according to claim 7 is characterized in that: The sending of the version update instruction to the document storage system so that the document storage system replaces the target paragraph according to the incremental update data and verifies the integrity of the replaced paragraph according to the rollback verification code includes: Creating a temporary version branch in the document storage system, and writing the incremental update data into the temporary version branch; Performing an integrity check on the temporary version branch, the integrity check including a section identifier continuity check and a semantic logic conflict detection; If the integrity check passes, the temporary version branch is merged into the main version branch, and the last updated timestamp is updated; If a semantic logic conflict is detected, the write operation of the temporary version branch is undone according to the rollback verification code, and a manual intervention process is triggered.

9. The context reconstruction method based on parent-child document matching according to claim 1, characterized in that: Before obtaining the target parent document and its associated child document set, the method further includes: Configuring a document association rule library, wherein the document association rule library defines matching conditions of parent and child documents, wherein the matching conditions include file format consistency, subject keyword overlap rate, and author authority identifier; Monitor the document input channel, and when a new document is detected to be uploaded, extract metadata features of the new document, wherein the metadata features include a creator identifier, a document subject tag, and a format type; Matching the metadata features of the new document with the metadata features of the existing parent document according to the document association rule library, and if the match is successful, adding the new document to the child document set of the corresponding parent document; If the match fails, create an independent parent document identifier for the new document and initialize its child document collection to be empty; The configuration document association rule base includes: Define a file format consistency rule, wherein the file format consistency rule requires that the format type of the parent and child documents is the same and the version compatibility identifier is consistent; Setting a topic keyword overlap rate threshold, wherein the topic keyword overlap rate threshold is dynamically adjusted according to the topic distribution statistics of the historical document collection; Bind author permission identifiers so that the creator identifier of a child document must be included in the list of authorized editors of the parent document; Assign a weight coefficient to each matching condition, and determine the matching priority of the parent and child documents based on the weighted sum of the weight coefficients; When multiple parent documents meet the matching conditions, the parent document with the highest matching priority is selected for association.

10. A context reconstruction system, characterized in that: It includes a processor and a computer-readable storage medium, wherein the computer-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by the processor, the context reconstruction method based on parent-child document matching described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Method for analyzing medical insurance data document and automatically incremental database

    CN116894029A

  • Medical data document classification and marking system

    CN119621972A

  • LLM-based document structuring automatic processing method and system

    CN119782503A

  • Document consistency comparison method based on semantic analysis and keyword driving

    CN119886103A

  • Method, system and computer program product for management and review of product consumer medicine information

    US20250005828A1

Cited By

  • AI-driven efficient data redundancy detection and cleaning method and system

    CN120408042A